Method, system for generating compound intervention scheme based on pre-trained model and application thereof

By using a SEMO feature quantification method combining PPI networks and compounds, the problem of the comprehensive impact of compounds on the human immune system was solved, providing personalized compound intervention programs and improving the treatment effect of novel coronavirus infection.

CN117766054BActive Publication Date: 2026-05-08BEIJING DEEP METHYL HEALTH TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING DEEP METHYL HEALTH TECH CO LTD
Filing Date
2023-02-21
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively study the combined effects of multiple compounds on the human immune system, thus hindering the role of these compounds in the management of novel coronavirus infection.

Method used

By defining the SEMO characteristics formed by PPI networks and compound combinations, and using compound action quantification methods combined with DNA methylation or transcriptome data, SEMO values ​​are calculated to screen phenotype-related compound combinations and provide personalized intervention plans.

Benefits of technology

This enables the quantitative characterization of compound effects, obtains biomarkers and therapeutic targets related to disease phenotypes, provides personalized treatment plans, and improves the treatment efficacy for immune-related diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_4
    Figure SMS_4
  • Figure SMS_7
    Figure SMS_7
Patent Text Reader

Abstract

The application discloses a method and system for generating a compound intervention scheme based on a pre-trained model and application thereof, wherein the compound refers to a nutritional supplement, a traditional Chinese medicine, a natural product with a medicinal value to be explored, or an old drug with a new use compound, and particularly refers to a method and system for generating a personalized nutritional supplement combination intervention scheme. The method of the application can successfully quantify and characterize the action of the compound, and the method of the application can also be applied to guide the medication of patients or a clinical nutrition scheme, for example, a patient infected with a novel coronavirus.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biotechnology, specifically relating to methods, systems, and applications for generating compound intervention schemes based on pre-trained models. Background Technology

[0002] The relationship between compounds (such as drugs and nutritional supplements) and the human immune system is highly complex, and focusing on single compounds has hindered research into their role in the management of novel coronavirus infection. While drugs, nutritional supplements, and other compounds have significant effects on the immune system, it is difficult to conduct trials that simultaneously study multiple compounds.

[0003] Therefore, there is an urgent need in this field for a method to provide subjects with compound intervention programs. Summary of the Invention

[0004] As mentioned above, there is an urgent need in the field for a method to provide subjects with compound intervention programs.

[0005] The inventors defined the SEMO (Semo) features formed by PPI (Proteinizing Ingredient) networks and compound combinations. By detecting individual DNA methylation or transcriptome data, they calculated the PPI network-compound SEMO value, achieving quantitative characterization of compound effects. Using this quantitative characterization method, they analyzed phenotypic data to obtain phenotype-related SEMO features, which were then used to guide medication regimens for subjects. Thus, this invention was achieved.

[0006] In a first aspect, the present invention provides a method for the quantitative characterization of compound interactions, the method comprising the following steps:

[0007] (1) Obtain a set of targets for a single compound, including but not limited to drugs and nutritional supplements;

[0008] (2) Obtain all genes that interact with a single gene, and use the single gene and all genes that interact with it as the gene set of the protein interaction network of the single gene.

[0009] (3) Calculate the SEMO value of the target set of the single compound and the gene set of the protein-protein interaction network of the single gene, wherein the SEMO value is calculated by the following formula:

[0010]

[0011] in, The mean of the eigenvalues ​​of the target gene. S represents the mean of the eigenvalues ​​of non-target genes. x 2 S is the variance of the eigenvalues ​​of the target gene. y2 denoted as , where is the variance of the eigenvalues ​​of non-target genes, n is the number of target genes, and m is the number of non-target genes; the target genes are the intersection of the gene set of the protein-protein interaction network and the target set of the compound, and the non-target genes are genes that belong to the gene set of the protein-protein interaction network but do not belong to the target set of the compound.

[0012] The SEMO value is a characterization of the effect of a single compound by using the protein-protein interaction network of the single gene.

[0013] In a second aspect, the present invention provides a system for the quantitative characterization of compound interactions, the system implementing the method of the first aspect, the system comprising:

[0014] The gene acquisition module is used to acquire the target set of a single compound and the gene set of the protein-protein interaction network of a single gene, thereby obtaining the target gene and non-target gene.

[0015] The data acquisition module is used to acquire feature values ​​of the target genes and non-target genes from the subjects;

[0016] The calculation module is used to calculate the SEMO value based on the feature values ​​of the target gene and non-target genes.

[0017] In a third aspect, the present invention provides a medium for the quantitative characterization of compound interactions, the medium comprising a program for implementing the method of the first aspect, the medium comprising:

[0018] (1) Obtain the target set of a single compound and the gene set of the protein-protein interaction network of a single gene, and then obtain the target gene and non-target gene.

[0019] (2) Read the characteristic values ​​of the target genes and non-target genes of the subject;

[0020] (3) Calculate the semo value based on the feature values ​​of the target gene and non-target gene.

[0021] In a fourth aspect, the present invention provides a method for providing a subject with a compound intervention regimen, the method using the method described in the first aspect, the method comprising the following steps:

[0022] (1) Obtain protein interaction network data and compound target data;

[0023] (2) Based on the protein interaction network data and compound target data, protein interaction network and compound combination are screened, i.e., SEMO features, wherein the protein interaction network and compound combination is the intersection of the gene set of the protein interaction network and the gene set of the compound target is significant.

[0024] (3) Based on the selected SEMO features, an evaluation model for the phenotypic data is established using the phenotypic data, wherein the phenotypic data is a feature value × sample matrix;

[0025] The establishment of the evaluation model includes the following steps:

[0026] (i) Based on the selected SEMO features, obtain the SEMO values ​​of the SEMO features in the phenotypic data, and finally obtain the SEMO feature × sample matrix;

[0027] (ii) Based on the semo features × sample matrix in step (i), obtain semo features that are significantly correlated with the phenotype, for example, through machine learning algorithms;

[0028] (4) For the SEMO feature that is significantly correlated with the phenotype, calculate the SEMO value of the subject and provide the subject with a compound intervention program based on the SEMO value.

[0029] In a fifth aspect, the present invention provides a system for providing a compound intervention to a subject, the system implementing the method of the fourth aspect, the system comprising:

[0030] Data acquisition unit, used to acquire protein interaction network data, compound target data, and phenotypic data;

[0031] The semo filter is used to filter valid semo features.

[0032] A computational analyzer is used to calculate the SEMO value of the effective SEMO feature in the phenotypic data, obtain the SEMO feature that is significantly correlated with the phenotypic, and provide a compound intervention plan for the subject based on the SEMO value of the SEMO feature that is significantly correlated with the phenotypic.

[0033] In a sixth aspect, the present invention provides a medium for providing a compound intervention protocol to a subject, the medium comprising procedures for implementing the method of the fourth aspect, the medium comprising:

[0034] (1) Obtain protein interaction network data, compound target data, and phenotypic data;

[0035] (2) Filter valid SEMO features;

[0036] (3) Calculate the SEMO value of the effective SEMO feature in the phenotypic data, obtain the SEMO feature that is significantly correlated with the phenotypic, and provide the subject with a compound intervention plan based on the SEMO value of the SEMO feature that is significantly correlated with the phenotypic.

[0037] In a seventh aspect, the present invention provides a method for providing a compound intervention regimen to a subject infected with the novel coronavirus, the method using the method described in the fourth aspect, wherein the SEM features positively correlated with the novel coronavirus infection outcome include:

[0038] The single gene is FUT4, and the compound is L-Alanine;

[0039] SEMO features negatively associated with COVID-19 infection outcomes include:

[0040] The single gene is IL6, the compound is L-Tryptophan; and / or

[0041] The single gene is CXCL8, and the compound is Vitamin E; and / or

[0042] The single gene is ITGA2, the compound is L-Cystine; and / or

[0043] The single gene is CXCL8, the compound is Calcitriol; and / or

[0044] The single gene is VCAM1, the compound is L-Citrulline; and / or

[0045] The single gene is CTSB, the compound is Vitamin A; and / or

[0046] The single gene is VCAM1, and the compound is Spermine; and / or

[0047] The single gene is ITGAM, the compound is L-Tryptophan; and / or

[0048] The single gene is ITGA2, the compound is L-Citrulline; and / or

[0049] The single gene is ITGAM, the compound is Glycine betaine; and / or

[0050] The single gene is MMP9, the compound is L-Aspartic Acid; and / or

[0051] The single gene is TFRC, the compound is L-Tryptophan; and / or

[0052] The single gene is LCK, the compound is L-Tyrosine; and / or

[0053] The single gene is ITGAM, the compound is L-Arginine; and / or

[0054] The single gene is GZMB, and the compound is Vitamin A; and / or

[0055] The single gene is ITGAM, the compound is L-Isoleucine; and / or

[0056] The single gene is CXCL8, the compound is L-Tryptophan; and / or

[0057] The single gene is MMP9, the compound is L-Citrulline; and / or

[0058] The single gene is ITGAM, the compound is L-Valine; and / or

[0059] The single gene is MMP9, the compound is L-Tryptophan; and / or

[0060] The single gene is VCAM1, the compound is Clopidogrel; and / or

[0061] The single gene is CD40LG, the compound is Niacin; and / or

[0062] The single gene is CSF2, and the compound is Choline; and / or

[0063] The single gene is IL6, the compound is L-Citrulline; and / or

[0064] The single gene is IL15, and the compound is Vitamin A; and / or

[0065] The single gene is MMP9, the compound is Glycine betaine; and / or

[0066] The single gene is MMP2, and the compound is Melatonin; and / or

[0067] The single gene is MMP2, the compound is L-Valine; and / or

[0068] The single gene is ALB, and the compound is N-Acetyl-D-glucosamine; and / or

[0069] The single gene is MMP2, the compound is Glycine betaine; and / or

[0070] The single gene is CD40LG, the compound is L-Tryptophan; and / or

[0071] The single gene is VCAM1, the compound is L-Tryptophan; and / or

[0072] The single gene is CXCL8, and the compound is N-Acetyl-D-glucosamine; and / or

[0073] The single gene is CSF3, the compound is Calcitriol; and / or

[0074] The single gene is ITGAM, the compound is L-Citrulline; and / or

[0075] The single gene is MMP9, and the compound is Spermine; and / or

[0076] The single gene is TLR2, the compound is Lipoic Acid; and / or

[0077] The single gene is CD40, the compound is Lipoic Acid; and / or

[0078] The single gene is ITGAM, the compound is Clopidogrel; and / or

[0079] The single gene is CXCL8, the compound is L-Citrulline; and / or

[0080] The single gene is SPI1, the compound is L-Cysteine; and / or

[0081] The single gene is IL4, and the compound is Choline; and / or

[0082] The single gene is MMP2, and the compound is Spermine; and / or

[0083] The single gene is STAT3, the compound is L-Citrulline; and / or

[0084] The single gene is ITGAX, the compound is Choline; and / or

[0085] The single gene is IL4, and the compound is Lipoic Acid; and / or

[0086] The single gene is MMP2, the compound is L-Aspartic Acid; and / or

[0087] The single gene is MMP2, the compound is L-Tryptophan; and / or

[0088] The single gene is CD8A, the compound is Choline; and / or

[0089] The single gene is MMP2, the compound is L-Citrulline; and / or

[0090] The single gene is STAT1, the compound is Lipoic Acid; and / or

[0091] The single gene is CD8A, and the compound is Lipoic Acid;

[0092] The novel coronavirus infection outcome indicates the efficacy of treatment with the compound in the subjects.

[0093] In an eighth aspect, the present invention provides a method for providing a compound intervention regimen to a subject infected with the novel coronavirus, the method using the method described in the fourth aspect.

[0094] The single gene is CD68:

[0095] SEM features positively correlated with COVID-19 infection outcomes include:

[0096] The compounds are Succinic acid, Magnesium, and Niacin.

[0097] SEMO features negatively associated with COVID-19 infection outcomes include:

[0098] The compound is Lipoic Acid;

[0099] The single gene in question is CD4:

[0100] SEM features positively correlated with COVID-19 infection outcomes include:

[0101] The compounds are Calcium and Magnesium.

[0102] SEMO features negatively associated with COVID-19 infection outcomes include:

[0103] The compounds are Tretinoin and Calcitriol;

[0104] The single gene is IL6:

[0105] SEM features positively correlated with COVID-19 infection outcomes include:

[0106] The compound is Glycine.

[0107] SEMO features negatively associated with COVID-19 infection outcomes include:

[0108] The compounds are L-Citrulline, L-Valine, and L-Tryptophan;

[0109] The single gene is CXCL8:

[0110] SEMO features negatively associated with COVID-19 infection outcomes include:

[0111] The compounds are Glucosamine, Calcitriol, Tretinoin, and L-Citrulline;

[0112] The single gene in question is FOXP3:

[0113] SEM features positively correlated with COVID-19 infection outcomes include:

[0114] The compounds are Vitamin E, Succinic acid, Magnesium, and Glutathione.

[0115] The novel coronavirus infection outcome indicates the efficacy of treatment with the compound in the subjects.

[0116] The beneficial effects of this invention are that it can successfully quantify and characterize the effects of compounds, thereby obtaining biomarkers, therapeutic targets, candidate drugs, and personalized treatment plans related to disease phenotypes; for immune-related diseases such as novel coronavirus infection, the method of this invention can predict candidate drugs and nutritional supplements for treating the disease, thereby better treating and caring for patients with immune diseases. Attached Figure Description

[0117] To more clearly illustrate the technical solutions in the specific embodiments of the present invention, the accompanying drawings in the specific embodiments will be briefly described below.

[0118] Figure 1 A flowchart of the method of the present invention and its specific applications are shown.

[0119] Figure 2 The SEMO characteristics of the severe group vs. the mild group in Data 1 of Example 2 of the present invention are shown, wherein the compounds in the SEMO characteristics are nutritional supplements.

[0120] Figure 3 The ROC curves of the model trained using the lasso machine learning algorithm in Embodiment 3 of the present invention are shown in the training set and the independent validation set.

[0121] Figure 4The diagram shows a PPI network with high frequency in the SEMO features of the top 200 cases in the severe group vs. mild group in Data 1 of Embodiment 4 of the present invention, which shows a significant difference.

[0122] Figure 5 The data in Example 5 of the present invention shows the nutritional supplements with higher frequency of SEM features in the top 200 of the severe group vs. the mild group, which showed significant differences.

[0123] Figure 6 The SEMO characteristics of the severe group vs. the mild group in Data 1 of Example 6 of the present invention are shown, wherein the compounds in the SEMO characteristics are FDA-approved drugs.

[0124] Figure 7 The SEMO characteristics of the severe group vs. the mild group in Data 1 of Example 7 of the present invention are shown, wherein the compounds in the SEMO characteristics are herbal compounds.

[0125] Figure 8 The data from Example 7 of this invention show the high frequency of herbal compound components in the SEMO features of the top 200 patients in the severe group vs. mild group who showed significant differences.

[0126] Figure 9 The SEM risk profiles of three individuals (including one mild case and two severe cases) in Embodiment 8 of the present invention are shown.

[0127] Figure 10 The personal SEMO risk profile of a critically ill patient after ranking is shown in Embodiment 8 of the present invention.

[0128] Figure 11 The semo features associated with the PPI network CD68.n in Data 1 of Embodiment 9 of the present invention show significant differences between the severe group and the mild group. Detailed Implementation

[0129] The present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the following description is merely illustrative and is not intended to limit the scope of the invention; the scope of protection of the invention is defined by the appended claims. Furthermore, those skilled in the art will understand that modifications can be made to the technical solutions of the present invention without departing from its spirit and intent. Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art.

[0130] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the subject matter pertains. Before a detailed description of the invention, the following definitions are provided to better understand it.

[0131] In the context of this invention, many embodiments use the expressions "comprising," "including," or "basically / mainly composed of." The expressions "comprising," "including," or "basically / mainly composed of" are generally understood as open-ended expressions, indicating that they include not only the elements, components, parts, or method steps specifically listed after the expression, but also other elements, components, parts, or method steps. However, in this document, the expressions "comprising," "including," or "basically / mainly composed of" can also be understood as closed-ended expressions in certain cases, indicating that they only include the elements, components, parts, or method steps specifically listed after the expression, and do not include any other elements, components, parts, or method steps. In this case, the expression is equivalent to the expression "composed of."

[0132] As previously stated, the present invention aims to provide a method for providing a compound intervention program to a subject.

[0133] Therefore, in a first aspect, the present invention provides a method for the quantitative characterization of compound interactions, the method comprising the following steps:

[0134] (1) Obtain a set of targets for a single compound, including but not limited to drugs and nutritional supplements;

[0135] (2) Obtain all genes that interact with a single gene, and use the single gene and all genes that interact with it as the gene set of the protein interaction network of the single gene.

[0136] (3) Calculate the SEMO value of the target set of the single compound and the gene set of the protein-protein interaction network of the single gene, wherein the SEMO value is calculated by the following formula:

[0137]

[0138] in, The mean of the eigenvalues ​​of the target gene. S represents the mean of the eigenvalues ​​of non-target genes. x 2 S is the variance of the eigenvalues ​​of the target gene. y 2denoted as , where is the variance of the eigenvalues ​​of non-target genes, n is the number of target genes, and m is the number of non-target genes; the target genes are the intersection of the gene set of the protein-protein interaction network and the target set of the compound, and the non-target genes are genes that belong to the gene set of the protein-protein interaction network but do not belong to the target set of the compound.

[0139] The SEMO value is a characterization of the effect of a single compound by using the protein-protein interaction network of the single gene.

[0140] In one embodiment, the feature value is selected from, but is not limited to, gene methylation feature values ​​and expression values. The methylation feature value is the average of the methylation beta values ​​or the SIMPO value of all methylation sites in the gene body region. In a preferred embodiment, the methylation feature value is the average of the methylation beta values ​​or the SIMPO value of all methylation sites in the gene promoter region.

[0141] The SIMPO value is calculated as follows: The inventors used a previously developed SIMPO method for EWAS (Epigenome-wide Association Study) research. This method converts the beta value of methylation sites into a SIMPO value for a gene. The principle of this method is that the difference in DNA methylation between the gene body and the promoter region is significantly correlated with gene expression, with a correlation coefficient as high as 0.67, suggesting that the difference in methylation between the gene body and the promoter region is a biologically significant feature that can be used to predict gene expression. Specifically, high-order features at the gene level are extracted based on the low-level CpG site methylation data. After this conversion, the beta values ​​of multiple methylation sites in a gene are transformed into a single SIMPO value for each gene, which is correlated with the gene's expression activity.

[0142] In one implementation, the number of genes in the gene set of the protein-protein interaction network of the single gene is ≥20.

[0143] In a second aspect, the present invention provides a system for the quantitative characterization of compound interactions, the system implementing the method of the first aspect, the system comprising:

[0144] The gene acquisition module is used to acquire the target set of a single compound and the gene set of the protein-protein interaction network of a single gene, thereby obtaining the target gene and non-target gene.

[0145] The data acquisition module is used to acquire feature values ​​of the target genes and non-target genes from the subjects;

[0146] The calculation module is used to calculate the SEMO value based on the feature values ​​of the target gene and non-target genes.

[0147] In a third aspect, the present invention provides a medium for the quantitative characterization of compound interactions, the medium comprising a program for implementing the method of the first aspect, the medium comprising:

[0148] (1) Obtain the target set of a single compound and the gene set of the protein-protein interaction network of a single gene, and then obtain the target gene and non-target gene.

[0149] (2) Read the characteristic values ​​of the target genes and non-target genes of the subject;

[0150] (3) Calculate the semo value based on the feature values ​​of the target gene and non-target gene.

[0151] In a fourth aspect, the present invention provides a method for providing a subject with a compound intervention regimen, the method using the method described in the first aspect, the method comprising the following steps:

[0152] (1) Obtain protein interaction network data and compound target data;

[0153] (2) Based on the protein interaction network data and compound target data, protein interaction network and compound combination are screened, i.e., SEMO features, wherein the protein interaction network and compound combination is the intersection of the gene set of the protein interaction network and the gene set of the compound target is significant.

[0154] (3) Based on the selected SEMO features, an evaluation model for the phenotypic data is established using the phenotypic data, wherein the phenotypic data is a feature value × sample matrix;

[0155] The establishment of the evaluation model includes the following steps:

[0156] (i) Based on the selected SEMO features, obtain the SEMO values ​​of the SEMO features in the phenotypic data, and finally obtain the SEMO feature × sample matrix;

[0157] (ii) Based on the semo features × sample matrix in step (i), obtain semo features that are significantly correlated with the phenotype, for example, through machine learning algorithms;

[0158] (4) For the SEMO feature that is significantly correlated with the phenotype, calculate the SEMO value of the subject and provide the subject with a compound intervention program based on the SEMO value.

[0159] In a fifth aspect, the present invention provides a system for providing a compound intervention to a subject, the system implementing the method of the fourth aspect, the system comprising:

[0160] Data acquisition unit, used to acquire protein interaction network data, compound target data, and phenotypic data;

[0161] The semo filter is used to filter valid semo features.

[0162] A computational analyzer is used to calculate the SEMO value of the effective SEMO feature in the phenotypic data, obtain the SEMO feature that is significantly correlated with the phenotypic, and provide a compound intervention plan for the subject based on the SEMO value of the SEMO feature that is significantly correlated with the phenotypic.

[0163] In a sixth aspect, the present invention provides a medium for providing a compound intervention protocol to a subject, the medium comprising procedures for implementing the method of the fourth aspect, the medium comprising:

[0164] (1) Obtain protein interaction network data, compound target data, and phenotypic data;

[0165] (2) Filter valid SEMO features;

[0166] (3) Calculate the SEMO value of the effective SEMO feature in the phenotypic data, obtain the SEMO feature that is significantly correlated with the phenotypic, and provide the subject with a compound intervention plan based on the SEMO value of the SEMO feature that is significantly correlated with the phenotypic.

[0167] In a seventh aspect, the present invention provides a method for providing a compound intervention regimen to a subject infected with the novel coronavirus, the method using the method described in the fourth aspect, wherein the SEM features positively correlated with the novel coronavirus infection outcome include:

[0168] The single gene is FUT4, and the compound is L-Alanine;

[0169] SEMO features negatively associated with COVID-19 infection outcomes include:

[0170] The single gene is IL6, the compound is L-Tryptophan; and / or

[0171] The single gene is CXCL8, and the compound is Vitamin E; and / or

[0172] The single gene is ITGA2, the compound is L-Cystine; and / or

[0173] The single gene is CXCL8, the compound is Calcitriol; and / or

[0174] The single gene is VCAM1, the compound is L-Citrulline; and / or

[0175] The single gene is CTSB, the compound is Vitamin A; and / or

[0176] The single gene is VCAM1, and the compound is Spermine; and / or

[0177] The single gene is ITGAM, the compound is L-Tryptophan; and / or

[0178] The single gene is ITGA2, the compound is L-Citrulline; and / or

[0179] The single gene is ITGAM, the compound is Glycine betaine; and / or

[0180] The single gene is MMP9, the compound is L-Aspartic Acid; and / or

[0181] The single gene is TFRC, the compound is L-Tryptophan; and / or

[0182] The single gene is LCK, the compound is L-Tyrosine; and / or

[0183] The single gene is ITGAM, the compound is L-Arginine; and / or

[0184] The single gene is GZMB, and the compound is Vitamin A; and / or

[0185] The single gene is ITGAM, the compound is L-Isoleucine; and / or

[0186] The single gene is CXCL8, the compound is L-Tryptophan; and / or

[0187] The single gene is MMP9, the compound is L-Citrulline; and / or

[0188] The single gene is ITGAM, the compound is L-Valine; and / or

[0189] The single gene is MMP9, the compound is L-Tryptophan; and / or

[0190] The single gene is VCAM1, the compound is Clopidogrel; and / or

[0191] The single gene is CD40LG, the compound is Niacin; and / or

[0192] The single gene is CSF2, and the compound is Choline; and / or

[0193] The single gene is IL6, the compound is L-Citrulline; and / or

[0194] The single gene is IL15, and the compound is Vitamin A; and / or

[0195] The single gene is MMP9, the compound is Glycine betaine; and / or

[0196] The single gene is MMP2, and the compound is Melatonin; and / or

[0197] The single gene is MMP2, the compound is L-Valine; and / or

[0198] The single gene is ALB, and the compound is N-Acetyl-D-glucosamine; and / or

[0199] The single gene is MMP2, the compound is Glycine betaine; and / or

[0200] The single gene is CD40LG, the compound is L-Tryptophan; and / or

[0201] The single gene is VCAM1, the compound is L-Tryptophan; and / or

[0202] The single gene is CXCL8, and the compound is N-Acetyl-D-glucosamine; and / or

[0203] The single gene is CSF3, the compound is Calcitriol; and / or

[0204] The single gene is ITGAM, the compound is L-Citrulline; and / or

[0205] The single gene is MMP9, and the compound is Spermine; and / or

[0206] The single gene is TLR2, the compound is Lipoic Acid; and / or

[0207] The single gene is CD40, the compound is Lipoic Acid; and / or

[0208] The single gene is ITGAM, the compound is Clopidogrel; and / or

[0209] The single gene is CXCL8, the compound is L-Citrulline; and / or

[0210] The single gene is SPI1, the compound is L-Cysteine; and / or

[0211] The single gene is IL4, and the compound is Choline; and / or

[0212] The single gene is MMP2, and the compound is Spermine; and / or

[0213] The single gene is STAT3, the compound is L-Citrulline; and / or

[0214] The single gene is ITGAX, the compound is Choline; and / or

[0215] The single gene is IL4, and the compound is Lipoic Acid; and / or

[0216] The single gene is MMP2, the compound is L-Aspartic Acid; and / or

[0217] The single gene is MMP2, the compound is L-Tryptophan; and / or

[0218] The single gene is CD8A, the compound is Choline; and / or

[0219] The single gene is MMP2, the compound is L-Citrulline; and / or

[0220] The single gene is STAT1, the compound is Lipoic Acid; and / or

[0221] The single gene is CD8A, and the compound is Lipoic Acid;

[0222] The novel coronavirus infection outcome indicates the efficacy of treatment with the compound in the subjects.

[0223] This section uses "CD8A-Choline" as an example to illustrate the method of providing compound intervention to subjects infected with the novel coronavirus. Specifically, all proteins that interact with CD8A are obtained, resulting in gene set A of the CD8A protein interaction network; the target set B of the compound Choline is obtained; the intersection of set A and set B is defined as the target genes, and genes belonging to set A but not to set B are non-target genes; the eigenvalues ​​of the target and non-target genes are used to calculate the SEMO value corresponding to "CD8A-Choline". If the SEMO value corresponding to "CD8A-Choline" for subject 1 is greater than that for subject 2, since "CD8A-Choline" is a SEMO feature negatively correlated with the outcome of novel coronavirus infection, subject 2 is determined to be more suitable for Choline treatment after infection with the novel coronavirus than subject 1.

[0224] In an eighth aspect, the present invention provides a method for providing a compound intervention regimen to a subject infected with the novel coronavirus, the method using the method described in the fourth aspect.

[0225] The single gene is CD68:

[0226] SEM features positively correlated with COVID-19 infection outcomes include:

[0227] The compounds are Succinic acid, Magnesium, and Niacin.

[0228] SEMO features negatively associated with COVID-19 infection outcomes include:

[0229] The compound is Lipoic Acid;

[0230] The single gene in question is CD4:

[0231] SEM features positively correlated with COVID-19 infection outcomes include:

[0232] The compounds are Calcium and Magnesium.

[0233] SEMO features negatively associated with COVID-19 infection outcomes include:

[0234] The compounds are Tretinoin and Calcitriol;

[0235] The single gene is IL6:

[0236] SEM features positively correlated with COVID-19 infection outcomes include:

[0237] The compound is Glycine.

[0238] SEMO features negatively associated with COVID-19 infection outcomes include:

[0239] The compounds are L-Citrulline, L-Valine, and L-Tryptophan;

[0240] The single gene is CXCL8:

[0241] SEMO features negatively associated with COVID-19 infection outcomes include:

[0242] The compounds are Glucosamine, Calcitriol, Tretinoin, and L-Citrulline;

[0243] The single gene in question is FOXP3:

[0244] SEM features positively correlated with COVID-19 infection outcomes include:

[0245] The compounds are Vitamin E, Succinic acid, Magnesium, and Glutathione.

[0246] The novel coronavirus infection outcome indicates the efficacy of treatment with the compound in the subjects.

[0247] Figure 1 This paper describes a complete data analysis and application process. The process includes three main stages: pre-training stage 1, pre-training stage 2, and application stage. In stage 1, gene set 1 (protein network) and gene set 2 (target genes of compounds) are input to screen effective PPI-compound combinations (i.e., SEMO features). In stage 2, the method of this invention is used to screen SEMO features significantly related to the phenotype (disease) using phenotypic (disease)-related omics data. These SEMO features can be used to obtain biomarkers, rank disease treatment targets, and rank compounds (drugs, nutritional supplements, etc.). Furthermore, the above process can also be applied to calculate individual disease risk, thereby obtaining individual disease treatment plans. In a specific implementation, after screening SEMO features significantly related to the phenotype (disease) using the above method, they are stored in a cloud server. Subject data is obtained through a subject data collection module (e.g., collecting DNA methylation characteristic values ​​from the subject's saliva). The subject's health optimization goals (treatment goals) are obtained through a human-computer dialogue module. Personalized intervention measures (treatment plans) are obtained using the SEMO features stored on the server and the calculation program.

[0248] Example

[0249] It should be noted that the terminology used in this specification is for the purpose of describing specific embodiments only and is not intended to limit the invention. The foregoing summary section and the following detailed description are for illustrative purposes only and are not intended to limit the invention in any way. The scope of the invention is defined by the appended claims without departing from its spirit and intent.

[0250] Example 1: Establishing a pre-trained model

[0251] Protein-protein interaction (PPI) data was obtained. In this embodiment, PPI data refers to a gene set in the form of "PPI:TP53", specifically the set of genes that have protein-protein interaction relationships with gene TP53 (denoted as TP53.n), totaling 1306 genes. The PPI data in this embodiment was obtained from the HPRD database (www.hprd.org / ). In this embodiment, genes interacting with ≥20 proteins were included in subsequent calculations.

[0252] This embodiment includes three categories of compounds: (1) nutraceuticals, defined from the DrugBank database (https: / / www.drug-bank.ca); (2) approved drugs, i.e., FDA-approved drugs in the DrugBank database; and (3) herbal compound components, data from the TCMID database (http: / / www.megabionet.org / tcmid / ). Target information for all compounds was obtained from the STITCH database (http: / / stitch.embl.de), with a STITCH score ≥ 200 as the screening criterion.

[0253] Screening PPI-compound combinations. Effective PPI-compound combinations are defined as those where the intersection of the PPI gene set and the compound target set is statistically significant. A small number of overlapping genes indicates a less significant combination; a large number of overlapping genes suggests a potentially important combination. Hypergeometric tests are used to determine statistical significance.

[0254]

[0255] Where x represents the number of genes in the intersection of the PPI gene set and the compound target set, K represents the number of genes in the PPI gene set, N represents the number of genes in the compound target set, and M represents the total number of genes in the genome. This embodiment only uses PPI-compound combinations with hypergeometric p-values ​​< 0.01 as valid candidate sequences for subsequent calculations.

[0256] Example 2: Establishment of a SEMO set related to the severity of novel coronavirus infection

[0257] Based on the aforementioned method, a set of semos related to the severity of SARS-CoV-2 infection was established in two whole blood DNA methylation data related to SARS-CoV-2 infection.

[0258] Data 1 is GSE179325 (https: / / www.ncbi.nlm.nih.gov / geo / query / acc.cgi?acc=GSE179325), which includes whole blood DNA methylation microarray data from 473 SARS-CoV-2 positive patients and 101 negative subjects. Among the positive patients, there were 360 ​​mild cases and 113 severe cases. Patients were classified according to the World Health Organization Clinical Rating Scale (WHO COVID-19 Therapeutic Trial Synopsis. R&D Blueprint (2020)). Mild cases were defined as those with a rating of 1-4, and severe cases were defined as those with a rating of 5-8 or those who died. DNA methylation of whole blood samples was detected using the Illumina Infinium Human Methylation EPIC BeadChip.

[0259] Data 2, GSE167202 (https: / / www.ncbi.nlm.nih.gov / geo / query / acc.cgi?acc=GSE167202), includes whole blood DNA methylation microarray data from 525 individuals. Of these, 164 were infected with the novel coronavirus, 296 were not infected, and 65 were infected with other non-novel coronaviruses. Based on the novel coronavirus infection severity score (SS), the groups are categorized as follows: 0 represents the uninfected group; 1 represents the home group; 2 represents the hospitalized group; 3 represents the ICU group; and 4 represents the deceased group. DNA methylation of whole blood samples was detected using the Illumina Infinium Human Methylation EPICBeadChip.

[0260] The data analysis and calculation process for the SEMO is as follows:

[0261] Gene feature extraction. A DNA methylation microarray detects DNA methylation at 485,512 cytosine sites in the human genome, the vast majority (482,421) of which are CpG sites. The methylation level of each methylation site is represented by a beta value, which represents the proportion of cytosine methylated at a specific site. Multiple methylation sites are distributed within the same functional region of the same gene in the microarray, and different sites will show different methylation values. The methylation values ​​in the microarray are normalized according to the manufacturer's instructions. Since each gene has multiple methylation sites, the methylation feature value of each gene is calculated based on the above-mentioned normalized methylation microarray data, i.e., methylation sites × subject sample matrix (CpG sites × Sample), to obtain the subject's methylation feature value map, i.e., gene × subject sample matrix (Gene × Sample). There are several ways to calculate the methylation feature value of each gene: Method 1, taking the average of the methylation beta values ​​of all methylation sites in the promoter region; Method 2, using the SIMPO method previously developed by the inventors. To simplify the data analysis process, subsequent procedures all adopt Method 1, which uses the average value of all methylation sites in the promoter region as the methylation characteristic value of the gene. After this transformation, the beta values ​​of multiple methylation sites in a gene are converted into a single characteristic value for each gene, forming a gene × sample matrix (Gene × Sample).

[0262] SEMO calculation. Based on the candidate SEMO features obtained in Example 1, the combination of PPIs and nutritional supplements is considered. Each PPI gene set and each nutritional supplement target set is enumerated, and a SEMO value is calculated. This SEMO value is derived from the T-test result, i.e.:

[0263]

[0264] in, This represents the mean methylation characteristic value of the target gene. S represents the mean methylation characteristic values ​​of non-target genes. x 2 S represents the variance of the methylation eigenvalues ​​of the target gene. y 2 denoted as , where is the variance of the methylation characteristic values ​​of non-target genes, n is the number of target genes, and m is the number of non-target genes; the target genes are the intersection of the gene set of the protein-protein interaction network and the target set of the compound, and the non-target genes are genes that belong to the gene set of the protein-protein interaction network but do not belong to the target set of the compound.

[0265] After the above calculations, the methylation data of the subjects is transformed into a sequence value × subject sample matrix (semo × Sample).

[0266] Then, the differences of each SEMOA feature between the severe group and the mild group in Data 1, and between the hospitalized group and the home-based group in Data 2, were calculated. The four SEMOA features with the most significant differences in Data 1 are as follows: Figure 2 As shown, they are STAT1-Lipoic acid, MMP2-L-Citrulline, CXCL8-Glucosamin, and CSF3-Calcitriol (also known as vitamin D3).

[0267] Further calculation of the odds ratio (OR) is used to evaluate the classification performance of each SEMo value on the phenotype (e.g., severe vs. mild), i.e., predictive ability. For example, for a given SEMo, let x be the SEMo value variable of 473 cases (360 mild cases and 113 severe cases), and xm be the mean of x. A simple predictor is set up that predicts cases with x values ​​higher than the mean xm as severe cases and cases with x values ​​lower than the mean xm as mild cases. The odds ratio (OR) can then be calculated to evaluate the predictor's performance: Let a be the number of subjects predicted as severe and actually being severe; b be the number of subjects predicted as severe and actually being mild; c be the number of subjects predicted as mild and actually being severe; and d be the number of subjects predicted as mild and actually being mild. Then the odds ratio (OR) = (a / b) / (c / d). If the odds ratio is <1, it indicates that individuals with higher SEMOS scores have a lower risk of severe illness than individuals with lower SEMOS scores. The further the odds ratio is from 1, the stronger the predictive ability of the predictor. Figure 2 Taking the first subplot as an example, OR = 0.11, which means that individuals with a semo value higher than the average have a risk that is 0.11 times higher than other individuals.

[0268] Obtain the intersection of the top 100 SEM features that are significantly related to the phenotype in Data 1 and Data 2, as shown in Table 1 (53 rows in total).

[0269] Table 1: Intersection of SEM images with predictive ability for novel coronavirus infection outcomes in Data 1 and Data 2

[0270]

[0271]

[0272]

[0273] Example 3: Generation of Disease Biomarkers and Predictive Models

[0274] Using the Lasso machine learning algorithm, with Data 1 as the training set and Data 2 as the independent validation dataset, a predictive model was built and validated. An optimized predictive model was obtained, and its specific parameters are shown in Table 2. The Lasso algorithm (leastabsolute shrinkage and selection operator algorithm, also known as the lasso algorithm) is a commonly used method in bioinformatics for building classification models. This algorithm aims to screen variables to reduce model complexity and avoid overfitting. After variable optimization, a total of 6 SEM variables were retained in the final predictive model. The last column in Table 2 shows the weighted value of each variable in the final linear model; the larger the absolute value, the greater its contribution to the predictive model. Table 2 is sorted in descending order of the absolute value of the weighted values.

[0275] Table 2: Parameter table of the optimized model obtained using the machine learning algorithm lasso

[0276]

[0277] The model exhibits good generalization ability on both the training and independent validation datasets; that is, the model obtained on the training set does not show a significant performance degradation on the independent validation dataset. Figure 3 As shown, the model has an AUC of 0.81 in the training set and an AUC of 0.79 in the independent validation set.

[0278] Example 4: Generation of Intervention Targets

[0279] Based on the T-test p-value (Data 1), all SEMO features were sorted, and the frequency of all PPIs and nutritional supplements appearing in the top 200 SEMO features was counted. The frequency of PPIs in the network is as follows: Figure 4As shown, the higher the frequency of a PPI network, the more nutritional supplements can target (intervene) that PPI network. The most frequent PPI is the ITGAM gene network, with ITGAM, the integrin subunit αm, as its central gene. ITGAM plays a role in stimulating endothelial cell adhesion in neutrophils and monocytes and has been reported to be associated with long-term pulmonary complications in patients after COVID-19 infection, indicating that this gene network is related to the severity of COVID-19 infection. The MMP9 (matrix metalloproteinase 9) and MMP2 networks play important roles in regulating inflammatory responses and have also been reported to be associated with mortality in patients with COVID-19 infection. Vascular cell adhesion molecule 1 (VCAM-1), which mediates leukocyte endothelial cell adhesion, has been reported to be associated with adverse patient outcomes and death in the ICU. Another group of high-frequency PPIs includes inflammatory cytokines such as CXCL8, IL4, and IL6; dysregulation of these protein networks may also lead to central nervous system inflammatory syndrome in patients with COVID-19 infection. The CD8A and GZMB gene networks are closely related to CD8+ T cell function and the severity of COVID-19 infection. These PPIs indicate that T cell activation, endothelial regulation, and cytokine / chemokine-related processes are crucial in the severity of SARS-CoV-2 infection and can be modulated through nutritional intervention. The importance of these biological processes is consistent with several characterization analyses of SARS-CoV-2 infection severity.

[0280] Example 5: Nutritional Supplement Intervention Ranking

[0281] In the SEMO features related to the severity of novel coronavirus infection in Data 1, the frequency of nutritional supplements appearing in the top 200 SEMO features is as follows: Figure 5 As shown, the higher the frequency of a nutritional supplement, the more PPI networks it can target (intervene in).

[0282] Among frequently occurring nutritional supplements, L-Tryptophan was the most frequently observed metabolite, indicating that its accumulation is an important biomarker. L-Citrulline ranked second; it has been reported as an indicator of the prognosis of severe COVID-19 infection. L-Isoleucine was significantly correlated with disease severity and gut microbiota in COVID-19 patients, showing impaired L-isovalerate biosynthesis. These severity-related metabolites are biomarkers of disease severity, suggesting that interventions along these metabolic pathways may be potential therapeutic candidates.

[0283] More importantly, Figure 5The study directly highlighted a group of modulators that regulate the severity or treat COVID-19 infection. Several top-ranked nutritional supplements, such as lipoic acid (ranked 3rd), glucosamine (ranked 4th), vitamin D3 (calcitriol (ranked 5th), vitamin A (ranked 7th) / retinoic acid (ranked 8th), and spermine (ranked 9th), have been reported in the literature to be effective in symptom management and treatment of COVID-19 infection.

[0284] Example 6: Generation of candidate drugs

[0285] The same analysis was performed on FDA-approved drugs, and the SEM features associated with the severity of COVID-19 infection in Data 1 were as follows: Figure 6 The drugs shown are MMP9-HEPARIN, CXCL8-HEPARIN, CXCL8-CERIVASTATIN, and MMP2-INDOMETHACIN. Heparin has been reported to affect matrix metalloproteinases and may benefit severely ill COVID-19 patients through anticoagulation, prompting a series of studies and clinical trials. Indomethacin is an anti-inflammatory and broad-spectrum antiviral drug; clinical trials have demonstrated its benefit in the treatment and management of symptoms related to COVID-19 infection.

[0286] Therefore, SEMO analysis may help identify new indications for approved drugs. Since this SEMO analysis was guided by phenotypic information from severe vs. mild cases, the drugs identified were primarily related to the regulation of the severity of COVID-19 infection.

[0287] Example 7: Generation of Natural Product Drug Candidates

[0288] The same analysis was performed on the compound components of traditional Chinese medicine as described above. The SEMO features associated with the severity of novel coronavirus infection in Data 1 are as follows: Figure 7 As shown, they are CSF2-Parthenolide, CSF2-Andrographolide, TLR2-Parthenolide, and LCK-Leurosidine, respectively.

[0289] Silva lactones are important active components of medicinal plants such as chrysanthemum and tansy. Belonging to the sesquiterpene lactone class of compounds, they are NF-κB inhibitors exhibiting anti-inflammatory, antitumor, and antiviral activities. They can inhibit IL-6 production and suppress the papain-like protease activity of the novel coronavirus. Andrographolide is the main active compound of the plant Andrographis paniculata, widely used in traditional Chinese medicine. It has anti-inflammatory effects and can reduce cytokine storms. Semo analysis suggests that, from a protein network perspective, the mechanism of action of these natural products mainly involves regulating cytokine production through the CSF2 and TLR2 protein networks, thereby affecting the human body's response to viruses.

[0290] In Data 1, the frequency of herbal compound components in the top 200 SEMs related to the severity of novel coronavirus infection was analyzed, such as... Figure 8 As shown, ternolactone and andrographolide, mentioned above, rank 2nd and 3rd respectively in this ranking. Notably, ursodeoxycholic acid, ranked 4th, has recently been reported in the literature as a potential therapeutic for the novel coronavirus. Semo analysis revealed that ursodeoxycholic acid may target CD4+ and CD69+ cells.

[0291] Therefore, SEMO analysis may help identify potential drug candidates from natural products. Since this SEMO analysis was guided by phenotypic information from severe vs. mild cases, the drugs identified were primarily related to the regulation of the severity of COVID-19 infection.

[0292] Example 8: Personal Disease Risk Score and Intervention Recommendations

[0293] From an individual perspective, it is necessary to compare the relative impact and risk of different SEM characteristics in order to propose potential interventions and preventive measures. To this end, a logistic regression model was used to predict the phenotypic label y for each SEM (severe cases are labeled y=1, mild cases are labeled y=0). Subsequently, the SEM × patient matrix was converted into a risk score × patient matrix, where the values ​​are risk scores based on the SEM, ranging from 0 to 1.

[0294] The relative contribution of SEM (Semo) characteristics to severity risk was assessed for three individuals (including patient 1 with mild symptoms, patient 11 with severe symptoms, and patient 21 with severe symptoms) using the method described above. Figure 9As shown in the left subplot, the maximum risk score for each SEMO dimension is <0.1; in the middle and right subplots, the risk scores for most SEMO dimensions are >0.4. Even among patients in the severe illness group, there are differences in risk scores across the various SEMO dimensions. In the middle subplot, the highest risk comes from the SEMO "CXCL8-Glucosamine," while the right subplot shows two important risk factors: "STAT1-Lipoic acid" and "CSF2-Magnesium." This analysis demonstrates that SEMO analysis can resolve the heterogeneity between individuals within the same disease group; therefore, it can be used to rank nutritional interventions or drug treatment regimens suitable for each individual.

[0295] Furthermore, the SEMOR dimensions of the critically ill patient No. 21 were reordered according to risk values, such as... Figure 10 As shown, among the top 15 SEMO dimensions, Lipoic acid (4 times) and L-Citrulline (2 times) appeared frequently. Both of them target MMP2.n, indicating that a combination of certain nutritional supplements can target a set of related protein networks, forming the basis of a personalized intervention plan.

[0296] Example 9: Detection and intervention protocols targeting specific immune cells and signaling pathways

[0297] Semo analysis can also rank interventions for specific cell types and signaling pathways, including immune cells closely related to the process of novel coronavirus infection and the mechanism of post-infection sequelae, such as T lymphocytes and monocytes / macrophages.

[0298] CD68 is a hallmark molecule of monocytes / macrophages, and its protein-protein interaction network plays an important role in macrophage signaling pathways. Based on data 1, a t-test was used to identify four CD68-containing SEMO features that showed the most significant differences between the severe and mild cases, such as... Figure 11The images show CD68-Lipoic acid, CD68-Succinic acid, CD68-Magnesium, and CD68-Niacin, respectively. Further investigation revealed the intersection between the CD68 protein-protein interaction network gene set and the lipoic acid target set, including IL6 and IL10, suggesting that lipoic acid can influence the expression levels of IL6 and IL10 through DNA methylation-related mechanisms. IL6 and IL10 are key biomarkers closely related to the cytokine storm in SARS-CoV-2 infection. Lipoic acid has also been shown to alleviate the severity of SARS-CoV-2 infection by inhibiting the excessive production of reactive oxygen species (ROS) and inflammatory cytokines.

[0299] These results also suggest that SEMO analysis has application value in nutritional interventions targeting specific immune cells, especially in screening epigenetic reprogramming reagents for immune cells / stem cells.

[0300] Table 3 shows a set of SEMO features targeting specific immune cells. Using the results in Table 3, SEMO analysis can be used to rank various candidate nutritional supplements if intervention in specific cells is needed to influence the outcome of COVID-19 infection. For example, in clinical practice, if CD68 regulation is required, the decision to use medication can be made based on the results in Table 3 combined with the practical experience of medical staff or literature reports.

[0301] Table 3: Semo characteristics of nutritional interventions targeting key immune pathways CD68, CD4, IL6, CXCL8, and FOXP3

[0302]

[0303]

[0304] The methods, systems, media, and applications provided by this invention have been described in detail above. Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the methods and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. A method for the quantitative characterization of compound interactions, characterized in that, The method includes the following steps: (1) Obtain a target set of a single compound, said compound including drugs and nutritional supplements; (2) Obtain all genes that interact with a single gene, and use the single gene and all genes that interact with it as the gene set of the protein interaction network of the single gene; (3) Calculate the SEMO value of the target set of the single compound and the gene set of the protein-protein interaction network of the single gene, wherein the SEMO value is calculated by the following formula: seed = , in, The mean of the eigenvalues ​​of the target gene. The mean of the eigenvalues ​​of non-target genes. S is the variance of the eigenvalues ​​of the target gene. y 2 denoted as , where is the variance of the eigenvalues ​​of non-target genes, n is the number of target genes, and m is the number of non-target genes; the target genes are the intersection of the gene set of the protein-protein interaction network and the target set of the compound, and the non-target genes are genes that belong to the gene set of the protein-protein interaction network but do not belong to the target set of the compound. The SEMO value is a characterization of the effect of a single compound by using the protein-protein interaction network of the single gene.

2. The method according to claim 1, characterized in that, The feature values ​​are selected from gene methylation feature values ​​and expression values.

3. The method according to claim 2, characterized in that, The methylation characteristic value is the average of the methylation beta values ​​or the SIMPO value of all methylation sites in the gene body region.

4. The method according to claim 3, characterized in that, The methylation characteristic value is the average of the methylation beta values ​​or the SIMPO value of all methylation sites in the gene promoter region.

5. The method according to any one of claims 1-4, characterized in that, The number of genes in the gene set of the protein-protein interaction network of the single gene is ≥20.

6. A system for the quantitative characterization of compound interactions, characterized in that, The system implements the method according to any one of claims 1-5, and the system comprises: The gene acquisition module is used to acquire the target set of a single compound and the gene set of the protein-protein interaction network of a single gene, thereby obtaining the target gene and non-target gene. The data acquisition module is used to acquire feature values ​​of the target genes and non-target genes from the subjects; The calculation module is used to calculate the SEMO value based on the feature values ​​of the target gene and non-target genes.

7. A medium for quantitative characterization of compound interactions, characterized in that, The medium includes a program for implementing the method of any one of claims 1-5, the medium comprising: (1) Obtain the target set of a single compound and the gene set of the protein-protein interaction network of a single gene, and then obtain the target gene and non-target gene. (2) Read the characteristic values ​​of the target genes and non-target genes of the subject; (3) Calculate the SEMO value based on the feature values ​​of the target gene and non-target genes.

8. A method for providing a compound intervention regimen to a subject, characterized in that, The method uses the method described in any one of claims 1-5, and the method includes the following steps. (1) Obtain protein-protein interaction network data and compound target data; (2) Based on the protein interaction network data and compound target data, protein interaction network and compound combination are screened, i.e., SEMO features, wherein the protein interaction network and compound combination is the intersection of the gene set of the protein interaction network and the gene set of the compound target is significant. (3) Based on the selected SEMO features, an evaluation model for the phenotypic data is established using the phenotypic data, wherein the phenotypic data is a feature value × sample matrix; The establishment of the evaluation model includes the following steps: (i) Based on the selected SEMO features, obtain the SEMO values ​​of the SEMO features in the phenotypic data, and finally obtain the SEMO feature × sample matrix; (ii) Based on the semo features × sample matrix in step (i), obtain semo features that are significantly correlated with the phenotype through machine learning algorithms; (4) For the SEMO feature that is significantly correlated with the phenotype, calculate the SEMO value of the subject and provide the subject with a compound intervention program based on the SEMO value.

9. A system for providing a compound intervention regimen to a subject, characterized in that, The system implements the method of claim 8, the system comprising: Data acquisition unit, used to acquire protein interaction network data, compound target data, and phenotypic data; The semo filter is used to filter valid semo features. A computational analyzer is used to calculate the SEMO value of the effective SEMO feature in the phenotypic data, obtain the SEMO feature that is significantly correlated with the phenotypic, and provide a compound intervention plan for the subject based on the SEMO value of the SEMO feature that is significantly correlated with the phenotypic.

10. A medium for providing a compound intervention regimen to a subject, characterized in that, The medium includes a program for implementing the method of claim 8, the medium comprising: (1) Obtain protein interaction network data, compound target data, and phenotypic data; (2) Filter valid SEMO features; (3) Calculate the SEMO value of the effective SEMO feature in the phenotypic data, obtain the SEMO feature that is significantly correlated with the phenotypic, and provide the subject with a compound intervention plan based on the SEMO value of the SEMO feature that is significantly correlated with the phenotypic.

11. A method for providing a compound intervention regimen to subjects infected with the novel coronavirus, characterized in that, The method described herein uses the method described in claim 8. SEM features positively correlated with COVID-19 infection outcomes include: The single gene is FUT4, and the compound is L-Alanine; SEMO features negatively associated with COVID-19 infection outcomes include: The single gene is IL6, the compound is L-Tryptophan; and / or The single gene is CXCL8, and the compound is Vitamin E; and / or The single gene is ITGA2, the compound is L-Cystine; and / or The single gene is CXCL8, the compound is Calcitriol; and / or The single gene is VCAM1, the compound is L-Citrulline; and / or The single gene is CTSB, the compound is Vitamin A; and / or The single gene is VCAM1, and the compound is Spermine; and / or The single gene is ITGAM, the compound is L-Tryptophan; and / or The single gene is ITGA2, the compound is L-Citrulline; and / or The single gene is ITGAM, the compound is Glycine betaine; and / or The single gene is MMP9, the compound is L-Aspartic Acid; and / or The single gene is TFRC, the compound is L-Tryptophan; and / or The single gene is LCK, the compound is L-Tyrosine; and / or The single gene is ITGAM, the compound is L-Arginine; and / or The single gene is GZMB, and the compound is Vitamin A; and / or The single gene is ITGAM, the compound is L-Isoleucine; and / or The single gene is CXCL8, the compound is L-Tryptophan; and / or The single gene is MMP9, the compound is L-Citrulline; and / or The single gene is ITGAM, the compound is L-Valine; and / or The single gene is MMP9, the compound is L-Tryptophan; and / or The single gene is VCAM1, the compound is Clopidogrel; and / or The single gene is CD40LG, the compound is Niacin; and / or The single gene is CSF2, and the compound is Choline; and / or The single gene is IL6, the compound is L-Citrulline; and / or The single gene is IL15, and the compound is Vitamin A; and / or The single gene is MMP9, the compound is Glycine betaine; and / or The single gene is MMP2, and the compound is Melatonin; and / or The single gene is MMP2, the compound is L-Valine; and / or The single gene is ALB, and the compound is N-Acetyl-D-glucosamine; and / or The single gene is MMP2, the compound is Glycine betaine; and / or The single gene is CD40LG, the compound is L-Tryptophan; and / or The single gene is VCAM1, the compound is L-Tryptophan; and / or The single gene is CXCL8, and the compound is N-Acetyl-D-glucosamine; and / or The single gene is CSF3, the compound is Calcitriol; and / or The single gene is ITGAM, the compound is L-Citrulline; and / or The single gene is MMP9, and the compound is Spermine; and / or The single gene is TLR2, the compound is Lipoic Acid; and / or The single gene is CD40, the compound is Lipoic Acid; and / or The single gene is ITGAM, the compound is Clopidogrel; and / or The single gene is CXCL8, the compound is L-Citrulline; and / or The single gene is SPI1, the compound is L-Cysteine; and / or The single gene is IL4, and the compound is Choline; and / or The single gene is MMP2, and the compound is Spermine; and / or The single gene is STAT3, the compound is L-Citrulline; and / or The single gene is ITGAX, the compound is Choline; and / or The single gene is IL4, and the compound is Lipoic Acid; and / or The single gene is MMP2, the compound is L-Aspartic Acid; and / or The single gene is MMP2, the compound is L-Tryptophan; and / or The single gene is CD8A, the compound is Choline; and / or The single gene is MMP2, the compound is L-Citrulline; and / or The single gene is STAT1, the compound is Lipoic Acid; and / or The single gene is CD8A, and the compound is Lipoic Acid; The novel coronavirus infection outcome indicates the efficacy of treatment with the compound in the subjects.

12. A method for providing a compound intervention regimen to subjects infected with the novel coronavirus, characterized in that, The method described herein uses the method described in claim 8. The single gene is CD68: SEM features positively correlated with COVID-19 infection outcomes include: The compounds are Succinic acid, Magnesium, and Niacin. SEMO features negatively associated with COVID-19 infection outcomes include: The compound is Lipoic Acid; The single gene in question is CD4: SEM features positively correlated with COVID-19 infection outcomes include: The compounds are Calcium and Magnesium. SEMO features negatively associated with COVID-19 infection outcomes include: The compounds are Tretinoin and Calcitriol; The single gene is IL6: SEM features positively correlated with COVID-19 infection outcomes include: The compound is Glycine. SEMO features negatively associated with COVID-19 infection outcomes include: The compounds are L-Citrulline, L-Valine, and L-Tryptophan; The single gene is CXCL8: SEMO features negatively associated with COVID-19 infection outcomes include: The compounds are Glucosamine, Calcitriol, Tretinoin, and L-Citrulline; The single gene in question is FOXP3: SEM features positively correlated with COVID-19 infection outcomes include: The compounds are Vitamin E, Succinic acid, Magnesium, and Glutathione. The novel coronavirus infection outcome indicates the efficacy of treatment with the compound in the subjects.