Method, device and kit for predicting efficacy based on expression profile of small number of genes

By using clustering and deep neural networks to screen out a small number of representative genes, a drug action mechanism prediction model was constructed. This solved the problems of data stability and low detection efficiency caused by too many gene types in the L1000 technology, and achieved efficient and low-cost drug action mechanism prediction.

CN115905898BActive Publication Date: 2026-01-09ACADEMY OF MILITARY MEDICAL SCIENCES
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211278396.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-19
Publication Date
2026-01-09
Estimated Expiration
2042-10-19

AI Technical Summary

Technical Problem

The existing L1000 technology has too many gene types when detecting gene expression profiles, resulting in poor data stability, low detection efficiency, and high complexity in drug association analysis, making it difficult to effectively predict drug action mechanisms with a small number of genes.

Method used

N types (90≤N≤150) of central genes were screened using clustering methods. An expression matrix with redundant features was constructed. An MOA prediction model was established using the LR model. Drug and gene expression profile data were trained using deep neural networks to screen 100 representative genes, which were then detected using Luminex liquid-phase chip technology.

Benefits of technology

It reduces the number of gene types to be detected, improves detection efficiency, reduces costs, and maintains the accuracy of drug action mechanism prediction, making it suitable for large-scale compound transcriptome data collection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115905898B_ABST
    Figure CN115905898B_ABST
Patent Text Reader

Abstract

A method, device and kit for predicting drug efficacy based on a small number of gene expression profiles. The screening method for gene profiles for drug prediction comprises: constructing a 978*978 gene expression correlation matrix using 978 landmark genes in the public data of L1000; clustering the 978 genes into N classes using a clustering method, and selecting a central gene in each class cluster, i.e. the gene with the highest average correlation with other genes in the same cluster as the landmark gene; wherein 90<=N<=150. Compared with the L1000 technology, the method of the present application reduces the number of genes to be detected, and does not require the use of a magnetic bead to detect two genes in proportion, which can effectively reduce the detection cost, improve the detection efficiency, and is suitable for large-scale compound transcriptome data acquisition; the prediction accuracy of the present application is high, while the detection cost is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of drug action mechanism prediction and drug repositioning, and in particular to a drug efficacy prediction method, device and kit based on an expression profile of a small number of genes. BACKGROUND

[0002] L1000 technology is a revolutionary technology developed by Broad Institute, which is a combination of liquid phase chip, gene correlation and artificial intelligence. The theoretical basis is to reduce the number of detected genes to 978 genes by using the high correlation between gene expressions, and then to obtain the expression information of the whole transcriptome by calculating the expression profile of the 978 genes. The principle is to use gene-specific probes containing universal primer sequences and coding sequences for ligation-mediated amplification (LMA), and combine Luminex liquid chip technology and a method of detecting two genes with large expression differences in a fixed ratio (2:1) to detect the expression data of 978 genes, so that the whole L1000 technology is more suitable for large-scale data acquisition. However, since L1000 technology involves multiple reactions of about 1000 genes and a magnetic bead detecting two genes, it is easy to cause poor data stability and low detection efficiency. In addition, one of the most important application scenarios of L1000 is to analyze the correlation between drugs based on the similarity between expression profiles, that is, the "guilt by association" strategy. The basic assumption is that if drugs have similar transcriptomic characteristics, they may also have similar targets and indications. In the past, many large-scale similarity correlation analysis studies have been conducted, which mainly compare the target drug with a large number of reference drugs that have annotated information such as targets, mechanisms of action and indications, to mine the potential unknown relationship between drugs and drugs, so as to divide the possible new action categories of the target drug. However, due to the co-expression characteristics of genes, this similarity correlation does not always need to be performed on the whole gene transcriptome level. Studies including L1000 have shown that only a limited number of gene expression characteristics can be used to evaluate the similarity between expression profiles. Therefore, it may not be necessary to detect 978 or more genes to achieve similarity evaluation, and thus to achieve drug action mechanism prediction and drug repositioning. Therefore, the industry urgently needs to develop a drug efficacy prediction method and device based on an expression profile of far fewer than 978 genes, in order to reduce the complexity of L1000 multiple reactions, improve detection efficiency and reduce costs. SUMMARY

[0003] Therefore, the main purpose of the present application is to provide a drug efficacy prediction method, device and kit based on an expression profile of a small number of genes, in order to at least partially solve the above technical problems.

[0004] To achieve the above object, as a first aspect of the present application, a gene set screening method for drug prediction is proposed, comprising the following steps:

[0005] 978 × 978 gene expression correlation matrix is constructed using 978 landmark genes in the public data of L1000;

[0006] The 978 genes are clustered into N classes using a clustering method, and a central gene is selected in each class cluster, that is, the gene with the highest average correlation with other genes in the same cluster is selected as the landmark gene; wherein 90 ≤ N ≤ 150.

[0007] Among them, the selection of the central gene uses the co-expression characteristics of the gene to select a representative central gene as the representative gene of the co-expression gene, thereby constructing a de-redundant expression matrix;

[0008] Among them, the screening process uses an LR model to establish and train an MOA prediction model reflecting the relationship between the drug and the gene expression profile data of the drug, the MOA prediction model is a binary classification model, and the score refers to the probability value of predicting whether each drug has the same or similar drug action mechanism as the drug for establishing the prediction model using the prediction model; The binary classifier training set includes two class samples: "positive set" and "negative set", the "positive set" label refers to the drug MOA annotation of the MedChemExpress library and the drug repositioning center, and the "negative set" selects compounds with low transcriptional activity and no MOA annotation, and assumes that they have no drug properties that can be reflected at the transcriptome level.

[0009] Among them, the clustering method uses K-means algorithm and cosine similarity measure to realize;

[0010] Among them, the t-distributed stochastic neighbor embedding algorithm is used in the clustering method, and the data is initialized by principal component analysis.

[0011] Among them, 3-fold cross-validation is used to evaluate the performance of the MOA prediction model; In cross-validation, the positive and negative drug sets are randomly divided into K parts by stratified sampling method, 2 / 3 of the samples are used as the training set to train the MOA prediction model, and the sensitivity and specificity are evaluated by testing the remaining 1 / 3 of the samples; This process is performed 3 times, with the average AUROC as the evaluation index, and the probability score of the model to the negative and positive sample set in each cross-training process is recorded, and the model with average AUROC ≥ 0.6 is regarded as a well-trained model in this process.

[0012] wherein the annotation information of the drug function of the drug is obtained from a Drug Repurposing Hub information base, including a drug action mechanism; and / or

[0013] obtaining gene expression profile data of the drug from a LINCS expression profile data set; and / or

[0014] sorting 103 drug sets with specific drug action mechanisms from a drug set of Medchemexpress Company;

[0015] taking the drug action mechanism of each drug set as a true label, training the MOA prediction model by using the gene expression profile data of the drug in each drug set to obtain a prediction model of each drug set; and then analyzing the gene expression profile data by using the prediction model of each drug to obtain the MOA prediction score of the gene expression profile data, and sorting the MOA prediction score, thereby screening and clustering to obtain the landmark gene.

[0016] As a second aspect of the present application, a gene set for drug prediction obtained according to the screening method as described above is also provided.

[0017] Among them, the gene set is: RNMT, TOPBP1, CBR3, IL1B, HADH, DHRS7, UBE2J1, NUDT9, CASC3, PGRMC1, KDM5B, DAG1, NUP62, CCNA2, NUP88, ALAS1, FAH, LYN, TRAPPC6A, MEST, NENF, GDPD5, HSPA1A, ICAM3, DNMT1, CDC25A, TSC22D3, PCMT1, SCARB1, BLVRA, POLR2K, KIAA0196, GFPT1, GAA, SLC35B1, LIG1, IKBKB, LYPLA1, SKP1, UBE3C, PRAF2, DDB2, AKAP8, IER3, FOXJ3, AKAP8L, GATA2, FBXO21, DECR1, PTPRF, RAE1, SPTLC2, SMARCA4, USP22, PSME1, HMOX1, CALM3, G3BP1, HSD17B10, AURKA, GADD45A, TSKU, ARNT2, CCNB2, IGF2R, CDK1, AURKB, DRAP1, CCNF, DERA, IFNAR1, DNM1, FOXO3, ASAH1, FOXO4, SCAND1, GABPB1, CLIC4, CALU, MAP3K4, RSU1, ALDH7A1, TBP, BUB1B, ABCF1, ANO10, PCBD1, PSMG1, ITGB5, NIPSNAP1, MYBL2, SACM1L, DNTTIP2, INSIG1, MTHFD2, TRIB3, RBM15B, ECH1, SCRN1, DUSP4.

[0018] As a third aspect of the present application, a drug prediction method using the expression profile of the gene set as described above is also proposed, comprising the following steps:

[0019] A classification prediction model of the association between drugs and the gene set is established using the expression profile of at least the genes in the gene set as described above, and the total number of genes is not more than one-third of L1000, more preferably not more than 150 genes, the classification prediction model is trained using the characteristics that drugs with the same or similar drug action mechanism have the same or similar efficacy, and the classification prediction model is used to predict whether the drug has efficacy.

[0020] As a fourth aspect of the present application, a kit is also provided, which uses at least the universal primer sequences and gene-specific probes of the gene set as described above, with the total number of genes not exceeding one-third of L1000, more preferably not exceeding 150 genes, to detect the expression data of the gene set, so that fewer detection genes can be used to achieve drug prediction.

[0021] The kit comprises:

[0022] A PCR plate coated with poly-dT primers for capturing mRNA in the sample to be tested;

[0023] An upstream and downstream probe set containing upstream and downstream probes for specific capture of the amplified target detection genes present in the sample to be tested on the PCR plate, using the gene expression profile as described above;

[0024] A result presentation unit for presenting the detection results in the form of fluorescence based on the results of capturing the amplified target detection genes by the upstream and downstream probe set.

[0025] The kit further comprises:

[0026] A cell lysis solution containing 1% β-mercaptoethanol for lysing cells to prepare a crude cell lysis solution; and / or

[0027] A liquid chip coupled with corresponding capture probes for detecting the prepared PCR products.

[0028] The result presentation unit comprises Luminex magnetic beads for coupling the captured amplified target detection genes, and the surface emits fluorescence signals with corresponding intensity values.

[0029] As a fifth aspect of the present application, a method for detecting using the kit as described above is also provided, comprising the following steps:

[0030] Capturing mRNA of the sample to be tested using a PCR plate coated with dT primers and reverse transcribing to generate a first strand of cDNA;

[0031] Synthesizing upstream and downstream probes using specific fragments of each target nucleic acid, wherein the upstream probe contains a T7 universal primer sequence, a Barcode sequence, and an upstream primer specific to the specific fragment of the target nucleic acid, and the downstream probe contains a downstream primer specific to the specific fragment of the target nucleic acid and a T3 universal primer sequence;

[0032] Hybridizing the upstream and downstream probes with the target cDNA formed by reverse transcription to form a ligation hybrid;

[0033] The above-mentioned connecting hybrid is used as a template, a T7 universal primer sequence with a biotin label is used as a forward primer, and a reverse T3 universal primer sequence is used as a reverse primer for amplification;

[0034] The amplification product labeled with biotin is hybridized with magnetic beads coupled with corresponding antisense Barcode sequences, and SAPE staining is performed by using a biotin-streptavidin cascade amplification reaction principle;

[0035] The average fluorescence signal value on each magnetic bead is detected by using a Luminex 3D instrument to obtain expression information of each gene.

[0036] Based on the above technical solution, the pharmacodynamic prediction method, device and kit based on the expression profile of a small number of genes of the present application at least have one of the following beneficial effects relative to the prior art:

[0037] As shown in Figure 4 Compared with the L1000 technology, the D100 technology of the present application reduces the types of genes to be detected (from 978 genes of L1000 to 100 genes), and does not need to use a method of detecting two genes in proportion by using one magnetic bead, thereby effectively reducing the detection cost, improving the detection efficiency, and having similar compound action mechanism prediction performance as L1000, and being suitable for large-scale compound transcriptome data acquisition.

[0038] As shown in Figure 9 、 10 It is found by experiments that, when 6H of 16 randomly selected drugs containing high, medium and low transcriptional activities is detected, 12 of the 16 detected drugs are predicted to have their original functions, and the detection accuracy is about 75%; when 24H of the drugs is detected, 12 of the 16 detected drugs are predicted to have their original functions, and the detection accuracy is about 75%; and when the results of 24H and 6H of the drugs are combined, 14 of the 16 detected drugs are predicted to have their original functions, and the prediction accuracy is 87.5%. BRIEF DESCRIPTION OF DRAWINGS

[0039] The method and device of the present application will be further described below in combination with the drawings and examples:

[0040] Figures 1A-1F is a schematic diagram of the screening process and prediction performance test of the D100 genes of the present application, wherein Figure 1A 、 1B is the prediction performance of GPAR and CS when the number of landmark genes is different; Figure 1C is the difference between the MOA positive set and the MOA negative set when 10-900 landmark genes are predicted; Figure 1Dis the co-expression relationship matrix of 978 landmark genes of L1000 technology in 310,000+ expression profiles, and the genes with co-expression relationship are divided into 100 cluster clusters by K-means clustering; Figure 1E is the data visualization of D100 features and L1000 features (in the figure, the 9 MOAs with the highest prediction performance of D100 and L1000 are colored and displayed); Figure 1F is the prediction performance comparison of D100 data and L1000 data;

[0041] Figures 2A-2B is a schematic diagram of MOA prediction performance test of the D100 technology of the application; wherein Figure 2A is the number comparison of MOA prediction accuracy of L1000, D100 and L1000FWD on 284 drugs; Figure 2B is the ROC curve of L1000 and D100 for MOA prediction of all test drugs;

[0042] Figure 3 is a schematic diagram of the detection principle of the D100 technology of the application;

[0043] Figure 4 is a schematic diagram of the comparison analysis of D100 of the application and L1000 of the prior art;

[0044] Figures 5A-5C is a reliability and stability analysis chart of the D100 technology of the application after 6H of drug action; wherein Figure 5A is a background value analysis and pollution analysis schematic diagram; Figure 5B is a well-to-well fold change analysis schematic diagram; Figure 5C is a CV value analysis schematic diagram; Figure 5D is a control well similarity analysis schematic diagram;

[0045] Figures 6A-6D is a reliability and stability analysis chart of the D100 technology of the application after 24H of drug action; wherein Figure 6A is a background value analysis and pollution analysis schematic diagram; 6B is a well-to-well fold change analysis schematic diagram; Figure 6C is a CV value analysis schematic diagram; Figure 6D is a control well similarity analysis schematic diagram;

[0046] Figure 7 is a cosine similarity matrix schematic diagram of the D100 technology of the application after 6H of drug action;

[0047] Figure 8 is a cosine similarity matrix schematic diagram of the D100 technology of the application after 24H of drug action;

[0048] Figure 9 is the prediction result schematic diagram of the D100 technology of the application after the drug acts for 6H to collect data;

[0049] Figure 10 is the prediction result schematic diagram of the D100 technology of the application after the drug acts for 24H to collect data. DETAILED DESCRIPTION

[0050] In order to make the objects, technical solutions and advantages of the present application clearer, further detailed description will be made to the present application in combination with specific embodiments and with reference to the drawings.

[0051] L1000 technology is a revolutionary technology developed by Broad Institute, which is a combination of liquid phase chip, gene correlation and artificial intelligence. The theoretical basis of L1000 technology is to reduce the detected gene types to 978 genes by using the high correlation between gene expressions, and then to obtain the expression information of the whole transcriptome by calculating and inferring the expression profile of the 978 genes. However, 978 genes are still too many, and since L1000 technology involves multiple reactions of about 1000 genes and a magnetic bead detection of two genes, it is easy to cause poor data stability and low detection efficiency. The present inventors have found through a large number of studies that due to the co-expression characteristics of genes, such similarity correlation does not always need to be carried out on the whole transcriptome level. Many studies including L1000 have shown that only through limited gene expression characteristics can the similarity evaluation between expression profiles be realized. After combining with the GPAR technology researched by the present inventors, the expression profile of 100 genes (much less than 978 genes) is used for the pharmacodynamic prediction method and device, which not only reduces the complexity of L1000 multiple reactions, but also improves the detection efficiency and greatly reduces the cost.

[0052] Therefore, the present inventors propose a device and method for determining cell biological reactions by detecting the expression of much less than 978 genes, so as to realize effective prediction of the mechanism of compound action with extremely low transcriptome detection cost and extremely high detection efficiency. The device and method are obtained by more compact clustering of the detected gene types. Clustering is a conventional technical means, but the clustering method in the present application is original.

[0053] Specifically, the 100 genes (D100) are screened by the following method:

[0054] I. Detection gene selection:

[0055] 1. Experimental method:

[0056] 1.1 Screening of D100 landmark genes

[0057] The screening of central genes is to screen a representative central gene as a representative gene of co-expression genes by using the co-expression characteristics of genes (i.e. the characteristics that a large number of genes usually exist simultaneously expressed or simultaneously inhibited in the process of expression) to construct an expression matrix with a de-redundant feature. Firstly, the pairwise correlation (cosine similarity) expression analysis between genes is carried out by using 978 landmark genes in the public data of L1000 phase1 and phase2 to construct a 978x978 gene expression correlation matrix, then the K-means clustering method is used to cluster 978 genes into 100 classes, and the central gene is selected in each class cluster, i.e. the gene with the highest average correlation with other genes in the same cluster is selected as the landmark gene of D100. The K-means algorithm and cosine similarity measure are realized by the python package scikit learn. The D100 data visualization adopts the t-Distributed Stochastic Neighbor Embedding (t-SNE) algorithm, and the data is initialized by the Principal Component Analysis (PCA) and the cosine similarity and the rest of the default parameters.

[0058] 1.2 Establishment of prediction model

[0059] The LR model is used to train the MOA prediction model, and the binary classification model will assign different weights to different genes when mapping D100 genes to MOA. The establishment of the algorithm model is realized by using the logistics regression function (the parameter is default) of the scikit-learn package. The training set of the binary classifier includes two class samples: "positive set" and "negative set". The "positive set" label refers to the MOA annotation of drugs in the MCE (MedChemExpress) library and the drug repurposing hub. The present inventors delete some positive molecules with too large difference from other "positive set" through repeated cross-validation process. In addition, the present inventors also select 6220 compounds (low transcriptional activity and no MOA annotation) as the unchanged "negative set", and assume that they have no drug properties that can be reflected at the transcriptome level.

[0060] Model evaluation and selection of default model: the performance of the prediction model is evaluated by 3-fold cross-validation. In cross-validation, the positive and negative drug sets are randomly divided into K parts by stratified sampling method. 2 / 3 of the samples are used as the training set to train the prediction model, and the sensitivity and specificity are evaluated by testing the remaining 1 / 3 of the samples. This process is performed 3 times, and the average AUROC is used as the evaluation index, and the probability score of the model for the negative and positive sample set is recorded in each cross-training process. The model with an average AUROC≥0.6 is considered to be a well-trained model.

[0061] Preferably, the MOA prediction model (GPAR prediction model) uses a GPAR prediction tool, the specific content of which is as follows:

[0062] The MOA prediction model uses a deep neural network (DNN), which can be understood as a neural network including many hidden layers. The first layer of the deep neural network is the input layer (input), and the last layer of the deep neural network is the output layer (output). The layers between the input layer and the output layer are called hidden layers (hidden), and the layers are fully connected. That is, any neuron in the ith layer is connected to any neuron in the i+1th layer. The prediction model uses the drug mechanism of action as the true label, or assumes that the drug is a known positive compound (Positive Compound), and sets the positive drug as the true label. The gene expression profile data of known drugs is input into the deep neural network for training to obtain the prediction model of these drugs, and the prediction model of the drug corresponds to the drug mechanism of action. Therefore, the trained prediction model of the drug can be used to predict the predicted drug or potential drug that may have the same or similar drug mechanism of action as the drug, and a variety of drugs with the same or similar drug mechanism of action can be used to score the gene expression profile data of the prediction model, so that the prediction model can be constructed into an MOA model using the LR model to predict the accuracy of the gene expression profile data, and the gene expression profile data with the highest prediction accuracy is selected to screen the candidate landmark genes, thereby realizing the unique clustering of the present application.

[0063] The drug identifier is used to uniquely identify a drug, which can be the generic name of the drug, the trade name of the drug, or the drug identifier (Pubchem ID) in the compound database Pubchem. Since different drug identifiers from different databases have different definition rules, the present application provides a self-defined drug identifier (denoted as ID BROAD), and establishes a correspondence table between ID BROAD, the generic name of the drug, and Pubchem ID to unify the drug identifier and facilitate user operation.

[0064] The gene expression profile data of the drug can be obtained from a public database, which can be the Library of Integrated Network-based Cellular Signatures (LINCS) expression profile dataset. The LINCS expression profile dataset is obtained by detecting the expression of 978 marker genes using the L1000 technology and extrapolating the expression of other genes by constructing a model. Under the premise of reducing cost and ensuring data quality, gene expression profiles of different cell lines under three types of perturbations, including gene silencing (RNA Interference), gene overexpression (Overexpression), and small-molecule compounds (Small-molecule Compounds), are obtained and disclosed. As of July 2018, the scale of the expression profile data disclosed by the LINCS project has exceeded one million, which includes the expression profiles of 41847 small-molecule compounds under multi-gene perturbation, and 396 datasets, and the main cell perturbation expression profiles are detected on different cancer cell lines, mainly including breast cancer, colon cancer, liver cancer, lung cancer, melanoma, and prostate cancer. The gene expression profile data of a plurality of drugs with the same or similar drug action mechanism are analyzed by the prediction model to obtain the score of the gene expression profile data of each drug by the prediction model. The score refers to the probability value of predicting that each drug (second drug) has the same or similar drug action mechanism as the drug (first drug) used to establish the prediction model. Specifically, the second drug is tested under different conditions of dose / time / cell line, which produces different gene expression profiles. The gene expression profile data of a plurality of second drugs are obtained from a public database, and the number of gene expression profile data of each second drug is greater than or equal to 1. The gene expression profile data of each second drug with a number greater than or equal to 1 is analyzed by the prediction model of the first drug to obtain the corresponding probability value. The average probability value of the obtained several probability values is obtained to obtain the score of the gene expression profile data of each second drug by the prediction model of the first drug. The score is denoted as AVG PROB (Average probability that all gene expression profile of a drug is judged to be positive). Since the drug action mechanism of the second drug and the first drug is known, it can be evaluated in reverse which genes in the gene expression profile data have better prediction effect, so that the landmark genes of D100 are screened out.

[0065] The drug action mechanism can be obtained from some public drug information library. For example, the annotation information of the drug function of the drug can be obtained from the DrugRepurposing Hub information library, including the drug action mechanism. For example, the gene expression profile database is the LINCS expression profile dataset. The drug identifiers are obtained from the drug set of the MCE (Medchemexpress) company to generate a drug identifier list. Then the drug identifiers in the drug identifier list are matched with the drug expressions in the LINCS expression profile dataset, including not only complete matching of the drug name or drug identifier, but also other semantic, format and various matching modes. The gene expression profile data of each drug is obtained from the LINCS expression profile dataset. In a specific embodiment, before the deep neural network is trained using the drug action mechanism and the gene expression profile data of each drug respectively to obtain the prediction model of each drug, a plurality of drug sets can also be obtained according to the drug action mechanism of each drug. Among them, the drugs in the drug set have the same or similar drug action mechanism. Specifically, according to the drug identifiers in the drug identifier list, the drug action mechanism of each drug is obtained, and each drug is classified according to the drug name, format or compound suffix. The drugs with similar or identical drug action mechanism are collected together to form a drug set. For example, 103 drug sets with specific drug action mechanisms can be sorted from the drug set of the MCE (Medchemexpress) company. The drug action mechanism of each drug set is used as the true label, and the deep neural network is trained using the gene expression profile data of the drugs in each drug set to obtain the prediction model of each drug set. By analyzing the drug action mechanism of the gene expression profile data through the prediction model of each drug, it can be judged which gene expression profile data has consistent MOA prediction scores, so that the least but precise landmark genes can be screened and clustered. In a specific embodiment, after the deep neural network is trained using the drug action mechanism and the gene expression profile data of each drug set respectively to obtain the prediction model of each drug set, the performance indicators of the prediction model of each drug set can also be evaluated. Specifically, the performance indicators of the prediction model of each drug set are evaluated using the ROC curve and AUC value. The AUC of the ROC is to verify the performance and generalization ability of the prediction model by using the external test set. The ROC curve is also called the sensitivity curve, and the reason for this name is that each point on the curve reflects the same sensitivity, which is the response to the same drug molecule stimulation, but only the results obtained under different judgment criteria. The ROC curve is a coordinate graph composed of the false alarm probability as the horizontal axis and the hit probability as the vertical axis, and the curve drawn by the different results of the subjects under the specific stimulation conditions due to the use of different judgment criteria. The ROC curve has a good property: when the distribution of positive and negative samples in the test set changes, the ROC curve remains unchanged.In actual data sets, sample class imbalance often occurs, i.e., the positive and negative sample ratio difference is large, and the positive and negative samples in the test data may also change over time.

[0066] According to the evaluation result, a plurality of prediction models meeting a preset condition can be selected from the prediction models of each drug set. The preset condition is used to select a model with prediction value from each prediction model. For example, the preset condition can be a limitation on the AUC value, such as a model with an AUC value greater than 0.6 is a model with prediction value.

[0067] For example, according to the drug name input by the user, the corresponding gene expression profile data is obtained from the Broad Institute PHASE L1000 platform. When extracting the expression profile data from the Broad Institute PHASE L1000 platform, the user can select three drug identifiers, namely ID BROAD, Pubchem ID, and Alternative names. When the drug name input by the user does not match the drug name of the Broad Institute PHASE I L1000, a mismatch information is prompted, and which drug name input does not match is prompted. At the same time, the expression profile of only sensitive cell lines can be extracted, and 72 cell lines are provided for selection, namely A549, VCAP, ASC, PHH, PC3, HEC108, HT29, HA1E, A375, SKB, NEU, SNGM, HCC515, FIBRNPC, MCF7, HEPG2, MDAMB231, HT115, A673, PL21, OV7, MDST8, SKLU1, SNU1040, THP1, BT20, NPC, WSUDLCL2, AGS, SKM1, SKMEL1, SW620, HUH7, T3M10, SKMEL28, U937, CL34, MCF10A, NCIH1836, RMUGS, RKO, NCIH1694, SNUC4, SW480, CORL23, NEU.KCL, DV90, HEK293T, HCT116, LOVO, JHUEM2, HCC15, NOMO1, H1299, NCIH2073, NCIH596, RMGI, SNUC5, NCIH508, SKBR3, TYKNU, COV644, NKDBA, EFO27, SW948, U266, HL60, JURKAT, CD34, HS578T, HS27A, MCH58. It should be noted that the gene expression profile file with the same or similar drug gene function can also be uploaded by receiving the user. The file format can be defined, including: one row for each gene, one column for each drug, using Entrez ID for gene, using Z-Score for expression profile, etc. The submitted file needs to be verified to find possible mismatch conditions, including: incorrect file format, less than 90% matching genes. If the genes covered by the expression profile file uploaded by the user are greater than or equal to 90% of the 978 marker genes selected by the L1000 technology, the verification is passed.

[0068] The predicted output result includes four columns of information ES (Enrichment score calculated by Kolmogorov-Smirnov test), AVG PROB, P value and gene expression profile number REP in addition to the three identifiers of the drug, and the predicted output result is arranged in descending order of AVG PROB.

[0069] In the model evaluation using the ROC curve and the AUC value, the tensorflow package and the sklearn package in the python language are used. The standard for selecting the number of drugs is: 5 >= the number of drugs >= 2, setting the number of folds as the number of drugs; 10 > the number of drugs >= 5, setting the number of folds as 5; and the number of drugs >= 10, setting the number of folds as 10. The verification fold method is sklearn.model_selection.StratifiedKFold, and the specific parameters are n_splits = number of folds, shuffle = True, and random_state = 0. The model evaluation method is classifier.evaluate, classifier.predict_proba, roc_curve, classifier.predict_classes, sklearn.metrics.f1_score, etc. The prediction effect of the constructed model can also be judged by MeanROC. According to the evaluation results, multiple prediction models that meet the preset conditions are selected from the prediction models of each drug set. In a preferred embodiment, 3-fold cross-validation is used to evaluate the performance of the prediction model. In cross-validation, the positive and negative drug sets are randomly divided into K parts by stratified sampling. 2 / 3 of the samples are used as the training set to train the prediction model, and the sensitivity and specificity are evaluated by testing the remaining 1 / 3 of the samples. This process is performed 3 times, and the average AUROC (area under the receiver operating characteristic curve) is used as the evaluation index, and the probability scores of the model for the negative and positive sample sets in each cross-training process are recorded. In this process, the model with an average AUROC >= 0.6 is considered to be a well-trained model.

[0070] For the user-uploaded drug gene expression profile file, it can be processed, such as median filling for missing genes, and the specific calling method is the sklearn.impute.SimpleImputer package of python.

[0071] The number of clusters can not only be 100, but also for example, 90-150. The smaller the number of clusters, the greater the error of the prediction result. The number of clusters is not as large as possible, but is selected in a range where the error is acceptable, and the test steps and costs can be greatly saved.

[0072] It should be understood that, although the steps in the above flowchart are shown in sequence according to the direction of the arrows, these steps are not necessarily executed in the order according to the direction of the arrows. Unless explicitly stated herein, there is no strict order limitation for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowchart can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be round-robin or alternately executed with at least a part of other steps or sub-steps or stages of other steps.

[0073] 2. MOA prediction and accuracy analysis of L1000 GPAR, D100 GPAR and L1000 FWD

[0074] After the GPAR prediction model (based on L1000 and D100) is established, by inputting the target expression profile into the model of each MOA, the probability of the target expression profile belonging to the classification of the model is calculated, and the top ten MOA predictions are retained in descending order. The MOA prediction of L1000 FWD is crawled from its website, containing 4,938 molecules, and each molecule retains the top 10 MOA with the highest prediction score.

[0075] Experimental Results

[0076] 1. Screening of D100 landmark genes and performance test of drug prediction based on D100

[0077] The number of features N of the training set is tested from the interval [10, 800]. In the interval of [10, 100] and in the interval of [100, 800], N genes are randomly extracted for 10 times, and the GPAR model is trained and evaluated. By calculating the area under the receiver operating characteristic curve (AUROC) and the average precision score (AP score) of GPAR and cosine similarity (CS) respectively, it is found that the AUROC and AP score of GPAR and CS are improved with the increase of the number of genes, but when the number of features exceeds 100, the performance improvement slows down. Considering that if the number of input genes is too small, it will significantly affect the prediction performance, and if the number of input genes is too large, it will also lead to a significant increase in detection cost, therefore, in order to ensure better prediction performance while having lower detection cost, 100 genes are finally selected as input. Then, by analyzing the co-expression characteristics of the L1000 landmark genes on 310,000+ expression profiles, and using the K-means algorithm (K=100) to screen 100 co-expression clusters, the center point genes of each co-expression cluster are selected for combination, thereby screening out immobilized 100 genes, that is, the inventors have unexpectedly found that the results obtained by the above-mentioned prediction model using clustering method are very stable, and the 100 genes screened each time are the same, which also proves the stability of the drug efficacy prediction method using 100 genes expression profile of the present application, and thus the inventors name the technology of detecting the expression of the 100 genes as D100.

[0078] Table 1 D100 landmark gene information table

[0079]

[0080]

[0081]

[0082] Table 2 D100 detection system upstream and downstream probe information

[0083] (wherein the capture probe is an antisense Barcode sequence probe, the left probe is an upstream probe, and the right probe is a downstream probe)

[0084]

[0085]

[0086]

[0087]

[0088]

[0089]

[0090]

[0091]

[0092]

[0093]

[0094]

[0095] Figures 1A-1F is a schematic diagram of the screening process and predictive performance test of D100 genes of the present application, wherein Figure 1A , 1B is the predictive performance of GPAR and CS when the number of landmark genes is different; Figure 1C is the difference between the MOA positive set and the MOA negative set when 10-900 landmark genes are predicted; Figure 1D is the co-expression relationship matrix of 978 landmark genes of L1000 technology in 310,000+ expression profiles, and the genes with co-expression relationship are divided into 100 cluster clusters by K-means clustering; Figure 1E is the data visualization of D100 features and L1000 features (the 9 MOAs with the highest prediction performance of D100 and L1000 are colored and displayed in the figure); Figure 1F is the comparison of the prediction performance of D100 data and L1000 data. In order to further verify the prediction performance of the fixed 100 genes (D100), the inventors selected 103 drug MOAs to test the GPAR modeling of D100, and compared it with the GPAR modeling test results of L1000. The results show (as shown in Figures 1A-1FAs shown in Table 2, although the overall prediction performance of the D100-based GPAR model was not as good as that of the L1000-based GPAR model, there were still 80 drugs with AUROC > 0.6, which was similar to the number of effective GPAR modeling based on L1000 (83), indicating that the D100-based GPAR modeling could effectively predict most drugs with specific MOA. Further, the inventors selected 284 drugs with known MOA to evaluate the accuracy of MOA prediction. On the one hand, the expression profiles of the 284 drugs were scored and ranked for MOA using the D100-based GPAR model and the L1000-based GPAR model. On the other hand, the MOA prediction results of the 284 drugs based on the CS similarity between L1000 expression profiles were obtained using the official tool L1000FWD of LINCS. Then, the MOA predicted by the three methods was compared with the true MOA of the drugs to evaluate the accuracy. The results are shown in Table 3. Figure 2A 、 2B As shown in Table 3, among the 284 drugs, the L1000-based GPAR model accurately predicted the MOA of 139 drugs, with the highest accuracy (48.9%), while the D100-based GPAR model accurately predicted the MOA of 124 drugs, with a prediction accuracy (43.6%) similar to that of L1000FWD (41.2%).

[0096] II. Technical principles and implementation schemes of D100

[0097] 1、 Detection principle:

[0098] The detection principle of D100 technology is shown in Figure 1. Figure 3 Specifically, the mRNA of the sample to be tested is captured and reverse transcribed to generate a cDNA first strand on a PCR plate (such as a TurboCapture plate) coated with a dT primer. Then, an upstream probe and a downstream probe are synthesized using a specific fragment of each target nucleic acid (the upstream probe contains a T7 universal primer sequence, a Barcode sequence, and an upstream primer specific to the target nucleic acid; the downstream probe contains a downstream primer specific to the target nucleic acid and a T3 universal primer sequence), and the probes are hybridized to the target cDNA. When both the upstream and downstream probes are connected to the target cDNA, the upstream and downstream probes can be connected together to form a ligation hybrid. Then, the ligation hybrid is used as a template, the T7 universal primer sequence with a biotin label is used as a forward primer, and the reverse T3 universal primer sequence is used as a reverse primer for amplification. Finally, the amplified product labeled with biotin is hybridized to magnetic beads coupled with the corresponding antisense Barcode sequence (i.e., capture probe), and a biotin-streptavidin cascade amplification reaction principle is used for SAPE staining. Finally, the average fluorescence signal value on each magnetic bead is detected using a Luminex 3D instrument to obtain the expression information of each gene.

[0099] The decomposition step, as shown, includes: Figure 3

[0100] (1) Prepare a TurboCapture plate coated with dT primers, leaving a certain number of wells as blank wells or control wells;

[0101] (2) Capture: The TurboCapture plate captures mRNA in the sample to be tested;

[0102] (3) Reverse transcription: With oligo(dT) as a primer, mRNA as a template, and under the action of reverse transcriptase, free dNTP in the solution is connected to form a first strand of cDNA (40 pb gene-specific sequence);

[0103] (4) Annealing: Use the upper and lower probes with gene-specific sequences to bind to complementary landmark genes, for example, the upper probe is a universal primer sequence 1 + Barcode sequence, and the lower probe is a universal primer sequence 2;

[0104] (5) Ligation: Based on the dehydration condensation reaction, use DNA ligase to connect the upper and lower probes bound to the same cDNA strand;

[0105] (6) Amplification: Based on the three steps of denaturation, annealing, and extension of PCR, use T7 PCR universal primers with biotin labels and T3 PCR universal primers to amplify the connected upper and lower probes;

[0106] (7) Hybridization: Prepare the hybridization reaction system according to the hybridization reaction mixture (D6) construction scheme;

[0107] (8) Staining: Use magnetic separation of Luminex magnetic beads to perform the washing step, and use biotin-streptavidin cascade amplification reaction to fluorescently stain the hybridization products;

[0108] (9) Detection: Use Luminex's dual-color laser to identify the type of magnetic bead and detect the median fluorescence signal intensity value on the surface of each magnetic bead; wherein the carboxyl magnetic bead becomes an activated carboxyl magnetic bead after encountering the antisense Barcode sequence probe and EDC, and further generates a coupled magnetic bead.

[0109] In this way, by making the above D100 gene into a TurboCapture plate (kit) and detecting the median fluorescence signal intensity value on the surface of each magnetic bead after the above steps, the corresponding expression data can be obtained.

[0110] 2. Reagent consumables and instruments:

[0111] Table III Consumables

[0112]

[0113] Table Four Reagents

[0114]

[0115]

[0116] Table Five Instruments

[0117]

[0118] Table Six Mix Configuration

[0119] Reverse transcription mix (D1)

[0120]

[0121]

[0122] Ligation mix (D3)

[0123]

[0124] PCR mix (D4)

[0125]

[0126] 1.5x TMAC hybridization solution (D5)

[0127]

[0128] Hybridization reaction mix (D6)

[0129]

[0130] Low strength buffer (D7)

[0131]

[0132]

[0133] High strength buffer (D8)

[0134]

[0135] SAPE mix (D9)

[0136]

[0137] 3. Sample preparation

[0138] Cells in exponential growth phase after passage were digested with 0.25% trypsin (B19) to prepare a cell suspension, which was counted using a cell counter (C3) and inoculated into 96-well cell culture plates (A2) at 10,000-20,000 cells per well (to achieve a 90% cell confluence in the 96-well plates after 24 hours), and cultured in medium (B17) containing 10% FBS (B18) for 24 hours, then the corresponding test drugs were added (with serum culture) and cultured for 6 hours or 24 hours, the supernatant was discarded, and the cells were washed twice with PBS (B20). Then 86 μl of TCL lysis solution (TurboCapture 96 mRNA Kit (B2) provided by the manufacturer) was added to each well, and the plate was placed on a shaker for 30 minutes. The plate was stored in a -80°C refrigerator (C5) for later use.

[0139] 4. Magnetic bead coupling:

[0140] Objective: Coupling of antisense Barcode sequence to Luminex magnetic beads (A12) (one Barcode sequence corresponds to one Luminex magnetic bead)

[0141] Principle: Condensation reaction of carboxyl and amino groups to connect the antisense Barcode sequence probe (-NH2 labeled) to the surface of the magnetic beads (-COOH). First, the carboxyl group forms a -CO-EDC (carboxyl activation process) with EDC (B22) under the condition of pH 4.5, then the -NH2 labeled antisense Barcode sequence probe is added to form an amide bond.

[0142] 5. Experimental steps

[0143] Remove the Luminex magnetic beads, and ultrasonicate and vortex for 20 seconds.

[0144] Take 200 μl (2.5 x 106) of each magnetic bead and place it in a low-protein adsorption PCR tube (A3). After magnetic separation using a magnetic separator (C7), the supernatant is discarded.

[0145] Add 25 μl of 0.1 M MES (B21) (pH 4.5) to suspend the magnetic beads, and add 1 μl (100 μM) of the corresponding antisense Barcode sequence probe.

[0146] Fresh and dry EDC powder (B22) is taken from -20°C and dissolved to a concentration of 10 mg / ml (freshly prepared 10 mg / ml EDC solution using dH2O).

[0147] Add 2.5 μl of EDC solution to the magnetic bead solution (10 μg or 1 μg / μl final, discarded after use), vortex to mix, and incubate at room temperature in the dark for 30 minutes.

[0148] Add 2.5 μΐ of freshly prepared 10 mg / mL EDC solution to the magnetic bead solution (10 μg or 1 μg / μΐ final, discard after use) again, vortex mix, and incubate at room temperature for 30 min in the dark.

[0149] Add 400 μΐ of 0.02% Tween-20 (B23) to the reaction solution, mix, and after magnetic separation, aspirate the supernatant.

[0150] Add 400 μΐ of 0.1% SDS (B24) solution, mix, and after magnetic separation, aspirate the supernatant.

[0151] Add 400 μΐ of 1 x TE buffer (pH 8.0), mix, and after magnetic separation, aspirate the supernatant.

[0152] Resuspend the coupled magnetic beads in 50 μΐ of 1 x TE (pH 8.0) buffer.

[0153] Take 1 μΐ of the coupled magnetic bead stock solution, dilute it 100-fold, and use a hemocytometer (A4) to count the beads and record the data. The final coupled magnetic bead stock solution is stored at 2-8 °C in the dark.

[0154] Prepare an equal amount of the coupled magnetic bead mixture before each use (the prepared coupled magnetic bead mixture must be used up within one week, otherwise it is discarded).

[0155] 6. Ligation-mediated amplification

[0156] § capture

[0157] Objective: Isolate cellular mRNA.

[0158] Principle: Since eukaryotic mRNA has a unique polyA tail structure, this part mainly uses the base complementary pairing between oligo(dT) primer and the polyA tail structure of mRNA to isolate cellular mRNA.

[0159] Experimental steps:

[0160] Take the sample cell lysate from the -80 °C refrigerator (C5), and equilibrate to room temperature.

[0161] Mix the cell lysate well, and transfer it to the TurboCapture plate (A8) (20 μΐ / well) (reserve a certain number of wells as blank wells or control wells as needed).

[0162] Seal the above TurboCapture plate with a silica gel sealing film (A7)

[0163] Quickly spin the TurboCapture plate (1000 r / min, 1 min) at room temperature using the centrifuge (C9).

[0164] After the TurboCapture plate is free of bubbles, incubate at room temperature for 60-90 min.

[0165] § reverse transcription

[0166] Objective: Synthesize the first strand of cDNA.

[0167] Principle: Use oligo(dT) as a primer to connect free dNTP in solution to form the first strand of cDNA under the action of reverse transcriptase with mRNA as a template.

[0168] Experimental steps:

[0169] Remove the TurboCapture plate from the above room temperature and cover it with a water-absorbing cloth (A6).

[0170] Invert the TurboCapture plate and quickly spin the TurboCapture plate (1000 r / min, 1 min) at room temperature using the centrifuge (C9) to remove the cell lysate.

[0171] Remove the water-absorbing cloth covering the TurboCapture plate and add 25 μl of reverse transcription mixture (D1) to each well using the fully automatic pipetting workstation (C8).

[0172] Seal the TurboCapture plate with a silica gel sealing plate film (A7).

[0173] Quickly spin the TurboCapture plate (1000 r / min, 1 min) at room temperature using the centrifuge (C9).

[0174] After the TurboCapture plate is free of bubbles, incubate at 37°C for 90 min using the gradient PCR instrument (C10).

[0175] § annealing

[0176] Objective: Combine the complementary landmark genes with upstream and downstream probes.

[0177] Principle: Combine the complementary landmark genes with upstream and downstream probes with gene-specific sequences.

[0178] Experimental steps:

[0179] Remove the TurboCapture plate from the above 37°C and cover it with a water-absorbing cloth.

[0180] Invert the TurboCapture plate and spin the TurboCapture plate at room temperature (1000 rpm, 1 min) to remove the cell lysis solution.

[0181] Remove the blotting paper covering the TurboCapture plate and add 20 μΐ of annealing mix (D2) to each well using the automated pipetting station.

[0182] Seal the TurboCapture plate with a silicon sealing membrane.

[0183] Spin the TurboCapture plate at room temperature (1000 rpm, 1 min) using the centrifuge.

[0184] After the TurboCapture plate is free of air bubbles, incubate the plate at 95°C for 5 min using the gradient PCR machine, then decrease the temperature from 70°C to 40°C at a constant rate over 6 h, and finally keep the plate at 4°C until the next step (the plate can also be incubated overnight at the optimal temperature of the probe (55°C) after the temperature program is completed).

[0185] § ligation

[0186] Purpose: To ligate the upstream and downstream probes bound to the same cDNA strand.

[0187] Principle: Based on the dehydration condensation reaction, ligate the upstream and downstream probes bound to the same cDNA strand using DNA ligase.

[0188] Experimental procedure:

[0189] Remove the blotting paper covering the TurboCapture plate and add 20 μΐ of annealing mix (D2) to each well using the automated pipetting station.

[0190] Invert the TurboCapture plate and spin the TurboCapture plate at room temperature (1000 rpm, 1 min) to remove the cell lysis solution.

[0191] Remove the blotting paper covering the TurboCapture plate and add 20 μΐ of annealing mix (D2) to each well using the automated pipetting station.

[0192] Seal the TurboCapture plate with a silicon sealing membrane.

[0193] Spin the TurboCapture plate at room temperature (1000 rpm, 1 min) using the centrifuge.

[0194] After no air bubbles in the TurboCapture plate, use gradient PCR instrument to incubate at 45°C for 60 min, and at 75°C for 10 min.

[0195] § PCR

[0196] Purpose: Amplification of the connected upstream and downstream probes.

[0197] Principle: Based on the three steps of denaturation, annealing and extension of PCR, the connected upstream and downstream probes are amplified by T7 PCR universal primer with biotin label and T3 PCR universal primer.

[0198] Experimental steps:

[0199] Take the TurboCapture plate out of the above 75°C, and cover it with a water-absorbing cloth.

[0200] Invert the TurboCapture plate and use the centrifuge to quickly rotate the TurboCapture plate at room temperature (1000 r / min, 1 min) to remove the cell lysate.

[0201] Remove the water-absorbing cloth covering the TurboCapture plate, and use the full-automatic pipetting workstation to add 50 μl of PCR mixture (D4) to each well.

[0202] Use aluminum sealing plate film (A11) to seal the TurboCapture plate.

[0203] Use the centrifuge to quickly rotate the TurboCapture plate at room temperature (1000 r / min, 1 min).

[0204] After no air bubbles in the TurboCapture plate, put it into the gradient PCR instrument, and perform the following PCR standard scheme. After completion, transfer it to the -20°C refrigerator (C11) for frozen preservation.

[0205] PCR standard scheme

[0206]

[0207] § hybridization detection

[0208] Experimental steps:

[0209] Take the TurboCapture plate out of the above -20°C refrigerator and thaw it at room temperature.

[0210] From each well of the TurboCapture plate, 5 μl was pipetted into the corresponding well of a new 96-well PCR plate, and the hybridization reaction mixture (D6) was prepared according to the protocol.

[0211] The plate was sealed with aluminum sealing film (Al) to prevent evaporation of water. The temperature was controlled using a constant temperature mixer (C12), and the hybridization program was 96°C, 2 min; 45°C, 60 min (the time can be extended); and 800 r / min shaking throughout.

[0212] § washing, staining

[0213] Objective: To remove the unhybridized PCR product, and to stain it with SAPE staining agent, and to remove the SAPE before detection.

[0214] Principle: The washing step was performed using magnetic separation of Luminex magnetic beads. The hybridization product was fluorescently stained using a biotin-streptavidin cascade amplification reaction.

[0215] Experimental procedure:

[0216] First washing: The conjugated magnetic beads were adsorbed using a magnetic plate (C13), and 35 μl of supernatant was discarded. 60 μl of low-concentration buffer (D7) was added, and the mixture was thoroughly mixed using a constant temperature mixer at 45°C and 600 r / min shaking. The conjugated magnetic beads were adsorbed using a magnetic plate, and 60 μl of supernatant was discarded. 60 μl of high-concentration buffer (D8) was added, and the mixture was thoroughly mixed using a constant temperature mixer at 45°C and 600 r / min shaking. The conjugated magnetic beads were adsorbed using a magnetic plate, and 60 μl of supernatant was discarded.

[0217] Staining: 50 μl of SAPE mixture (D9) was added and mixed, and the mixture was incubated at room temperature for 15 min in the dark. The conjugated magnetic beads were adsorbed using a magnetic plate, and 50 μl of supernatant was discarded.

[0218] Second washing: 60 μl of low-concentration buffer was added, and the mixture was thoroughly mixed using a constant temperature mixer at 45°C and 600 r / min shaking. The conjugated magnetic beads were adsorbed using a magnetic plate, and 60 μl of supernatant was discarded. 60 μl of sheath liquid (B29) was added, and the mixture was thoroughly mixed using a constant temperature mixer at 45°C and 600 r / min shaking. The conjugated magnetic beads were adsorbed using a magnetic plate, and 60 μl of supernatant was discarded. 60 μl of sheath liquid was added for detection.

[0219] § detection

[0220] Objective: To obtain the median fluorescence signal intensity value on each conjugated magnetic bead.

[0221] Principle: The Luminex dual-color laser was used to identify the type of magnetic beads and detect the median fluorescence signal intensity value on the surface of each magnetic bead.

[0222] Experimental steps:

[0223] Turn on the Luminex 3D instrument, preheat for 30 min, and calibrate the instrument using the FLEXMAP 3D Calibration Kit (B30) and the FLEXMAP 3D Performance Verification Kit (B31).

[0224] Set the detection conditions (magnetic microspheres; detection volume 60 μl; detection temperature 45℃; magnetic bead type according to the use of magnetic bead type; magnetic bead detection number 100 (the minimum settable number is 50); detection gate 7500-15000).

[0225] After detection, follow the instrument usage method to perform the routine pre-shutdown operation.

[0226] The application also discloses a drug prediction method based on D100 (a drug efficacy prediction method based on a small amount of gene expression profiles), which comprises the following steps:

[0227] The detected gene types are reduced to N genes by using the high correlation between gene expressions, and then the expression profiles of the N genes are detected and calculated to predict the potential action mechanism of the drug. The detection principle is to use gene-specific probes containing universal primer sequences and coding sequences for ligation-mediated amplification (LMA), and combine Luminex liquid chip technology and Luminex magnetic beads to detect the expression data of the 100 genes, so that the whole D100 technology is more suitable for large-scale data acquisition. The calculation prediction principle is to use the LR model to construct a classification prediction model for the correlation between the drug and the gene expression profile, train the classification prediction model by using the characteristic that drugs with the same or similar drug action mechanism have the same or similar efficacy, and use the trained classification prediction model to predict whether the drug has efficacy. The N genes at least include the genes screened by the above D100 technology, and can also include part of other genes in L1000, but the number of genes detected is much less than that of L1000, for example, no more than 2 / 3 of L1000, no more than 1 / 3, or even no more than 150, so that the number of gene detection and processing can be greatly reduced.

[0228] The application also discloses a kit using the above D100 technology, which uses gene-specific probes containing only universal primer sequences and coding sequences of the above D100 genes to detect the expression data of the N genes, so that fewer detection genes can be used to realize drug prediction.

[0229] More specifically, the kit comprises:

[0230] PCR plate coated with poly dT primer for capturing mRNA in the sample to be tested;

[0231] Upstream and downstream probe set containing upstream and downstream probes of N detection genes screened by the method described above, for specifically capturing the amplified target detection genes present in the sample to be tested on the PCR plate;

[0232] Result presentation unit for presenting the detection results in the form of fluorescence based on the results of capturing the amplified target detection genes by the upstream and downstream probe set.

[0233] In a preferred embodiment, the kit further comprises:

[0234] Cell lysis solution containing 1% β-mercaptoethanol for lysing cells to prepare crude cell lysate; and / or

[0235] Liquid chip coupled with corresponding capture probes for detecting the prepared PCR products.

[0236] In a preferred embodiment, the result presentation unit comprises Luminex magnetic beads for coupling the captured amplified target detection genes, and the surface emits fluorescence signals with corresponding intensity values.

[0237] The present application will be further described in detail below with reference to the examples, but the scope of the present application is not limited thereto. Obviously, the examples described below are only some of the embodiments of the present application, but not all the embodiments; based on the examples of the present application described below, all the embodiments obtained by those skilled in the art without creative labor fall within the scope of the present application.

[0238] Example of Embodiment

[0239] I. Information of the detection compound:

[0240]

[0241]

[0242] II. Cells

[0243] Human malignant melanoma cells (A375) were purchased from the China Typical Culture Collection Cell Library and stored in our laboratory.

[0244] III. Experimental materials and experimental procedures

[0245] See the experimental technical solution described above for screening D100. The drug action time is 6H and 24H, respectively. Each drug has three biological replicates.

[0246] IV. Experimental results

[0247] Reliability and Stability Analysis of D100 Data

[0248] Figures 5A-5C Figure 6A is a chart of reliability and stability analysis of D100 technology detection data after 6H of drug action of the present application; wherein Figure 5A Figure 6B is a chart of background value analysis and pollution analysis; Figures 6A-6D Figure 6C is a chart of reliability and stability analysis of D100 technology detection data after 24H of drug action of the present application; wherein Figure 6A Figure 6D is a chart of background value analysis and pollution analysis; the median signal value analysis of the blank wells found (as shown in Figure 5A and Figure 6A ), all of the coupled magnetic beads signal values were lower than the background value (MFI≦200), suggesting that there was no aerosol pollution in this experiment. Figure 5B Figure 6E is a chart of inter-well fold change analysis; Figure 5C Figure 6F is a chart of CV value analysis; Figure 6G is a chart of inter-well fold change analysis; Figure 6C Figure 6H is a chart of CV value analysis; further inter-well difference and coefficient of variation analysis found (as shown in Figure 5B , Figure 5C and Figure 6B , Figure 6C ), the inter-well difference of this data was ≦2, and the coefficient of variation was ≦20%, suggesting that this data had good stability. Figure 5D Figure 6I is a chart of control well similarity analysis; Figure 6D Figure 6J is a chart of control well similarity analysis; finally, the solvent control similarity analysis found (as shown in Figure 5D and Figure 6D ), the similarity scores of all solvent controls were above 0.80, further indicating that this data had good stability. The above results indicated that D100 technology had good stability and reliability.

[0249] Accuracy of D100 Data

[0250] In order to detect the prediction accuracy of D100, the present inventors randomly selected 16 drugs containing high, medium and low transcriptional activities for detection. The cosine similarity matrix between all data expression profiles was subjected to hierarchical clustering analysis (as shown in Figure 7 and Figure 8 ), drugs with the same or similar functions had higher correlation in expression profiles, and therefore tended to form clusters with closer relationships. Further, GPAR was used to analyze the functional prediction of drug expression profiles detected by D100 after 6H and 24H of drug action, and compared with the drug annotation function, the results (as shown in Figure 9 and Figure 10When the drug effect is 6H, 12 of the 16 detected drugs are predicted to have their original functions, and the detection accuracy is about 75%. When the drug effect is 24H, 12 of the 16 detected drugs are predicted to have their original functions, and the detection accuracy is about 75%. The results of 24H and 6H are combined, and it is found that 14 of the 16 detected drugs are predicted to have their original functions, and the prediction accuracy is 87.5%.

[0251] The above specific embodiments further illustrate the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A screening method for a gene profile for drug prediction, characterized by, The method comprises the following steps: 978 landmark genes in the public data of L1000 are used to construct a 978*978 gene expression correlation matrix; The 978 genes are clustered into N classes by using a clustering method, and a center gene is selected in each class cluster, that is, the gene with the highest average correlation with other genes in the same cluster is selected as a landmark gene; wherein, 90≤N≤150; The selection of the center gene uses the co-expression characteristics of the genes to select a representative center gene as a representative gene of the co-expression gene, thereby constructing a de-redundant expression matrix; The screening process uses an LR model to establish and train an MOA prediction model reflecting the relationship between the drug and the gene expression profile data of the drug, the MOA prediction model is a binary classification model, and the score refers to the probability value of predicting whether each drug has the same or similar drug action mechanism as the drug for establishing the prediction model by using the prediction model; The binary classifier training set includes two class samples: "positive set" and "negative set", the "positive set" label refers to the drug MOA annotation of the MedChemExpress library and the Drug Repurposing Hub, and the "negative set" selects compounds with low transcriptional activity and without MOA annotation, and it is assumed that they have no drug properties that can be reflected at the transcriptome level.

2. The screening method according to claim 1, wherein The clustering method uses a K-means algorithm and a cosine similarity measure.

3. The screening method according to claim 2, wherein In the clustering method, a t-distributed stochastic neighbor embedding algorithm is used, and the data is initialized by principal component analysis.

4. The screening method according to claim 2, wherein The performance of the MOA prediction model is evaluated by using 3-fold cross-validation; when cross-validation, the positive and negative drug sets are randomly divided into K parts by using stratified sampling method, 2 / 3 samples are used as the training set to train the MOA prediction model, and the sensitivity and specificity are evaluated by testing the remaining 1 / 3 samples; this process is performed for 3 times, and the average AUROC of 3 times is used as the evaluation index, and the probability score of the model to the negative and positive sample sets in each cross-training process is recorded, and the model with an average AUROC≥0.6 is regarded as a well-trained model in this process.

5. The screening method according to claim 1, wherein The annotation information of the drug function of the drug is obtained from the Drug Repurposing Hub information database, including the drug action mechanism; and / or The gene expression profile data of the drug is obtained from the LINCS expression profile data set; and / or 103 drugs with specific drug action mechanisms are sorted from the drug set of the Medchemexpress company; The mechanism of action of each drug set is taken as a true label, and the MOA prediction model is trained using the gene expression profile data of the drugs in each drug set to obtain a prediction model for each drug set; then the gene expression profile data is analyzed for the mechanism of action of each drug by the prediction model of each drug, and the MOA prediction score of the gene expression profile data is sorted to screen and cluster the landmark genes.

6. The gene set obtained by the screening method according to any one of claims 1-5 for drug prediction.

7. The gene set of claim 6, wherein, The gene set is: RNMT, TOPBP1, CBR3, IL1B, HADH, DHRS7, UBE2J1, NUDT9, CASC3, PGRMC1, KDM5B, DAG1, NUP62, CCNA2, NUP88, ALAS1, FAH, LYN, TRAPPC6A, MEST, NENF, GDPD5, HSPA1A, ICAM3, DNMT1, CDC25A, TSC22D3, PCMT1, SCARB1, BLVRA, POLR2K, KIAA0196, GFPT1, GAA, SLC35B1, LIG1, IKBKB, LYPLA1, SKP1, UBE3C, PRAF2, DDB2, AKAP8, IER3, FOXJ3, AKAP8L, GATA2, FBXO21, DECR1, PTPRF, RAE1, SPTLC2, SMARCA4, USP22, PSME1, HMOX1, CALM3, G3BP1, HSD17B10, AURKA, GADD45A, TSKU, ARNT2, CCNB2, IGF2R, CDK1, AURKB, DRAP1, CCNF, DERA, IFNAR1, DNM1, FOXO3, ASAH1, FOXO4, SCAND1, GABPB1, CLIC4, CALU, MAP3K4, RSU1, ALDH7A1, TBP, BUB1B, ABCF1, ANO10, PCBD1, PSMG1, ITGB5, NIPSNAP1, MYBL2, SACM1L, DNTTIP2, INSIG1, MTHFD2, TRIB3, RBM15B, ECH1, SCRN1, DUSP4.

8. A method for drug prediction using the expression profile of the gene set according to claim 6 or 7, characterized in that, The method comprises the following steps: A classification prediction model for the association between drugs and gene expression profiles is established using at least the genes in the gene set according to claim 6 or 7, and the total number of genes is not more than one-third of L1000, the classification prediction model is trained using the characteristic that drugs with the same or similar mechanism of action have the same or similar efficacy, and the classification prediction model is used to predict whether a drug has efficacy.

9. The method for drug prediction using the expression profile of the gene set according to claim 8, characterized in that, The total number of genes is not more than 150 genes.

10. A kit characterized in that, The kit uses universal primer sequences and gene-specific probes encoding sequences of at least the genes in the gene set according to claim 6 or 7, and the total number of genes is not more than one-third of L1000, so that fewer detection genes can be used to achieve drug prediction.

11. The kit of claim 10, wherein the total number of genes is not more than 150 genes. The kit comprises:

12. The kit of claim 10, wherein PCR plates coated with poly-dT primers for capturing mRNA in the sample to be tested; an upstream and downstream probe set containing upstream and downstream probes of the gene set according to claim 6 or 7 for specifically capturing the amplified target detection genes present in the sample to be tested on the PCR plate; a result presentation unit for presenting the detection results in the form of fluorescence based on the results of capturing the amplified target detection genes by the upstream and downstream probe set. The kit further comprises:

13. The kit of claim 10, wherein cell lysis solution containing 1% β-mercaptoethanol for lysing cells to prepare crude cell lysate; and / or liquid chip coupled with corresponding capture probes for detecting the prepared PCR products.

14. The kit of claim 12, wherein the result presentation unit comprises Luminex magnetic beads for coupling to capture the amplified target detection genes, and the surface emits fluorescence signals with corresponding intensity values. The kit comprises the following steps: using PCR plates coated with dT primers to capture mRNA in the sample to be tested and reverse transcribing to generate cDNA first strand; 15. A method of using a kit as claimed in any one of claims 10 to 13, characterised in that, synthesizing upstream and downstream probes using specific fragments of each target nucleic acid, wherein the upstream probe contains a T7 universal primer sequence, a Barcode sequence, and an upstream primer specific to the specific fragment of the target nucleic acid, and the downstream probe contains a downstream primer specific to the specific fragment of the target nucleic acid and a T3 universal primer sequence; hybridizing the upstream and downstream probes to the target cDNA formed by reverse transcription to form a ligation hybrid; using the ligation hybrid as a template, a biotin-labeled T7 universal primer sequence as a forward primer, and a reverse T3 universal primer sequence as a reverse primer for amplification; hybridizing the biotin-labeled amplification product with magnetic beads coupled with corresponding antisense Barcode sequences, and performing SAPE staining using biotin-streptavidin cascade amplification reaction principle; using Luminex 3D instrument to detect the average fluorescence signal value on each magnetic bead to obtain the expression information of each gene. ​ ​

Citation Information

Patent Citations

  • Construction method and equipment of predictive cell aging model, medium and program product

    CN120853690A