Drug Function Prediction Method, Device, Equipment and Storage Medium
By using gene expression profile correlation coefficient vectors and drug gene expression correlation change values in drug function prediction, the positive similarity score is calculated, and the problem of low accuracy in drug function prediction in the prior art is solved, and more accurate drug function prediction is achieved.
Patent Information
- Application Number
- CN202410739807.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-07
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-06-07
AI Technical Summary
The prior art has low accuracy in drug function prediction, making it difficult to accurately identify the deep expression relationship between gene characteristics, especially when there are large differences in the biological environment of cells.
By determining the correlation coefficient vectors of the experimental group and the control group based on the gene expression profiles of each cell line, the change value of drug gene expression correlation is calculated, the gene characteristics and reference examples are constructed, and the positive similarity score is calculated to predict the function of the drug to be detected.
Improves the accuracy of drug function prediction and enables cross-cell line prediction of drug action among different cell lines.
Smart Images

Figure CN118629500B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method, device, equipment and storage medium for predicting drug functions. Background Art
[0002] Currently, there are many defects in the traditional means of drug research and development, such as low throughput, long time, high failure rate, and high cost. The virtual screening technology based on calculation has gradually become popular in drug screening, reducing the number of small molecules that need to be experimentally verified, lowering the research and development cost, and shortening the research and development cycle. Currently, there are various methods applied in drug development. Among them, the feature matching method based on gene expression data reveals the similarity and mechanism of action between drugs by comparing the gene expression features caused by unknown drug perturbations with the gene features in a large-scale reference database, providing a powerful tool for drug discovery in theory. However, in practical applications, it shows certain limitations. Specifically, the existing feature matching methods only focus on the fold change of gene expression, which makes the constructed gene features unable to reflect the deep expression relationship between drugs acting on genes. In the case of large differences in the biological environment of cells, it is usually difficult to accurately identify the similarity relationship between two gene features, thus affecting the accurate prediction of drug effects and further resulting in low accuracy in predicting drug functions.
[0003] The above content is only used to assist in understanding the technical solution of the present invention, and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main object of the present invention is to provide a method, device, equipment and storage medium for predicting drug functions, aiming to solve the technical problem of low accuracy in predicting drug functions in the prior art.
[0005] To achieve the above object, the present invention provides a method for predicting drug functions, which includes the following steps:
[0006] Determine the experimental group correlation coefficient vector and the control group correlation coefficient vector according to the gene expression profiles of each cell line;
[0007] Calculate the change value of drug gene expression correlation according to the experimental group correlation coefficient vector and the control group correlation coefficient vector;
[0008] Construct gene features and reference instances according to the change value of drug gene expression correlation;
[0009] Calculate the positive similarity score between the gene features and the reference instances, and predict the function of the drug to be detected according to the positive similarity score.
[0010] Optionally, determining the correlation coefficient vector of the experimental group and the correlation coefficient vector of the control group according to the gene expression profiles of each cell line includes:
[0011] Obtaining a first cell line and a second cell line from each of the cell lines;
[0012] Treating the first cell line with the drug to be explored and treating the second cell line with a known drug;
[0013] Respectively measuring the cell gene expression level data in the treated first cell line and the cell gene expression level data in the second cell line;
[0014] Determining the gene expression profiles of each cell line according to the cell gene expression level data in the first cell line and the cell gene expression level data in the second cell line;
[0015] Calculating the correlation coefficient vector of the experimental group and the correlation coefficient vector of the control group according to the gene expression profiles.
[0016] Optionally, constructing gene features and reference examples according to the drug gene expression correlation change value includes:
[0017] Sorting the drug gene expression correlation change values of the first cell line according to a preset relationship;
[0018] Respectively extracting the top M gene pairs with the largest up-regulation and down-regulation of correlation from the sorting results;
[0019] Constructing gene features according to the top M gene pairs with the largest up-regulation and down-regulation of correlation;
[0020] Constructing reference examples according to the drug gene expression correlation change value.
[0021] Optionally, constructing reference examples according to the drug gene expression correlation change value includes:
[0022] Obtaining the drug gene expression correlation change values of each drug in the second cell line according to the drug gene expression correlation change value;
[0023] Distinguishing the positive and negative sets of the drug gene expression correlation change values of each drug in the second cell line to obtain an up-regulated gene pair set and a down-regulated gene pair set;
[0024] Respectively obtaining the gene correlation change characteristics in the up-regulated gene pair set and the down-regulated gene pair set;
[0025] Marking the gene pairs in the up-regulated gene pair set and the down-regulated gene pair set respectively according to the gene correlation change characteristics to obtain the ranking mark numbers of each gene pair;
[0026] Obtain the top N gene pairs with the smallest ranks in the up-regulated gene pair set and the top N gene pairs with the smallest ranks in the down-regulated gene pair set according to the ranking markers of each gene pair;
[0027] Construct a reference instance according to the top N gene pairs with the smallest ranks in the up-regulated gene pair set, the top N gene pairs with the smallest ranks in the down-regulated gene pair set, and the ranking markers of each gene pair.
[0028] Optionally, calculating the positive similarity score between the gene feature and the reference instance includes:
[0029] Obtain the types of known drugs;
[0030] Obtain the reference instances of all known drugs according to the types of the known drugs and the reference instance;
[0031] Generate a reference instance database according to the reference instances of all known drugs;
[0032] Calculate the positive similarity scores between the gene feature and each reference instance in the reference instance database respectively.
[0033] Optionally, the calculating the positive similarity scores between the gene feature and each reference instance in the reference instance database respectively includes:
[0034] Divide the gene pairs in the gene feature into a first gene pair set and a second gene pair set according to the positive and negative relationships of the correlation change values;
[0035] Divide the gene pairs in each reference instance in the reference instance database into a third gene pair set and a fourth gene pair set according to the positive and negative relationships of the correlation change values;
[0036] Take the intersection of the first gene pair set and the third gene pair set, and take the intersection of the second gene pair set and the fourth gene pair set;
[0037] Determine the ranks of each gene pair in the intersection in the reference instance;
[0038] Sort the gene pairs according to the ranks, and calculate the positive similarity scores between the gene feature and each reference instance in the reference instance database respectively according to the sorted gene pairs.
[0039] Optionally, the predicting the function of the drug to be detected according to the positive similarity score includes:
[0040] Sort the positive similarity scores;
[0041] Extract the top T positive similarity scores with the highest ranks from the score sorting results;
[0042] Determine the drug to be detected and the reference instance drug corresponding to the top T positive similarity scores;
[0043] Obtain the function of the reference instance drug;
[0044] Predict the function of the drug to be detected based on the function of the reference instance drug.
[0045] In addition, to achieve the above object, the present invention also provides a drug function prediction device, which includes:
[0046] A determination module, configured to determine an experimental group correlation coefficient vector and a control group correlation coefficient vector according to the gene expression profiles of each cell line;
[0047] A calculation module, configured to calculate a drug gene expression correlation change value according to the experimental group correlation coefficient vector and the control group correlation coefficient vector;
[0048] A construction module, configured to construct gene features and reference instances according to the drug gene expression correlation change value;
[0049] A prediction module, configured to calculate the positive similarity score between the gene features and the reference instances, and predict the function of the drug to be detected according to the positive similarity score.
[0050] In addition, to achieve the above object, the present invention also provides a drug function prediction device, which includes: a memory, a processor, and a drug function prediction program stored on the memory and executable on the processor, and the drug function prediction program is configured to implement the drug function prediction method as described above.
[0051] In addition, to achieve the above object, the present invention also provides a storage medium, on which a drug function prediction program is stored, and when the drug function prediction program is executed by a processor, it implements the drug function prediction method as described above.
[0052] The drug function prediction method proposed by the present invention determines the correlation coefficient vector of the experimental group and the correlation coefficient vector of the control group according to the gene expression profiles of each cell line; calculates the change value of drug gene expression correlation according to the correlation coefficient vector of the experimental group and the correlation coefficient vector of the control group; constructs gene features and reference examples according to the change value of drug gene expression correlation; calculates the positive similarity score between the gene features and the reference examples, and predicts the function of the drug to be detected according to the positive similarity score; in the above way, from the perspective of the gene network, uses the gene expression profiles of cell lines to determine the correlation coefficient vector, calculates the change value of drug gene expression correlation, then constructs gene features and reference examples, and performs drug identification between different cell lines to achieve cross-cell line prediction of drug effects, thereby effectively improving the accuracy of predicting drug functions. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 FIG. is a schematic structural diagram of a drug function prediction device in the hardware operating environment related to the embodiment solution of the present invention;
[0054] Figure 2 FIG. is a schematic flowchart of the first embodiment of the drug function prediction method of the present invention;
[0055] Figure 3 FIG. is a schematic flowchart of the second embodiment of the drug function prediction method of the present invention;
[0056] Figure 4 FIG. is a schematic overall flowchart of an embodiment of the drug function prediction method of the present invention;
[0057] Figure 5 FIG. is a schematic functional module diagram of the first embodiment of the drug function prediction device of the present invention.
[0058] The implementation, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0060] Refer to Figure 1 , Figure 1 FIG. is a schematic structural diagram of a drug function prediction device in the hardware operating environment related to the embodiment solution of the present invention.
[0061] As Figure 1As shown in the figure, the drug function prediction device may include: a processor 1001, such as a Central Processing Unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen and an input unit such as a keyboard. Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed Random Access Memory (RAM) or a stable Non-Volatile Memory (NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0062] Those skilled in the art can understand that Figure 1 the structure shown in the figure does not constitute a limitation on the drug function prediction device, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0063] As Figure 1 shown, the memory 1005, as a storage medium, may include an operating system, a network communication module, a user interface module, and a drug function prediction program.
[0064] In Figure 1 the drug function prediction device shown in the figure, the network interface 1004 is mainly used for data communication with a network integrated platform workstation; the user interface 1003 is mainly used for data interaction with users; the processor 1001 and the memory 1005 in the drug function prediction device of the present invention may be provided in the drug function prediction device. The drug function prediction device calls the drug function prediction program stored in the memory 1005 through the processor 1001 and executes the drug function prediction method provided in the embodiments of the present invention.
[0065] Based on the above hardware structure, an embodiment of the drug function prediction method of the present invention is proposed.
[0066] Referring to Figure 2 , Figure 2 it is a schematic flowchart of the first embodiment of the drug function prediction method of the present invention.
[0067] In the first embodiment, the drug function prediction method includes the following steps:
[0068] Step S10: Determine the correlation coefficient vector of the experimental group and the correlation coefficient vector of the control group according to the gene expression profiles of each cell line.
[0069] It should be noted that the execution subject of this embodiment is a drug function prediction device, and it can also be other devices that can achieve the same or similar functions, such as a drug function prediction device, etc. This embodiment does not limit this. In this embodiment, a drug function prediction device is taken as an example for illustration.
[0070] It should be understood that the correlation coefficient vector of the experimental group refers to the correlation coefficient vector of the experimental group in each cell line. Similarly, the control group refers to the group used for comparison with the experimental group, and the correlation coefficient vector of the control group refers to the correlation coefficient vector of the control group in each cell line. Each cell line includes a first cell line and a second cell line, and the object of the gene expression profile is the expression profiles of the experimental group and the control group in each cell line.
[0071] Further, step S10 includes: obtaining a first cell line and a second cell line from each cell line; treating the first cell line with the drug to be explored, and treating the second cell line with a known drug; respectively measuring the cell gene expression level data in the treated first cell line and the cell gene expression level data in the second cell line; determining the gene expression profiles of each cell line according to the cell gene expression level data in the first cell line and the cell gene expression level data in the second cell line; calculating the correlation coefficient vector of the experimental group and the correlation coefficient vector of the control group according to the gene expression profiles.
[0072] It can be understood that the drug to be explored refers to the drug whose function needs to be predicted. Under the specified cell background, the first cell line is treated with the drug to be explored, and the second cell line is treated with a known drug. In order to ensure that sufficient gene expression profiles are obtained, the above treatment of the cell line needs to be repeated at least three times. The concentrations used in the three repetitions can be the same or different. After the treatment is completed, the cell gene expression level data in the treated first cell line and the cell gene expression level data in the second cell line are respectively measured, and after determining the gene expression profiles, the correlation coefficient vector of the experimental group and the correlation coefficient vector of the control group are calculated. Specifically: for a certain treatment, assuming it is repeated p times and there are n genes in the gene expression profile, then for the expression matrix composed of n×p gene expression profiles, then calculate the correlation coefficient of each pair of genes. Specifically:
[0073]
[0074] Among them, r[G i ,G j represents the correlation coefficient of each pair of genes, Gi, respectively represent the expression value vectors of p dimensions. n different genes will generate (n 2 -n) / 2 correlation coefficients.
[0075] It should be understood that after obtaining multiple correlation coefficients, an experimental group correlation coefficient vector and a control group correlation coefficient vector are generated respectively according to the correlation coefficients. It should be noted that each correlation coefficient corresponds to a unique gene pair.
[0076] Step S20, calculate the change value of drug-gene expression correlation according to the experimental group correlation coefficient vector and the control group correlation coefficient vector.
[0077] It can be understood that the change value of drug-gene expression correlation refers to the gene expression correlation change value between the experimental group correlation coefficient vector and the control group correlation coefficient vector corresponding to each drug. Specifically:
[0078] V diff = V Drug - V Ctrl .
[0079] Among them, V diff represents the change value of drug-gene expression correlation, V Drug represents the experimental group correlation coefficient vector, and V Ctrl represents the control group correlation coefficient vector.
[0080] Step S30, construct gene features and reference examples according to the change value of drug-gene expression correlation.
[0081] It should be understood that gene features can be constructed through the change value of drug-gene expression correlation of the first cell line, and reference examples can be constructed through the change value of gene expression correlation of the second cell line. And for multiple reference gene pairs used to calculate the similarity score, after obtaining the change value of drug-gene expression correlation, gene features and reference examples are constructed.
[0082] Step S40, calculate the positive similarity score between the gene features and the reference examples, and predict the function of the drug to be detected according to the positive similarity score.
[0083] It can be understood that the positive similarity score refers to the positive score of the similarity between the gene features and the reference examples. The higher this positive similarity score, the more similar the gene features and the reference examples are. Then, the function of the drug to be detected is predicted using the calculated positive similarity score. For example, the function of the drug to be detected A is anti-inflammatory.
[0084] Further, calculating the positive similarity score between the gene feature and the reference instance includes: obtaining the types of known drugs; obtaining the reference instances of all known drugs according to the types of the known drugs and the reference instance; generating a reference instance database according to the reference instances of all known drugs; and respectively calculating the positive similarity scores between the gene feature and each reference instance in the reference instance database.
[0085] It should be understood that since a variety of known drugs are used when processing the second cell line, the types of the known drugs are multiple. Then, the reference instances of all known drugs are obtained by combining the constructed reference instance, and a reference instance database is generated according to the reference instances of all known drugs. Then, the positive similarity scores between the gene feature and each reference instance in the reference instance database are further calculated.
[0086] Further, respectively calculating the positive similarity scores between the gene feature and each reference instance in the reference instance database includes: dividing the gene pairs in the gene feature into a first gene pair set and a second gene pair set according to the positive and negative relationships of the correlation change values; dividing the gene pairs in each reference instance in the reference instance database into a third gene pair set and a fourth gene pair set according to the positive and negative relationships of the correlation change values; taking the intersection of the first gene pair set and the third gene pair set, and taking the intersection of the second gene pair set and the fourth gene pair set; determining the rankings of the gene pairs in the intersection in the reference instance; sorting the gene pairs according to the rankings, and respectively calculating the positive similarity scores between the gene feature and each reference instance in the reference instance database according to the ranked gene pairs.
[0087] It can be understood that taking a single calculation as an example, after obtaining the gene feature, the gene pairs in the gene feature are divided into a first gene pair set and a second gene pair set by using the positive and negative relationships of the correlation change values. For example, the first gene pair set is The second gene pair set is And the gene pairs in the reference instance are divided into a third gene pair set and a fourth gene pair set. For example, the third gene pair set is The fourth gene pair set is And the number of gene pairs in the above four gene sets is the same. Then, the intersection of the first gene pair set and the third gene pair set and the intersection of the second gene pair set and the fourth gene pair set are respectively taken, that is At this time, both u and d are gene pair sets. Among them, {u 1 , u 2 ,..., u nu} are the rankings corresponding to the gene pairs in u in the reference instance, {d 1 , d 2,...,d nd} is the rank corresponding to the gene pair in d in the reference instance, nu and nd represent the number of gene pairs in the two intersections, and then for {u 1 ,u 2 ,...,u nu} and {d 1 ,d 2 ,...,d nd} are sorted in ascending order to obtain {u (1) ,u (2) ,...,u (nu)} and {d (1) ,d (2) ,...,d (nd)}, and then {u (1) ,u (2) ,...,u (nu)} and {d (1) ,d (2) ,...,d (nd)} are input into the following formula:
[0088]
[0089] K = k up + k down .
[0090] Among them, the value of ρ can be 1, K represents the positive similarity score between the gene feature and the reference instance, and k up represents the similarity parameter after taking the intersection of the first gene pair set and the third gene pair set, and k down represents the similarity parameter after taking the intersection of the second gene pair set and the fourth gene pair set.
[0091] Furthermore, predicting the function of the drug to be detected according to the positive similarity score includes: sorting the positive similarity scores; extracting the top T positive similarity scores with the highest rankings from the score ranking results; determining the drugs to be detected and the reference instance drugs corresponding to the top T positive similarity scores; obtaining the functions of the reference instance drugs; predicting the functions of the drugs to be detected according to the functions of the reference instance drugs.
[0092] It should be understood that after obtaining multiple positive similarity scores, the multiple positive similarity scores are sorted in descending order, and the top T positive similarity scores with the highest rankings are extracted from the score ranking results. The value of T can be 1, which is the maximum positive similarity score, or an integer greater than 1, that is, multiple positive similarity scores are extracted. The score ranking results can be presented in the form of a list. The fields of the score ranking results include, but are not limited to, reference example drugs, similarity scores, rankings, etc. Then, the function of the drug to be detected is predicted using the function of the reference example drug corresponding to the maximum similarity score, that is, the closer the function of the drug to be detected is to the function of the reference example drug, the higher its ranking.
[0093] It should be noted that if exploring the treatment relationship between diseases and drugs, the negative similarity score between the gene signature and the reference example should be calculated, and when taking the intersection, the gene set needs to be swapped, that is Then continue to input the set obtained by taking the intersection into the above formula for calculation.
[0094] In this embodiment, the correlation coefficient vector of the experimental group and the correlation coefficient vector of the control group are determined according to the gene expression profiles of each cell line; the change value of the drug gene expression correlation is calculated according to the correlation coefficient vector of the experimental group and the correlation coefficient vector of the control group; the gene signature and the reference example are constructed according to the change value of the drug gene expression correlation; the positive similarity score between the gene signature and the reference example is calculated, and the function of the drug to be detected is predicted according to the positive similarity score; in the above way, from the perspective of the gene network, the correlation coefficient vector is determined using the gene expression profiles of the cell lines, the change value of the drug gene expression correlation is calculated, and then the gene signature and the reference example are constructed to identify drugs between different cell lines, realizing the cross-cell line prediction of drug effects, thereby effectively improving the accuracy of predicting drug functions.
[0095] In one embodiment, as Figure 3 shown, based on the first embodiment, the second embodiment of the drug function prediction method of the present invention is proposed. The step S30 includes:
[0096] Step S301, sort the change values of the drug gene expression correlations of the first cell line according to a preset relationship.
[0097] It can be understood that the preset relationship can be a relationship from large to small or from small to large. After sorting, the change values of the drug gene expression correlations of the first cell line are arranged according to the preset relationship.
[0098] Step S302, respectively extract the top M gene pairs with the largest up-regulation and down-regulation of correlations from the sorting results.
[0099] It should be understood that after the sorting is completed, the top M gene pairs with the largest up-regulation and down-regulation in the sorting result are respectively extracted, and the value of M can be 300,000.
[0100] Step S303, construct gene features according to the top M gene pairs with the largest up-regulation and down-regulation in terms of correlation.
[0101] It can be understood that after obtaining the top M gene pairs with the largest up-regulation and down-regulation in terms of correlation, gene features are constructed according to the top M gene pairs.
[0102] Step S304, construct a reference instance according to the change value of the drug-gene expression correlation.
[0103] Further, step S304 includes: obtaining the change value of the gene expression correlation of each drug in the second cell line according to the change value of the drug-gene expression correlation; distinguishing the positive and negative sets of the change values of the gene expression correlation of each drug in the second cell line to obtain an up-regulated gene pair set and a down-regulated gene pair set; respectively obtaining the gene correlation change characteristics in the up-regulated gene pair set and the down-regulated gene pair set; respectively marking the gene pairs in the up-regulated gene pair set and the down-regulated gene pair set according to the gene correlation change characteristics to obtain the ranking marker numbers of each gene pair; obtaining the top N gene pairs with the smallest ranking in the up-regulated gene pair set and the top N gene pairs with the smallest ranking in the down-regulated gene pair set according to the ranking marker numbers of each gene pair; constructing a reference instance according to the top N gene pairs with the smallest ranking in the up-regulated gene pair set, the top N gene pairs with the smallest ranking in the down-regulated gene pair set, and the ranking marker numbers of each gene pair.
[0104] It can be understood that after distinguishing the up-regulated gene pair set and the down-regulated gene pair set, the ranking marker numbers of each gene pair are marked by using the gene correlation change characteristics in the up-regulated gene pair set and the down-regulated gene pair set. For example, there are N + up-regulated gene pairs in the up-regulated gene pair set and N - gene pairs in the down-regulated gene pair set. At this time, N + +N - =(N 2 -N) / 2. In the up-regulated gene pair set, the gene pair with the most significant change in correlation is marked as 1, that is, the ranking marker number of this gene pair is 1, and the gene pair with the least significant change in correlation is marked as N + , that is, the ranking marker number of this gene pair is N + . Then, the gene pairs in the down-regulated gene pair set are marked in the same way. After the marking is completed, a reference instance is constructed by using the top N gene pairs with the smallest ranking in the up-regulated gene pair set, the top N gene pairs with the smallest ranking in the down-regulated gene pair set, and the ranking marker numbers of each gene pair.
[0105] Specifically, with reference to Figure 4 , Figure 4 is a schematic diagram of the overall process. Specifically: in a specified cell background, gene expression profiles under various drug treatments and their control groups are obtained. Then, the correlation coefficient of each pair of genes is calculated based on the number of treatments for each cell line and the number of all genes in the gene expression profile, and the experimental group correlation coefficient vector V Drug and the control group correlation coefficient vector V Ctrl are determined. The drug-gene expression correlation change value V diff is calculated using the experimental group correlation coefficient vector and the control group correlation coefficient vector. Gene pairs within the gene feature are divided into the first gene pair set and the second gene pair set Gene pairs within each reference instance in the reference instance database are divided into the third gene pair set and the fourth gene pair set The number of gene pairs in the above four gene sets is the same. Then, the intersection of the first gene pair set and the third gene pair set is taken, and the intersection of the second gene pair set and the fourth gene pair set is taken, that is In addition, when exploring the treatment relationship between diseases and drugs, the specific intersection taken is: Then, the negative similarity scores between the gene feature and each reference instance in the reference instance database are calculated respectively, and the function of the drug to be detected is predicted based on the negative similarity scores.
[0106] In this embodiment, the drug-gene expression correlation change values of the first cell line are sorted according to a preset relationship; the top M gene pairs with the largest up-regulation and down-regulation of correlation are respectively extracted from the sorting results; a gene feature is constructed based on the top M gene pairs with the largest up-regulation and down-regulation of correlation; a reference instance is constructed based on the drug-gene expression correlation change value. In this way, after the drug-gene expression correlation change value of the first cell line, the correlation change value is sorted, a gene feature is constructed based on the extracted top 2M gene pairs, and a reference instance is constructed based on the drug-gene expression correlation change value, thereby effectively improving the accuracy of constructing the gene feature and the reference instance.
[0107] In addition, an embodiment of the present invention also proposes a storage medium, on which a drug function prediction program is stored. When the drug function prediction program is executed by a processor, the steps of the drug function prediction method described above are implemented.
[0108] Since this storage medium adopts all the technical solutions of the above all embodiments, it at least has all the beneficial effects brought by the technical solutions of the above embodiments, which will not be elaborated one by one here.
[0109] In addition, with reference to Figure 5, an embodiment of the present invention also provides a drug function prediction device, which includes:
[0110] A determination module 10, configured to determine an experimental group correlation coefficient vector and a control group correlation coefficient vector according to the gene expression profiles of each cell line.
[0111] A calculation module 20, configured to calculate a drug gene expression correlation change value according to the experimental group correlation coefficient vector and the control group correlation coefficient vector.
[0112] A construction module 30, configured to construct gene features and reference instances according to the drug gene expression correlation change value.
[0113] A prediction module 40, configured to calculate a positive similarity score between the gene features and the reference instances, and predict the function of the drug to be detected according to the positive similarity score.
[0114] In this embodiment, an experimental group correlation coefficient vector and a control group correlation coefficient vector are determined according to the gene expression profiles of each cell line; a drug gene expression correlation change value is calculated according to the experimental group correlation coefficient vector and the control group correlation coefficient vector; gene features and reference instances are constructed according to the drug gene expression correlation change value; a positive similarity score between the gene features and the reference instances is calculated, and the function of the drug to be detected is predicted according to the positive similarity score; in the above manner, from the perspective of the gene network, the correlation coefficient vector is determined by using the gene expression profiles of cell lines, and the drug gene expression correlation change value is calculated, and then gene features and reference instances are constructed to identify drugs between different cell lines, realizing cross-cell line prediction of drug effects, thereby effectively improving the accuracy of predicting drug functions.
[0115] It should be noted that the above-described work process is only illustrative and does not limit the protection scope of the present invention. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and no limitation is made here.
[0116] In addition, for technical details not described in detail in this embodiment, reference can be made to the drug function prediction method provided in any embodiment of the present invention, and details will not be repeated here.
[0117] In one embodiment, the determining module 10 is further configured to obtain a first cell line and a second cell line according to the respective cell lines; process the first cell line with a drug to be explored, and process the second cell line with a known drug; measure the cell gene expression level data in the processed first cell line and the cell gene expression level data in the second cell line respectively; determine the gene expression profiles of the respective cell lines according to the cell gene expression level data in the first cell line and the cell gene expression level data in the second cell line; and calculate an experimental group correlation coefficient vector and a control group correlation coefficient vector according to the gene expression profiles.
[0118] In one embodiment, the constructing module 30 is further configured to sort the drug-gene expression correlation change values of the first cell line according to a preset relationship; extract the top M gene pairs with the largest up-regulation and down-regulation of correlation from the sorting result respectively; construct gene features according to the top M gene pairs with the largest up-regulation and down-regulation of correlation; and construct a reference instance according to the drug-gene expression correlation change values.
[0119] In one embodiment, the constructing module 30 is further configured to obtain the drug-gene expression correlation change values of each drug in the second cell line according to the drug-gene expression correlation change values; distinguish positive and negative sets of the drug-gene expression correlation change values of each drug in the second cell line to obtain an up-regulated gene pair set and a down-regulated gene pair set; obtain the gene correlation change characteristics in the up-regulated gene pair set and the down-regulated gene pair set respectively; mark the gene pairs in the up-regulated gene pair set and the down-regulated gene pair set respectively according to the gene correlation change characteristics to obtain the ranking mark numbers of each gene pair; obtain the top N gene pairs with the smallest ranking in the up-regulated gene pair set and the top N gene pairs with the smallest ranking in the down-regulated gene pair set according to the ranking mark numbers of each gene pair; and construct a reference instance according to the top N gene pairs with the smallest ranking in the up-regulated gene pair set, the top N gene pairs with the smallest ranking in the down-regulated gene pair set, and the ranking mark numbers of each gene pair.
[0120] In one embodiment, the predicting module 40 is further configured to obtain the types of known drugs; obtain the reference instances of all known drugs according to the types of known drugs and the reference instance; generate a reference instance database according to the reference instances of all known drugs; and calculate the positive similarity scores between the gene features and each reference instance in the reference instance database respectively.
[0121] In one embodiment, the prediction module 40 is further configured to divide the gene pairs within the gene feature into a first gene pair set and a second gene pair set according to the positive or negative relationship of the correlation change value; divide the gene pairs within each reference instance in the reference instance database into a third gene pair set and a fourth gene pair set according to the positive or negative relationship of the correlation change value; take the intersection of the first gene pair set and the third gene pair set, and take the intersection of the second gene pair set and the fourth gene pair set; determine the ranking of each gene pair in the reference instance in the intersection; sort the gene pairs according to the ranking, and calculate the positive similarity scores between the gene feature and each reference instance in the reference instance database according to the ranked gene pairs.
[0122] In one embodiment, the prediction module 40 is further configured to sort the positive similarity scores; extract the top T positive similarity scores with higher rankings from the score ranking result; determine the drugs to be detected and the reference instance drugs corresponding to the top T positive similarity scores; obtain the functions of the reference instance drugs; and predict the functions of the drugs to be detected according to the functions of the reference instance drugs.
[0123] For other embodiments or implementation methods of the drug function prediction device of the present invention, reference may be made to the above method embodiments, which will not be elaborated here.
[0124] It should be understood that although the steps in the flowchart in the embodiments of the present application are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps is not strictly limited in order, and they can be executed in other orders. Moreover, at least a part of the steps in the figure may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0125] In addition, it should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including the element.
[0126] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as a read only memory (ROM) / RAM, magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, integrated platform workstation, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0128] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for predicting drug function, characterized in that: The drug function prediction method comprises the following steps: Determine the correlation coefficient vector of the experimental group and the correlation coefficient vector of the control group according to the gene expression profile of each cell line; Calculating the drug gene expression correlation change value according to the experimental group correlation coefficient vector and the control group correlation coefficient vector; constructing gene signatures and reference instances according to the drug gene expression correlation change values; Calculating a positive similarity score between the gene signature and the reference instance, and predicting the function of the drug to be detected according to the positive similarity score; The constructing of gene signatures and reference examples according to the drug gene expression correlation change values comprises: Sorting the drug gene expression correlation change values of the first cell line according to a preset relationship; Extract the top M gene pairs with the largest up-regulated and down-regulated correlations from the ranking results; Constructing gene signatures based on the top M gene pairs with the largest up- and down-regulation in the correlation; A reference instance is constructed according to the drug gene expression correlation change value.
2. The drug function prediction method according to claim 1, characterized in that: The step of determining the correlation coefficient vector of the experimental group and the correlation coefficient vector of the control group according to the gene expression profile of each cell line comprises: Obtaining a first cell line and a second cell line according to the cell lines; treating the first cell line with a drug to be explored, and treating the second cell line with a known drug; respectively measuring the gene expression level data of cells in the first cell line and the gene expression level data of cells in the second cell line after treatment; Determine the gene expression profile of each cell line based on the gene expression level data of cells in the first cell line and the gene expression level data of cells in the second cell line; The experimental group correlation coefficient vector and the control group correlation coefficient vector are calculated according to the gene expression profile.
3. The drug function prediction method according to claim 1, characterized in that: The constructing a reference instance according to the drug gene expression correlation change value comprises: Obtaining a gene expression correlation change value of each drug in the second cell line according to the drug gene expression correlation change value; Distinguishing the gene expression correlation change values of each drug in the second cell line by positive and negative sets to obtain an up-regulated gene pair set and a down-regulated gene pair set; Respectively obtaining gene correlation change characteristics in the up-regulated gene pair set and the down-regulated gene pair set; According to the gene correlation change characteristics, the gene pairs in the up-regulated gene pair set and the down-regulated gene pair set are marked respectively to obtain the ranking mark number of each gene pair; Obtaining the top N gene pairs with the smallest ranking in the up-regulated gene pair set and the top N gene pairs with the smallest ranking in the down-regulated gene pair set according to the ranking marker number of each gene pair; A reference instance is constructed according to the top N gene pairs with the smallest ranking in the up-regulated gene pair set, the top N gene pairs with the smallest ranking in the down-regulated gene pair set, and the number of ranking markers of each gene pair.
4. The method for predicting drug function according to claim 1, characterized in that: The calculating the positive similarity score between the gene feature and the reference instance comprises: Obtain the types of known drugs; Obtaining reference examples of all known drugs according to the types of the known drugs and the reference examples; Generate a reference example database based on the reference examples of all the known drugs; The positive similarity scores between the gene signature and each reference instance in the reference instance database are calculated respectively.
5. The drug function prediction method according to claim 4, characterized in that: The respectively calculating the positive similarity scores between the gene signature and each reference instance in the reference instance database comprises: Dividing the gene pairs in the gene signature into a first gene pair set and a second gene pair set according to the positive and negative relationship of the correlation change value; Dividing the gene pairs in each reference example in the reference example database into a third gene pair set and a fourth gene pair set according to the positive and negative relationship of the correlation change value; Taking the intersection of the first gene pair set and the third gene pair set, and taking the intersection of the second gene pair set and the fourth gene pair set; Determine the ranking of each gene pair in the intersection among the reference instances; The gene pairs are sorted according to the ranking, and the positive similarity scores between the gene features and the reference instances in the reference instance database are calculated according to the ranked gene pairs.
6. The method for predicting drug function according to any one of claims 1 to 5, characterized in that: The method of predicting the function of the drug to be detected according to the positive similarity score comprises: sorting the positive similarity scores; Extract the top T positive similarity scores from the score ranking results; Determining the drugs to be detected and the reference example drugs corresponding to the first T positive similarity scores; A function of obtaining the reference example drug; The function of the drug to be detected is predicted according to the function of the reference example drug.
7. A drug function prediction device, characterized in that: The drug function prediction device comprises: A determination module, used for determining the correlation coefficient vector of the experimental group and the correlation coefficient vector of the control group according to the gene expression profile of each cell line; A calculation module, used for calculating the drug gene expression correlation change value according to the experimental group correlation coefficient vector and the control group correlation coefficient vector; A construction module, used to construct gene features and reference instances according to the drug gene expression correlation change values; A prediction module, used for calculating a positive similarity score between the gene signature and the reference instance, and predicting the function of the drug to be detected according to the positive similarity score; The construction module is also used to sort the drug gene expression correlation change values of the first cell line according to a preset relationship; extract the top M gene pairs with the largest up-regulation and down-regulation of correlation from the sorting results; construct gene features based on the top M gene pairs with the largest up-regulation and down-regulation of correlation; and construct a reference instance based on the drug gene expression correlation change values.
8. A drug function prediction device, characterized in that: The drug function prediction device comprises: a memory, a processor, and a drug function prediction program stored in the memory and executable on the processor, wherein the drug function prediction program is configured to implement the drug function prediction method according to any one of claims 1 to 6.
9. A storage medium, characterized in that: The storage medium stores a drug function prediction program, which, when executed by a processor, implements the drug function prediction method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Drug efficacy prediction method based on graph neural network and omics information
CN114649097A
Pharmacodynamic prediction method, device and kit based on expression profile of few genes
CN115905898A