Compound perturbation cell line gene transcription profile prediction method, device, equipment and medium
By constructing a sample data set containing compound structural information and gene transcription profile data, and using a deep neural network model, especially a model combining Transformer decoder and a gene association linear transformation layer, the problem of poor generalization of compound perturbation gene expression profile prediction in the prior art is solved, and accurate prediction of unknown compounds and cell lines is achieved.
Patent Information
- Application Number
- CN202510495727.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing methods for predicting compound perturbation gene expression profiles are generally generalized in the prediction of unknown compounds and cell lines, and it is difficult to achieve accurate predictions in different cell lines and compound perturbation scenarios.
A sample data set containing compound structure information and gene transcriptional profile data was constructed. A deep neural network model was used, especially a model combining Transformer decoder and a gene-associated linear transformation layer, and a combined loss function of mean square error, Pearson correlation coefficient loss and KL divergence loss was trained to achieve the prediction of the gene transcriptional profile of compound perturbation cell lines.
The predictive generalization of unknown compounds and cell lines is improved, ensuring that the pattern of gene expression change is highly correlated with real biological changes, and accurate predictions are achieved under different cell lines and compound perturbation scenarios.
Smart Images

Figure CN120015119A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of compound-perturbed gene transcription profile prediction, and in particular to a compound-perturbed cell line gene transcription profile prediction method, device, equipment and medium. Background Art
[0002] Comparative transcriptomics has become an important technical means to identify drug targets and analyze drug action mechanisms. By comparing the gene transcription profiles of cell lines (tissues, clinical biological samples) before and after drug intervention, the changes in RNA (ribonucleic acid) can be characterized on a genome-wide scale. Through pathway enrichment and gene module analysis, genes and signal pathways strongly associated with drug intervention are enriched, which has become a common technology in drug research such as drug action mechanisms and potential drug discovery. However, there are millions of known compound libraries at present. For example, the PubChem database (an organic small molecule biological activity database) contains more than 110 million compounds, and there are more than 400 known cell types in the human body. Comprehensive experimental characterization of the intervention effects of known compounds on cell lines is time-consuming, labor-intensive, and will cost a lot of financial resources. In addition, the batch effect of transcriptome data itself caused by factors such as technical platforms and cell line variations also makes it difficult to obtain stable drug perturbation characterization. Therefore, it is not appropriate to rely on traditional experimental methods to analyze large-scale drug-cell line combinations. Therefore, methods for training deep learning models based on existing large-scale drug perturbation gene profiles and extending them to other drug-cell line combinations are gradually emerging. Among the existing deep learning-based drug perturbation gene expression profile prediction tools, the variational autoencoder-based method has become the main training framework. This type of model can effectively solve the noise reduction problem of the expression profile and make high-accuracy predictions to a certain extent. However, there are also some problems. For example, CPA (Compositional Perturbationautoencoder) transforms the VAE (Variational Autoencoder) model and integrates the expression profile and drug structure information, but it can only predict the molecules already in the training data and cannot be extended to other unknown compounds; the CIGER (Chemical-induced Gene Expression Ranking) model predicts the data of 8 cell lines, and its generalization on cell lines is general. DLEPS (Deep Learning–based Efficacy Prediction System) uses a custom CTP (Changes in Transcriptional Profiles) value as a standard. The input of drug molecule information can predict the changes in CTP, but it cannot predict the perturbations of specific cell lines and cell types. In summary, the existing methods have only moderate generalizability in predicting unknown compounds and cell lines. Summary of the invention
[0003] The purpose of the present application is to provide a method, device, equipment and medium for predicting the gene transcription profile of a compound-perturbed cell line, so as to realize the prediction of the gene transcription profile of a cell line perturbed by any compound and improve the generalization of the prediction of unknown compounds and cell lines.
[0004] To achieve the above objectives, this application provides the following solutions.
[0005] In a first aspect, the present application provides a method for predicting gene transcription profiles of cell lines perturbed by compounds, comprising the following steps.
[0006] Constructing a sample data set; the sample data set includes multiple sample data; the sample data includes: compound structure information and gene transcription profile data, and the gene transcription profile data includes: cell line gene expression profile information before compound disturbance and cell line gene expression profile information after compound disturbance.
[0007] Construct a combined loss function that includes mean square error, Pearson correlation coefficient loss, and KL divergence loss.
[0008] A deep neural network model is trained using the sample data set and the combined loss function to obtain a trained deep neural network model as a gene transcription profile prediction model; the deep neural network model includes a Transformer decoder and a gene association linear transformation layer connected in sequence, the Transformer decoder is used to fuse the compound structure information and the cell line gene expression profile information before the compound disturbance to obtain a gene expression profile change vector of the fused compound structure information, and the gene association linear transformation layer is used to perform gene association calculation on the gene expression profile change vector to obtain the cell line gene expression profile information after the compound disturbance.
[0009] Gene transcription profile prediction model is used to predict gene transcription profile.
[0010] In a second aspect, the present application provides a device for predicting gene transcription profiles of compound-perturbed cell lines, wherein the device for predicting gene transcription profiles of compound-perturbed cell lines applies the above-mentioned method for predicting gene transcription profiles of compound-perturbed cell lines, and the device for predicting gene transcription profiles of compound-perturbed cell lines includes the following modules.
[0011] The sample data set construction module is used to construct a sample data set; the sample data set includes multiple sample data; the sample data includes: compound structure information and gene transcription spectrum data, and the gene transcription spectrum data includes: cell line gene expression spectrum information before compound disturbance and cell line gene expression spectrum information after compound disturbance.
[0012] The loss function construction module is used to construct a combined loss function including mean square error, Pearson correlation coefficient loss and KL divergence loss.
[0013] A model training module is used to train a deep neural network model using the sample data set and the combined loss function to obtain a trained deep neural network model as a gene transcription profile prediction model; the deep neural network model includes a Transformer decoder and a gene association linear transformation layer connected in sequence, the Transformer decoder is used to fuse the compound structure information and the cell line gene expression profile information before the compound disturbance to obtain a gene expression profile change vector of the fused compound structure information, and the gene association linear transformation layer is used to perform gene association calculation on the gene expression profile change vector to obtain the cell line gene expression profile information after the compound disturbance.
[0014] The gene transcription profile prediction module is used to predict the gene transcription profile using the gene transcription profile prediction model.
[0015] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned method for predicting gene transcription profiles of compound-perturbed cell lines.
[0016] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for predicting gene transcription profiles of compound-perturbed cell lines.
[0017] According to the specific embodiments provided in this application, this application has the following technical effects.
[0018] The present application provides a method, device, equipment and medium for predicting gene transcription profiles of compound-perturbed cell lines, constructs a combined loss function including mean square error, Pearson correlation coefficient loss and KL divergence loss, and obtains a gene transcription profile prediction model by training a deep neural network model including a Transformer decoder and a gene association linear transformation layer connected in sequence. The model uses a Transformer decoder to fuse compound structure information and cell line gene expression profile information before compound perturbation to obtain a gene expression profile change vector that fuses the compound structure information, and uses a gene association linear transformation layer to perform gene association calculation on the gene expression profile change vector to obtain the cell line gene expression profile information after compound perturbation. The present application is based on an optimization strategy of mean square error, Pearson correlation loss and KL divergence loss, and performs superior performance in both numerical accuracy and trend consistency. The method can not only accurately predict the expression level after gene perturbation, but also ensure that the gene expression change pattern is highly correlated with the real biological change, thereby improving the generalization ability in different cell lines and compound perturbation scenarios.
[0019] The gene transcription spectrum prediction model is further used to construct gene transcription map data of virtual drug perturbations, and a computational method is designed to accurately recommend potentially effective drugs. This can accurately and efficiently complete the prediction of gene expression spectra of large-scale drug perturbations, greatly improving the efficiency and quality of drug screening. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0021] Figure 1 A schematic diagram of a process for predicting gene transcription profiles of cell lines perturbed by compounds provided in one embodiment of the present application.
[0022] Figure 2 A connection score bar chart provided in accordance with an embodiment of the present application.
[0023] Figure 3 A gene pathway enrichment score dot plot provided in one embodiment of the present application.
[0024] Figure 4 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0025] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0026] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0027] In an exemplary embodiment, Figure 1 As shown, a method for predicting gene transcription profiles of cell lines perturbed by compounds is provided, comprising the following steps 101 to 104.
[0028] Step 101, constructing a sample data set; the sample data set includes multiple sample data; the sample data includes: compound structure information and gene transcription profile data, and the gene transcription profile data includes: cell line gene expression profile information before compound disturbance and cell line gene expression profile information after compound disturbance.
[0029] Step 102: Construct a combined loss function including mean square error, Pearson correlation coefficient loss and KL divergence loss.
[0030] Step 103: train a deep neural network model using the sample data set and the combined loss function to obtain the trained deep neural network model as a gene transcription profile prediction model; the deep neural network model includes a Transformer decoder and a gene association linear transformation layer connected in sequence, the Transformer decoder is used to fuse the compound structure information and the gene expression profile information of the cell line before the compound disturbance to obtain a gene expression profile change vector fused with the compound structure information, and the gene association linear transformation layer is used to perform gene association calculation on the gene expression profile change vector to obtain the gene expression profile information of the cell line after the compound disturbance.
[0031] Step 104: Use the gene transcription profile prediction model to predict the gene transcription profile.
[0032] By implementing the above steps 101 to 104, the gene transcription profile of a cell line disturbed by any compound can be predicted.
[0033] In another exemplary embodiment of the present application, in the above step 101, gene transcription profile data after large-scale compound perturbation is collected to establish a high-quality characterized sample data set. Specifically, the LINC (Library of Integrated Network-based Cellular Signatures) 2020 drug data set (level 3 data set) is downloaded from the clue.io database. This data set covers more than 1.8 million gene transcription profile data after drug intervention and its control data, including gene expression profiles before and after drug perturbation, cell line macro information, compound macro information, gene symbol macro information, and compound simplified molecular input linear representation (Simplified Molecular Input Line Entry System, Smiles). Its data features are transcription profile data of 978 genes before and after drug perturbation, including description information of each compound and cell line. The obtained data were cleaned to remove data that could not be characterized by smiles and had low cell line experimental batches. The Modified Zero-mean Normalized Cross-Correlation (MODZ) algorithm was used to deduplicate the data of the same biological replicates. All data were saved in the advanced H5 file format and stored in the local database. The H5 file format is the fifth generation version of the hierarchical data format (HierarchicalData Format 5).
[0034] In another exemplary embodiment, in order to further improve the quality of the above sample data set and facilitate computer calculation, the following steps 201 to 202 are set after the above step 101.
[0035] Step 201, collect the compound structure information in the data set, establish the molecular fingerprint data representation of the compound, that is, use the classic Morgan fingerprint algorithm to perform molecular fingerprint data representation, and convert the molecular text information into a computer-calculated molecular fingerprint data representation. Morgan fingerprint is a compound data representation at the atomic level. The embodiment of the present application uses Morgan molecular fingerprints to perform data representation of compounds, which can effectively learn the fusion representation of compounds and genes. The embodiment of the present application uses a software package based on RDKit (Open-Source Cheminformatics, open source chemical informatics toolkit) to perform data representation of compounds.
[0036] Step 202: Based on the variational autoencoder basic model, data representation and noise reduction are performed on the highly variable gene expression spectrum data to extract the inherent data pattern of the cell line gene expression.
[0037] The variational autoencoder in the above step 202 includes an encoder and a decoder. The encoder compresses high-dimensional data to low dimensions, and the decoder reconstructs the inherent data pattern of cell line gene expression from the low-dimensional data representation to complete the deep denoising of the gene transcription spectrum data. The basic cell line gene expression spectrum information not disturbed by the compound is input into the variational autoencoder for training, and the data noise caused by the sequencing platform and the inherent variation of the cell line is removed to robustly represent the transcription state of the cell line. The obtained variational autoencoder is stored in a local computer cluster. The cell line gene expression spectrum information not disturbed by the compound in this application represents the basic cell line gene expression spectrum information not disturbed by any compound, and has the characteristics of low noise. The cell line gene expression spectrum information before the above-mentioned compound disturbance refers to the cell line gene expression spectrum information not disturbed by the compound to be studied, which is used to compare with the cell line gene expression spectrum information after the compound disturbance to determine the effect of the compound to be studied.
[0038] In another exemplary embodiment, in the process of training the above-mentioned deep neural network model, a combined loss function is used to fit the gene expression spectrum of the cell line after the final compound perturbation. The combined loss function is specifically divided into three parts: first, by calculating the mean square error between the predicted value and the true value, the model's fitting ability for the gene expression level is directly measured; secondly, the present application designs a Pearson correlation coefficient loss to calculate the Pearson correlation coefficient between the predicted value and the true value, and minimize its deviation from 1. This loss can measure the consistency of the changing trends of the two variables; finally, the KL (Kullback-Leibler) divergence loss is used to regularize the latent space to ensure that the latent representation of the model is reasonably distributed and avoid overfitting.
[0039] In another exemplary embodiment, the gene transcription profile prediction model established in the above steps can be applied to large-scale virtual drug screening. For tissue types and cell lines corresponding to specific diseases, the corresponding cell line gene expression profile information is obtained as the cell line gene expression profile information in the current state, and virtual perturbations are performed using the compound structure information of different drugs, and the connection score and other calculation methods are combined to recommend compounds with strong reversal effects. The calculation results are presented in a visual way on a web page or a chart file.
[0040] For example, taking potential drugs for stroke as the research object, the obtained gene transcription profile prediction model is applied to a large-scale natural product database. On the basis of obtaining the gene expression profile information of the cell line after perturbation, it is associated with the expression profile of disease tissues and disease cell lines to calculate the connection score and gene pathway enrichment score. The connection score results are visualized using a bar chart, such as Figure 2 As shown, pathway analysis uses dot plots to visualize results, such as Figure 3The analysis result report is presented in the form of a web page or a file.
[0041] Based on the same inventive concept, the present application embodiment also provides a compound perturbed cell line gene transcription profile prediction device for implementing the compound perturbed cell line gene transcription profile prediction method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more compound perturbed cell line gene transcription profile prediction device embodiments provided below can refer to the above limitations on the compound perturbed cell line gene transcription profile prediction method, which will not be repeated here.
[0042] In an exemplary embodiment, a device for predicting gene transcription profiles of compound-perturbed cell lines is provided, comprising the following modules.
[0043] The sample data set construction module is used to construct a sample data set; the sample data set includes multiple sample data; the sample data includes: compound structure information and gene transcription spectrum data, and the gene transcription spectrum data includes: cell line gene expression spectrum information before compound disturbance and cell line gene expression spectrum information after compound disturbance.
[0044] The loss function construction module is used to construct a combined loss function including mean square error, Pearson correlation coefficient loss and KL divergence loss.
[0045] A model training module is used to train a deep neural network model using the sample data set and the combined loss function to obtain a trained deep neural network model as a gene transcription profile prediction model; the deep neural network model includes a Transformer decoder and a gene association linear transformation layer connected in sequence, the Transformer decoder is used to fuse the compound structure information and the cell line gene expression profile information before the compound disturbance to obtain a gene expression profile change vector of the fused compound structure information, and the gene association linear transformation layer is used to perform gene association calculation on the gene expression profile change vector to obtain the cell line gene expression profile information after the compound disturbance.
[0046] The gene transcription profile prediction module is used to predict the gene transcription profile using the gene transcription profile prediction model.
[0047] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 4As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for predicting the gene transcription profile of a compound-perturbed cell line is implemented.
[0048] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0049] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0050] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0051] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0052] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0053] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.
[0054] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0055] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for predicting gene transcription profiles of compound-perturbed cell lines, characterized in that: include: Construct a sample dataset; The sample data set includes a plurality of sample data; The sample data includes: compound structure information and gene transcription profile data, and the gene transcription profile data includes: gene expression profile information of the cell line before the compound disturbance and gene expression profile information of the cell line after the compound disturbance; Construct a combined loss function including mean square error, Pearson correlation coefficient loss and KL divergence loss; A deep neural network model is trained using the sample data set and the combined loss function to obtain a trained deep neural network model as a gene transcription profile prediction model; the deep neural network model includes a Transformer decoder and a gene association linear transformation layer connected in sequence, the Transformer decoder is used to fuse the compound structure information and the cell line gene expression profile information before the compound disturbance to obtain a gene expression profile change vector fused with the compound structure information, and the gene association linear transformation layer is used to perform gene association calculation on the gene expression profile change vector to obtain the cell line gene expression profile information after the compound disturbance; Gene transcription profile prediction model is used to predict gene transcription profile.
2. The method for predicting gene transcription profiles of compound-perturbed cell lines according to claim 1, characterized in that: Construct a sample dataset, and then include: A variational autoencoder is used to denoise the gene transcription profile data in the sample dataset.
3. The method for predicting gene transcription profiles of compound-perturbed cell lines according to claim 2, characterized in that: The variational autoencoder is trained based on gene expression profile information of a cell line that is not perturbed by a compound.
4. The method for predicting gene transcription profiles of compound-perturbed cell lines according to claim 2, characterized in that: The variational autoencoder includes an encoder and a decoder connected in sequence; The encoder is used to compress the gene transcription spectrum data to obtain compressed gene transcription spectrum data; The decoder is used to reconstruct the compressed gene transcription profile data according to the inherent data pattern of the cell line gene expression.
5. The method for predicting gene transcription profiles of compound-perturbed cell lines according to claim 1, characterized in that: Construct a sample dataset, and then include: The Morgan fingerprint algorithm is used to characterize the molecular fingerprint data of the compound structure information in the sample data set to obtain the molecular fingerprint data representation of the compound structure information.
6. The method for predicting gene transcription profiles of compound-perturbed cell lines according to claim 1, characterized in that: A cross attention mechanism is provided in the Transformer decoder; The cross-attention mechanism is used to perform matrix calculation on the query matrix and the bond matrix to obtain the attention weights between each local structure of the compound and each gene, and multiply the attention weights between each local structure of the compound and each gene with the value matrix to obtain the gene expression spectrum change vector that integrates the compound structure information; the query matrix is constructed based on the gene expression spectrum information of the cell line before the compound disturbance, and the bond matrix and the value matrix are both constructed based on the compound structure information.
7. The method for predicting gene transcription profiles of compound-perturbed cell lines according to claim 1, characterized in that: The deep neural network model is trained using the sample data set and the combined loss function to obtain the trained deep neural network model as a gene transcription profile prediction model, and then further comprising: Obtain compound structure information of different drugs; According to the compound structure information of different drugs and the gene expression profile information of the cell line in the current state, the gene transcription profile prediction model is used to obtain the gene expression profile information of the cell line after being disturbed by different drugs; Calculate the connection score and gene pathway enrichment score of the gene expression profile information of the cell lines after different drug perturbations; Drug selection is performed based on the connection score and the gene pathway enrichment score.
8. A device for predicting gene transcription profiles of compound-perturbed cell lines, characterized in that: The compound-perturbed cell line gene transcription profile prediction device applies the compound-perturbed cell line gene transcription profile prediction method according to any one of claims 1 to 7, and the compound-perturbed cell line gene transcription profile prediction device comprises: A sample data set construction module is used to construct a sample data set; the sample data set includes a plurality of sample data; the sample data includes: compound structure information and gene transcription profile data, and the gene transcription profile data includes: gene expression profile information of a cell line before compound disturbance and gene expression profile information of a cell line after compound disturbance; The loss function construction module is used to construct a combined loss function including mean square error, Pearson correlation coefficient loss and KL divergence loss; A model training module, used to train a deep neural network model using the sample data set and the combined loss function, and obtain the trained deep neural network model as a gene transcription profile prediction model; the deep neural network model includes a Transformer decoder and a gene association linear transformation layer connected in sequence, the Transformer decoder is used to fuse the compound structure information and the cell line gene expression profile information before the compound disturbance, and obtain the gene expression profile change vector of the fused compound structure information, and the gene association linear transformation layer is used to perform gene association calculation on the gene expression profile change vector to obtain the cell line gene expression profile information after the compound disturbance; The gene transcription profile prediction module is used to predict the gene transcription profile using the gene transcription profile prediction model.
9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for predicting gene transcription profiles of compound-perturbed cell lines as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for predicting gene transcription profiles of compound-perturbed cell lines described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Cancer driver gene prediction device based on graph attention network and multi-omics fusion
CN115171779A
Medicinal plant transcriptional regulation map prediction method
CN115223657A
Single cell transcription reaction prediction algorithm based on cross-domain feature cross migration
CN119252336A
Traditional Chinese medicine multi-target interaction prediction method based on Transform architecture
CN119479785A
Method for predicting transcriptional response to novel drug perturbation and virtual screening method and system
CN119763720A
Cited By
Evaluation method, device and system of tumor intervention small molecule effect and storage medium
CN121306599A