Gene expression dynamic prediction model establishment method and device, equipment and medium

By constructing a dynamic gene expression prediction model and utilizing deep learning and graph neural networks, the problem of capturing the dynamic relationship between drug and gene expression was solved, enabling efficient target discovery and new drug screening in drug development.

CN121838853APending Publication Date: 2026-04-10CHINA PHARM UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Current technologies lack effective models to capture the dynamic relationship between drug and gene expression, resulting in low efficiency in new drug discovery and target screening.

Method used

By constructing a dynamic gene expression prediction model, the nonlinear mapping relationship between drug-dosage-time-cell-gene expression is learned using multiple training samples. Combined with deep learning and graph neural networks, the dynamic changes in gene expression are predicted.

Benefits of technology

It enables dynamic prediction of drug-induced gene expression in cells while reducing costs, providing a quantitative reference for new drug discovery and target screening, shortening the R&D cycle and improving efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838853A_ABST
    Figure CN121838853A_ABST
Patent Text Reader

Abstract

The invention discloses a cell gene expression dynamic prediction model establishment method and device, equipment and a medium. The method mainly comprises the following steps: determining a plurality of training samples; training a model based on the plurality of training samples, and determining a target prediction gene expression dynamic prediction model; performing dynamic prediction on gene expressions corresponding to different drugs under different doses based on the gene expression dynamic prediction model; according to the method, the effect of dynamically predicting the gene expression of the drug acting on the specific cells on the basis of reducing the cost is achieved, and a reference is provided for drug research and development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical technology, specifically to a method, apparatus, equipment, and medium for proposing a dynamic gene expression prediction model. Background Technology

[0002] In the drug discovery phase, predicting the dynamic changes in cellular gene expression is crucial for drug development and target discovery. L1000 technology, by detecting the expression of select "anchor genes," infers the whole-genome expression profile, significantly improving the efficiency of acquiring large-scale gene expression data. This enables drug screening, disease mechanism research, and dynamic detection on cell turntables. However, despite the availability of abundant gene expression profiling data, models are still lacking to capture the complex relationships within these data for dynamic gene expression prediction, thus failing to provide a reference for new drug discovery and target screening.

[0003] Background of model building

[0004] In the fields of computational systems biology and pharmaceutical informatics, capturing the drug-gene expression relationship has always been a cutting-edge challenge and a hot research topic. The introduction of artificial intelligence-driven methods has quietly brought about significant improvements in efficiency and effectiveness in the drug development process. Unlike traditional statistical analysis or retrieval association methods, modern machine learning, especially deep learning models, is becoming an important tool for exploring the ternary relationship of "drug-cell-gene expression".

[0005] Gene expression plays a crucial role in understanding cellular responses to drug interventions. However, gene expression itself is noisy. This noisy gene expression poses a significant challenge to modeling gene-drug interactions. The L1000 assay is a low-cost, high-throughput gene expression profiling technique that measures a large number of drug-interfered transcriptome profiles. While these datasets provide chemical perturbations covering different cell types, dynamic models are still lacking to capture how drugs intervene in cellular processes over time. To capture this complex relationship, this study proposes an AI-driven dynamic prediction model for gene expression that learns from high-dimensional data, facilitating a fuller understanding of the complex relationship between drugs and changes in gene expression.

[0006] Significance of Model Research

[0007] Promoting the development and standardization of AI drug discovery methods

[0008] This invention is an application example of AI-driven gene modeling for drug intervention, representing a new paradigm for training and evaluating the relationship between drugs and dynamic gene expression. As an application example of drug intervention-gene modeling within AI for Drug discovery, this task can promote data standardization and, based on this, build AI models for prediction, thus facilitating the formation of standardized problem definitions and evaluation systems.

[0009] Capturing the drug-gene interaction

[0010] This invention introduces high-dimensional omics features and performs predictions, which can provide a reference for drug repositioning and the exploration of drug mechanisms. Furthermore, the model can predict gene expression in cells treated with different drugs and has the potential to be extended to personalized prediction.

[0011] Assisting in target discovery

[0012] Gene expression prediction has significant applications in target discovery. Since gene expression reflects cellular responses to external interventions such as drugs, predicting drug-induced changes in gene expression can indirectly reveal key regulatory pathways and potential targets. By using artificial intelligence models to model high-dimensional expression data, potential drug-gene interactions can be inferred, thus providing effective support for target discovery. Summary of the Invention

[0013] The purpose of this invention is to provide a method, apparatus, device, and medium for establishing a dynamic gene expression prediction model, so as to achieve, while reducing costs, fitting the predicted gene expression with the actual gene expression, dynamically predicting the gene expression of drug-treated cells, and thus providing a reference for new drug discovery and target screening, thereby solving the problems mentioned in the background art.

[0014] To achieve the above objectives, the present invention provides the following technical solution:

[0015] This invention provides a method for establishing a dynamic gene expression prediction model, the method comprising:

[0016] Identify multiple training samples:

[0017] The study aims to obtain cellular property characteristics, cell gene expression levels at different drug doses, cell gene expression levels at different time points, and cell gene expression levels under different drugs; the property characteristics include morphological characteristics, cell origin, tissue and organ affiliation, and disease type.

[0018] For each cell type, the data corresponding to the drug is processed according to a preset data processing method to obtain the gene expression level of cells under different drug doses and time points; wherein, the preset processing method includes data cleaning, data transformation, and data normalization;

[0019] Based on the data grouping at different times for each cell type, the data were classified and organized for several doses of different drugs to construct multiple training samples.

[0020] The training samples include, but are not limited to, high-dimensional features of drug molecules, morphological features of cells, cell origin, tissue and organ affiliation, disease type attributes, gene expression levels of cells treated with different drugs, gene expression levels of cells treated with drugs at different times, and gene expression levels of cells treated with drugs at different doses.

[0021] The gene expression dynamic prediction model is determined by training based on the training samples.

[0022] Based on the aforementioned gene expression dynamic prediction model, the gene expression levels of different drugs at different drug doses are dynamically predicted.

[0023] The method of the present invention determines the gene expression of the target drug intervention in at least one cell according to a preset design standard; and infers the dynamic changes of gene expression based on specific cells and drug doses corresponding to specific drugs.

[0024] The present invention also provides a device for establishing a dynamic gene expression prediction model, the device comprising:

[0025] The data input module is used to input multiple training samples; wherein, the training samples include, but are not limited to, attribute features such as high-dimensional features of drug molecules, morphological features of cells, cell origin, tissue and organ affiliation, disease type, gene expression levels corresponding to cells treated with different drugs, gene expression levels corresponding to cells treated with drugs at different time points, and gene expression levels corresponding to cells treated with drugs at different doses.

[0026] The training module trains the initial neural network model based on multiple training samples to obtain a dynamic prediction model of candidate gene expression containing learnable parameters.

[0027] The model determination module is used to determine key hyperparameters (based on the optimal model results) from the candidate gene expression dynamic prediction model to obtain the final gene expression dynamic prediction model. These key hyperparameters include layers, learning rate, batch size, epoch, dropout, and loss scale.

[0028] Furthermore, the data input module also includes a data acquisition unit, a data processing unit, and a training sample determination unit;

[0029] The data acquisition unit is used to acquire the cell's attribute characteristics, cell gene expression levels at different drug doses, cell gene expression levels at different times, and cell gene expression levels under different drugs;

[0030] The data processing unit is used to receive the cell, drug, and gene expression attribute characteristics acquired by the data acquisition unit, and process the data corresponding to the drug according to a preset data processing method to obtain the gene expression level of cells under different drugs at different doses and at different times; wherein, the preset processing method includes data cleaning, data transformation, and data normalization.

[0031] The training sample determination unit is used to construct training samples corresponding to different drug states at different times based on the gene expression level of the drug in its "initial state", the gene expression levels of different drugs and at different doses in specific cells, and the gene expression levels of specific cells, specific drugs, and specific doses at different times.

[0032] Furthermore, the training processing module includes: a forward propagation submodule, a loss calculation and backpropagation submodule, and a parameter update submodule;

[0033] The forward propagation submodule is responsible for inputting training samples into the neural network, calculating the activation value of each layer (forward propagation) and the final output prediction result through the network's layer structure.

[0034] The loss calculation and backpropagation submodule calls a preset loss function to evaluate the difference between predicted gene expression and actual gene expression. After obtaining the loss value, the module further executes the backpropagation process. The module automatically records the intermediate activation values ​​and weight information of each layer.

[0035] The parameter update submodule adjusts the parameters obtained during the backpropagation phase to gradually optimize the performance of the neural network.

[0036] Furthermore, the device also includes a gene expression dynamic prediction module, which is used to dynamically predict gene expression levels at other times based on attributes such as the high-dimensional characteristics of drug molecules, morphological characteristics of cells, cell origin, tissue and organ affiliation, disease type, and gene expression levels in the initial state of cells.

[0037] The present invention also provides an electronic device, the electronic device comprising:

[0038] At least one processor; and a memory communicatively connected to said at least one processor;

[0039] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the gene expression dynamic prediction model establishment method of the present invention.

[0040] The present invention also provides a computer-readable storage medium storing computer instructions for causing a processor to execute the gene expression dynamic prediction model establishment method described in the present invention.

[0041] Compared with the prior art, the beneficial effects of the present invention are:

[0042] This invention identifies multiple training samples with different cell types, drugs, drug dosages, and treatment times. Based on these training samples, a dynamic gene expression prediction model is trained, enabling the model to learn the nonlinear mapping relationship between "drug-dosage-time-cell-gene expression." Furthermore, a converged dynamic gene expression prediction model is determined, predicting gene expression levels at different drug dosages based on target cell characteristics, candidate drugs, their dosages, and intervention times. Through this approach, this invention not only captures the potential correlation between "drug-cell-gene" but also characterizes the dose-dependent and time-dependent dynamic changes during drug action. This provides a quantitative and scalable reference for candidate compound screening, dosing regimen optimization, and potential target discovery in drug development, thereby reducing experimental costs, shortening the development cycle, and improving development efficiency. Attached Figure Description

[0043] Figure 1 This is a flowchart of a method for establishing a dynamic prediction model of gene expression level according to Embodiment 1 of the present invention;

[0044] Figure 2 This is a schematic diagram of a device for establishing a dynamic prediction model of gene expression level according to Embodiment 2 of the present invention;

[0045] Figure 3 A schematic diagram of the structure of an electronic device for implementing the gene expression level dynamic prediction model establishment method of the present invention. Detailed Implementation

[0046] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0047] Example 1

[0048] Figure 1 This is a flowchart of a method for establishing a dynamic gene expression prediction model according to Embodiment 1 of the present invention. This embodiment is applicable to the dynamic prediction of gene expression in cells treated with different drugs. This method can be executed by a gene expression dynamic prediction model establishment device, which can be implemented in hardware and / or software. This bio-utilization prediction model establishment device can be configured in a terminal and / or server. Figure 1 As shown, the method includes:

[0049] S110. Determine multiple training samples.

[0050] The training samples include attributes such as high-dimensional features of drug molecules, morphological features of cells, cell origin, tissue and organ affiliation, and disease type; gene expression levels corresponding to cells treated with different drugs; gene expression levels corresponding to cells treated with drugs at different time points; and gene expression levels measured by cells treated with drugs at different doses.

[0051] Among the cell samples studied, most were cell lines derived from tumors or cancers, possessing high representativeness and research value. For example, the A549 and MCF7 cell lines are widely used in oncology research. The A549 cell line originates from human lung adenocarcinoma tissue and is a type of human epithelial cell belonging to the lung, specifically lung cancer. A549 cells not only play an important role in research on the biological mechanisms of lung cancer but are also frequently used in drug screening, cytotoxicity analysis, viral infection mechanisms, and other research fields. The MCF7 cell line is an epithelial cell line derived from human breast adenocarcinoma, also of human origin. It was initially isolated from breast cancer patients and is widely used in basic and translational medicine research related to breast cancer.

[0052] The high-dimensional structural features of drugs are digitally encoded using molecular fingerprints for subsequent processing and analysis in models. A classic and widely used type of molecular fingerprint is MACCS (Molecular Access System) Keys. MACCS is a Boolean encoding method based on structural rules that converts a chemical molecule into a fixed-length binary vector. Each bit in this fingerprint vector corresponds to a predefined chemical substructure or functional group feature, indicating whether the molecule contains a specific structural segment. In this way, multiple potential structural information of a chemical molecule are systematically mapped into a chemically meaningful vector representation that is easily processed in a computer.

[0053] The study employed a time-series design, using the unperturbed cell state as a baseline reference. Specifically, cell samples treated with DMSO (Dimethyl Sulfoxide) were selected as the unperturbed cell state. Furthermore, cell gene expression profiles at different time points (e.g., 6 hours and 24 hours after treatment) were obtained to constitute the training samples for the time series. In terms of data dimensions, expression data for at least 978 landmark genes were available at each time point. It should be noted that before constructing the training samples, the acquired parameters were processed to better utilize the data for subsequent analysis and model processing, and the training samples were constructed based on the processed parameters.

[0054] The preset data processing methods can include at least one of data cleaning, data encoding, data transformation, and data normalization. Data encoding can be understood as encoding the molecular structures contained in the data into vectors. Data transformation can be understood as transforming data into feature parameters that can be better recognized by a computer based on a preset data processing model. Data normalization can be understood as mapping data to a specified range, used to remove the dimensions and units of different dimensions of data.

[0055] In this embodiment, the parameters of drug structure and drug dosage can include various types. For example, suitable drug parameters and drug dosage parameters can be selected for a specific cell type. These drug parameters can be used to predict and analyze the gene expression level of cells under certain conditions, thereby revealing the transcriptional response patterns under drug action.

[0056] S120. Train the model based on multiple training samples and determine the dynamic prediction model of target gene expression.

[0057] The dynamic gene expression prediction model can be a neural network model capable of predicting gene expression, with model parameters set to initial or default values. The target gene expression dynamic prediction model is a trained neural network model that takes the molecular formula of the drug to be predicted, the target cell type, the drug dosage, and the gene expression level of the corresponding cells after drug treatment as input. The model dynamically predicts the gene expression level of the drug in the treated cells based on the input data.

[0058] In this embodiment, after obtaining multiple training samples, the gene expression dynamic prediction model to be trained can be trained based on these training samples. This results in a trained target gene expression dynamic prediction model.

[0059] Gene expression levels not only depend on the experimental data itself but also integrate prior knowledge of the biological background. For example, a pre-constructed gene regulatory network structure serves as a structural prior for modeling gene expression levels. This prior knowledge helps us capture potential dependencies between genes and upstream and downstream regulatory pathways, thereby improving the biological rationality and predictive ability of the model. Integrating prior knowledge can be achieved using graph neural networks (GNNs). GNNs can transmit information between nodes (genes), updating the expression representation of each gene node using the features of neighboring nodes. In this way, the model not only learns the individual expression characteristics of each gene but also captures its interaction patterns and regulatory dependencies within complex biological networks.

[0060] The predicted gene expression can be based on gene expression predictions obtained from a dynamic gene expression prediction model. A preset optimizer is used to update the training parameters in the neural network. Optionally, the preset optimizer may include Adam (Adaptive Moment Estimation), an adaptive learning rate optimization algorithm based on first- and second-order moment estimation. Adam dynamically adjusts the learning rate of each parameter in each iteration, thereby improving training stability and convergence speed. Preset evaluation metrics are used to evaluate the predictive performance of the dynamic gene expression prediction model. Optionally, preset evaluation metrics may include mean squared error (MSE), which measures the difference between the predicted results of the dynamic gene expression prediction model and the true values. Generally, the lower the MSE, the higher the accuracy of the prediction model.

[0061] S130. Based on the dynamic prediction model of gene expression, dynamically predict the gene expression of different drugs at different doses.

[0062] Among them, the gene expression dynamic prediction model can dynamically predict the gene expression level at other times based on the high-dimensional characteristics of drug molecules, the morphological characteristics of cells, cell origin, tissue and organ affiliation, disease type and other attributes, as well as the gene expression level of the initial state of the cell.

[0063] It should be noted that the gene expression dynamic prediction model can systematically predict gene expression levels at 6 h and 24 h, and can infer the dynamic changes in cellular gene expression under drug intervention. The gene expression dynamic prediction model can combine time series information to capture the mapping relationship between drugs and cellular genes, simulate the impact of drugs on gene regulation at different times, and thus infer the expression trend and level of genes at different time points.

[0064] The technical solution of this invention determines multiple training samples, then determines a dynamic gene expression prediction model based on the multiple training samples, and further, dynamically predicts the gene expression level of a certain cell corresponding to different drug doses based on the dynamic gene expression prediction model. This solves the problem that gene expression determination methods in related technologies are time-consuming and expensive, and achieves the effect of dynamically predicting the gene expression level of a drug acting on a specific cell while reducing costs, thereby providing a reference for drug development.

[0065] Example 2

[0066] Figure 2 is a schematic diagram of a gene expression dynamic prediction model establishment device provided in Embodiment 2 of the present invention. Figure 2 As shown, the device includes: a data input module 210, a training processing module 220, and a model determination module 230.

[0067] The data input module 210 is used to input multiple training samples. These training samples include attributes such as high-dimensional features of drug molecules, morphological features of cells, cell origin, tissue and organ affiliation, and disease type; gene expression levels corresponding to cells treated with different drugs; gene expression levels corresponding to cells treated with drugs at different time points; and gene expression levels corresponding to cells treated with drugs at different doses. The samples are divided into training and testing sets. The training processing module 220 constructs a neural network model and initializes weight parameters. During training, data is input into the network and predicted via forward propagation. The deviation between the prediction and the actual result is then measured using a loss function. This information is passed to the weights of each layer via backpropagation, and parameters are updated using Adam. The model determination module 230 determines key hyperparameters based on the initial model structure. Specifically, it can systematically adjust parameters affecting model performance, including but not limited to: the number of network layers, the selection of activation function types, and the Dropout ratio.

[0068] The technical solution of this invention determines multiple training samples, then determines a dynamic gene expression prediction model based on the multiple training samples, and further, dynamically predicts the gene expression level of a certain cell corresponding to different drug doses based on the dynamic gene expression prediction model. This solves the problem that gene expression determination methods in related technologies are time-consuming and expensive, and achieves the effect of dynamically predicting the gene expression level of a drug acting on a specific cell while reducing costs, thereby improving the success rate of drug development.

[0069] Optionally, the training sample determination module 210 includes: a parameter acquisition unit, a data processing unit, and a training sample determination unit.

[0070] Optionally, the data input module 210 includes a data acquisition unit, a training sample determination unit, and a data processing unit.

[0071] The data acquisition unit is used to acquire the gene expression levels of various cells in their "initial state", to acquire the gene expression levels of different drugs and different doses in specific cells, and to acquire the gene expression levels of specific cells, specific drugs, and specific doses at different times.

[0072] The training sample determination unit is used to construct training samples corresponding to different drug states at different times based on the gene expression level of the drug in its "initial state", the gene expression levels of different drugs and at different doses in specific cells, and the gene expression levels of specific cells, specific drugs, and specific doses at different times.

[0073] The data processing unit integrates training samples under different states to obtain training samples for model input; wherein the preset data processing method includes at least one of data cleaning, data encoding, data transformation and data normalization.

[0074] Optionally, the training processing module 220 includes: a forward propagation submodule, a loss calculation and backpropagation submodule, and a parameter update submodule.

[0075] The forward propagation submodule is responsible for inputting training samples through the network's layers, calculating the activation values ​​of each layer, and outputting the final prediction result. This module boasts excellent scalability and compatibility, supporting dynamic computation graph construction, batch parallel processing, and acceleration hardware (such as GPUs).

[0076] Loss Calculation and Backpropagation Submodule: After completing the prediction output in the forward propagation stage, this submodule first calls a preset loss function to evaluate the difference between predicted gene expression and actual gene expression. Mean Squared Error (MSE) is used to calculate the error. After obtaining the loss value, the module further executes the backpropagation process, automatically recording the intermediate activation values ​​and weight information of each layer.

[0077] Parameter Update Submodule: This submodule is primarily responsible for adjusting the parameters obtained during the backpropagation phase to gradually optimize the performance of the neural network. This module mainly relies on the Adam optimization algorithm to update the weights and biases of each layer in the network, enabling the model to converge continuously during training.

[0078] Optionally, the model determination module 230 determines key hyperparameters based on the initial model structure. Specifically, it can systematically adjust parameters affecting model performance, including but not limited to: flexibly setting the number of network layers and the number of neurons in each layer according to task complexity and sample size to regulate model capacity and expressive power; selecting appropriate activation function types based on nonlinear modeling requirements and gradient propagation stability to balance training efficiency and fitting accuracy; and setting the Dropout ratio to suppress overfitting risk and enhance the model's generalization ability.

[0079] Optionally, the device further includes a gene expression dynamic prediction module.

[0080] The gene expression dynamic prediction module is used to dynamically predict gene expression levels at other times based on attributes such as the high-dimensional characteristics of drug molecules, morphological characteristics of cells, cell origin, tissue and organ affiliation, disease type, and gene expression levels in the initial state of cells.

[0081] The gene expression dynamic prediction model establishment device provided in the embodiments of the present invention can execute the gene expression dynamic prediction model establishment method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0082] Example 3

[0083] Figure 3A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0084] like Figure 3 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0085] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0086] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the gene expression dynamic prediction model building method.

[0087] In some embodiments, the gene expression dynamic prediction model building method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the gene expression dynamic prediction model building method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to execute the gene expression dynamic prediction model building method by any other suitable means (e.g., by means of firmware).

[0088] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0089] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0090] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0091] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0092] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0093] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0094] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0095] It should be noted that the above content merely illustrates the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. For those skilled in the art, various improvements and modifications can be made without departing from the principle of the present invention, and all such improvements and modifications fall within the scope of protection of the claims of the present invention.

Claims

1. A method for establishing a dynamic prediction model of cell gene expression, characterized in that, include: Multiple training samples were obtained: the training samples included: cell attribute characteristics, high-dimensional structural fingerprints of drug molecules, and gene expression levels of cells under different drugs, different doses, and different times; the attribute characteristics included morphological characteristics, cell origin, tissue and organ affiliation, and disease type; For each training sample, the cell attribute features, the high-dimensional structural fingerprint features of the drug molecule, and the cell gene expression levels under different drug treatments, different doses, and different time points are encoded and input into the neural network model. The neural network model is trained according to the corresponding encoding to obtain the dynamic changes of the predicted gene expression levels corresponding to the training sample. The predicted gene expression level corresponding to the training sample and the actual gene expression level in the training sample are processed according to the preset evaluation index to obtain the evaluation result corresponding to the preset evaluation index, and the gene expression dynamic prediction model is determined according to the evaluation result. Based on the aforementioned gene expression dynamic prediction model, the gene expression of the test cells under different drugs, different doses, and different times is dynamically predicted.

2. The method for establishing a dynamic prediction model of cell gene expression according to claim 1, characterized in that, Obtaining multiple training samples includes the following steps: To obtain cell properties, cell gene expression levels at different drug doses, cell gene expression levels at different time points, and cell gene expression levels under different drugs; The attribute characteristics include morphological characteristics, cell origin, tissue and organ affiliation, and disease type; For each cell type, the data corresponding to the drug is processed according to a preset data processing method to obtain the gene expression levels of cells under different drug doses and time periods; wherein, the preset processing method includes data cleaning, data transformation, and data normalization; Based on the data grouping at different times for each cell type, the data were classified and organized for several doses of different drugs to construct multiple training samples.

3. The method for establishing a dynamic prediction model of cell gene expression according to claim 1, characterized in that, The preset evaluation index is used to evaluate the model prediction performance of the gene expression dynamic prediction model and to determine the gene expression dynamic model. The preset evaluation index is mean squared error.

4. A device for establishing a dynamic gene expression prediction model, characterized in that, include The data input module is used to input multiple training samples; wherein, the training samples include high-dimensional features of drug molecules, morphological features of cells, cell origin, tissue and organ affiliation, disease type, gene expression levels corresponding to cells treated with different drugs, gene expression levels corresponding to cells treated with drugs at different time points, and gene expression levels corresponding to cells treated with drugs at different doses. The training processing module is used to train the initial neural network model based on multiple training samples to obtain a candidate gene expression dynamic prediction model containing learnable parameters. The model determination module is used to determine key hyperparameters and optimize the model structure based on the candidate gene expression dynamic prediction model, thereby obtaining the final gene expression dynamic prediction model.

5. The apparatus for a gene expression dynamic prediction model according to claim 4, characterized in that, The data input module further includes a data acquisition unit, a data processing unit, and a training sample determination unit; The data acquisition unit is used to acquire the cell's attribute characteristics, cell gene expression levels at different drug doses, cell gene expression levels at different times, and cell gene expression levels under different drugs; The data processing unit processes the data corresponding to the drug according to a preset data processing method to obtain the gene expression level of cells under different doses and times for different drugs; wherein, the preset processing method includes data cleaning, data transformation, and data normalization. The training sample determination unit is used to construct training samples corresponding to different time states for different drugs and different dosages.

6. The apparatus for a gene expression dynamic prediction model according to claim 4, characterized in that, The device also includes a gene expression dynamic prediction module, which is used to dynamically predict gene expression levels at other times based on attributes such as the high-dimensional characteristics of drug molecules, morphological characteristics of cells, cell origin, tissue and organ affiliation, disease type, and gene expression levels in the initial state of cells.

7. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the gene expression dynamic prediction model establishment method according to any one of claims 1-3.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the gene expression dynamic prediction model establishment method according to any one of claims 1-3.