Evaluation method, device and system of tumor intervention small molecule effect and storage medium

By acquiring transcriptomic characterizations of tumors and normal tissues and using a neural network model based on the Transformer architecture to calculate cellular states after drug intervention, this approach solves the problem of existing technologies being unable to effectively assess the complexities of drug effects on heterogeneous tumors, enabling efficient, comprehensive evaluation and precise screening of drug efficacy.

CN121306599APending Publication Date: 2026-01-09ACADEMY OF MILITARY MEDICAL SCIENCES

Patent Information

Application Number
CN202511512423.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing tumor drug evaluation technologies cannot effectively simulate the complex situation of drugs in heterogeneous tumors, neglect the assessment of drug-resistant cell subpopulations, and lack objective and quantitative standards to measure the ability of drugs to reverse the state of tumor cells to a healthy state.

Method used

By acquiring transcriptomic characterizations of tumor and normal tissues, we used a neural network model based on the Transformer architecture to calculate the cellular state after drug intervention, evaluated the drug effect using distance measurement, and established an evaluation criterion based on "state recovery".

Benefits of technology

It enables efficient and comprehensive evaluation of drug efficacy, quickly screens drugs that can reverse tumor cell state to a healthy state, provides objective and quantitative drug evaluation standards, and improves the efficiency and accuracy of new drug development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121306599A_ABST
    Figure CN121306599A_ABST
Patent Text Reader

Abstract

The invention provides a method, device and system for evaluating the effect of intervening small molecules by tumors and a storage medium, and relates to the technical field of bioinformatics, the method comprises the following steps: obtaining a first transcriptome representation of a tumor cell population before target small molecule intervention and a second transcriptome representation of a normal tissue cell population corresponding to the tumors; inputting the first transcriptome representation and the condition vector into a prediction model to obtain a third transcriptome representation; calculating a first distance and a second distance; an intervention effect is evaluated based on a comparison of the first distance and the second distance. According to the method, the influence of the drug on the transcriptome of the whole cell population is efficiently predicted, and the distance between the transcriptome and the health state is quantitatively compared, so that a set of brand new standard for improving the evaluation target from the traditional'killing cell 'to'reversing disease state' is established, and an objective decision basis is provided for developing more accurate drugs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics technology, and in particular to methods, devices, systems and storage media for evaluating the effects of small molecules on tumor intervention. Background Technology

[0002] Bioinformatics is an interdisciplinary field that uses applied mathematics, informatics, statistics, and computer science to study biological problems. In modern life science research, the widespread adoption of high-throughput technologies such as single-cell sequencing has generated massive amounts of high-dimensional biological data. Extracting valuable biological insights from this complex data has become crucial for advancing fields such as precision medicine and drug development. Particularly in oncology research, the use of artificial intelligence, especially deep learning and other computational methods, to analyze and model single-cell transcriptome data offers unprecedented opportunities for understanding disease mechanisms, discovering drug targets, and evaluating candidate drugs, representing a cutting-edge direction in computational biology.

[0003] In traditional oncology drug development, high-throughput in vitro cell line screening experiments are typically relied upon to initially assess the effectiveness of small molecule compounds, such as by measuring the half-maximal inhibitory concentration (IC50) to determine the drug's ability to kill cancer cells. With the development of computational technology, existing techniques are beginning to employ machine learning models to assist this process. These computational methods typically attempt to establish correlations between drug molecular characteristics, cancer cell gene expression profiles, and drug responses (such as cell viability). By training on known drug-cell line response datasets, these models can make preliminary computational predictions about the effects of new candidate drugs or in different cell lines, thereby improving the efficiency of drug screening to some extent.

[0004] However, existing drug efficacy evaluation techniques have significant limitations. First, tumor tissues are highly heterogeneous, composed of numerous cell subpopulations with varying gene expression and phenotypes. Most existing computational models predict outcomes based on the average response of the cell population, such as overall cell survival, while neglecting the differential effects of drugs on different cell subpopulations. This simplistic approach cannot realistically simulate the complexities of drug action on heterogeneous tumors, particularly potentially overlooking the assessment of drug-resistant cell subpopulations that can lead to treatment failure and tumor recurrence. Second, single-cell transcriptomics data are inherently high-dimensional, highly sparsity, and susceptible to technical noise such as batch effects. Many existing models struggle to learn robust, transferable cell state representations that eliminate technical bias when processing this type of data, often limiting their generalization ability and predictive accuracy.

[0005] In summary, the current technical challenges in this field mainly focus on the following aspects: Firstly, there is a lack of computational models capable of effectively simulating the impact of drugs on the overall distribution of heterogeneous cell populations, leading to deviations between predicted results and actual biological processes. Secondly, regarding evaluation criteria, existing technologies primarily focus on assessing the cytotoxic toxicity of drugs, lacking an objective and quantitative standard to measure the extent to which a drug can reverse the entire tumor cell population from "disease" to "health." These issues collectively limit the depth and reliability of computational methods in the field of precision drug evaluation and screening. Summary of the Invention

[0006] In a first aspect, the present invention provides a method for evaluating the efficacy of small molecules in tumor intervention, comprising: Obtain the first transcriptome characterization of the tumor cell population before intervention with the target small molecule, and the second transcriptome characterization of the normal tissue cell population corresponding to the tumor. The first transcriptome characterization and the feature vector representing the target small molecule are input into the prediction model to obtain the third transcriptome characterization of the tumor cell population after intervention with the target small molecule. Calculate a first distance between the distribution of the first transcriptome representation and the distribution of the second transcriptome representation, and a second distance between the distribution of the third transcriptome representation and the distribution of the second transcriptome representation; The intervention effect of the target small molecule is evaluated based on the comparison between the first distance and the second distance.

[0007] In an optional implementation, evaluating the intervention effect of the target small molecule based on a comparison of the first distance and the second distance includes: Determine whether the target small molecule is an effective small molecule; If so, then calculate the intervention effect score of the target small molecule.

[0008] In an optional implementation, determining whether the target small molecule is a valid small molecule includes: Calculate the third distance between the distribution of the first transcriptome representation and the distribution of the third transcriptome representation, and calculate the p-value of the third distance; When the second distance is less than the first distance and the p-value of the third distance is less than 0.01, the small molecule is determined to be an effective interventional small molecule.

[0009] In an optional implementation, the calculation expression for the intervention effect score is: ; Among them, S 干预Dis represents the intervention effect score; Dis0 represents the first distance; Dis2 represents the second distance.

[0010] In an optional implementation, the method for constructing the prediction model includes: Obtain a training dataset; wherein the training dataset contains the first transcriptome characterization distribution of cell populations before various small molecule interventions, and the corresponding real transcriptome characterization distribution after interventions as observed in the experiment; A neural network model based on the Transformer architecture is constructed; wherein the model is configured to receive the transcriptome representations of a set of cells as input and output the corresponding predicted transcriptome representations; The model is trained by using a loss function based on the maximum mean difference, which minimizes the statistical distance between the transcriptome representation distribution predicted by the model based on the first transcriptome representation distribution and the corresponding real transcriptome representation distribution in the training dataset, thus obtaining the trained prediction model.

[0011] In an optional implementation, the method for constructing the training dataset includes: The raw single-cell gene expression data are preprocessed; the preprocessing includes at least one of screening protein-coding genes, normalizing cell sequencing depth, and logarithmic transformation. The preprocessed data is divided into a training subset and a test subset to form the training dataset; the training dataset includes all the data in the training subset and some perturbation data separated from the test subset.

[0012] In an optional implementation, the first transcriptome representation distribution and the real transcriptome representation distribution included in the training dataset are calculated by inputting single-cell gene expression data into a pre-trained base model based on the Transformer architecture.

[0013] In an optional implementation, the loss function based on the maximum mean difference is calculated using an energy distance kernel function; wherein the calculation expression of the kernel function is: ; Where k(u,v) represents the kernel function; u represents the transcriptome representation vector of the first cell; and v represents the transcriptome representation vector of the second cell.

[0014] In an optional implementation, the calculation expression for the loss function based on the maximum mean difference is: ; in, This represents the maximum mean difference loss; S represents the number of cells in the cell set. Transcriptome characterization representing cells from real, experimentally observed cell populations; represents the transcriptomic characterization of cells from the cell population generated by the model prediction; i and j represent the indices used in the summation calculation; , and Both represent kernel functions.

[0015] Secondly, the present invention provides an evaluation device for the efficacy of small molecules in tumor intervention, comprising: The acquisition module is used to acquire the first transcriptome characterization of the tumor cell population before intervention with the target small molecule, and the second transcriptome characterization of the normal tissue cell population corresponding to the tumor. The prediction module is used to input the first transcriptome characterization and the feature vector representing the target small molecule into the prediction model to obtain the third transcriptome characterization of the tumor cell population after intervention with the target small molecule. The calculation module is used to calculate a first distance between the distribution of the first transcriptome representation and the distribution of the second transcriptome representation, and a second distance between the distribution of the third transcriptome representation and the distribution of the second transcriptome representation; An evaluation module is used to evaluate the intervention effect of the target small molecule based on a comparison of the first distance and the second distance.

[0016] Thirdly, the present invention provides a computer system comprising a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the method for evaluating the effect of small molecules on tumor intervention as described in any of the foregoing embodiments.

[0017] Fourthly, the present invention provides a computer storage medium storing a computer program, which, when executed on a processor, implements a method for evaluating the effect of small molecules on tumor intervention according to any one of the foregoing embodiments.

[0018] This application provides a method, apparatus, system, and storage medium for evaluating the effects of small molecules on tumor intervention. The evaluation method achieves efficient virtual screening of candidate drugs by predicting the cell transcriptome state after small molecule intervention in a computer. It replaces traditional, resource-intensive, and time-consuming in vitro cell experiments with computational simulation, significantly shortening the early screening cycle in drug discovery. By inputting pre-intervention cell characterization and conditional vectors encoding the small molecule's identity, the predicted cell population state after drug action can be quickly obtained. This allows researchers to conduct preliminary evaluations of large-scale compound libraries at a lower cost and faster speed, thereby improving the overall efficiency of new drug development.

[0019] Secondly, this method provides a more comprehensive and in-depth dimension for evaluating drug efficacy. Traditional drug evaluation often relies on single or a few indicators, such as cell viability or changes in the expression of specific biomarkers. However, this method, by acquiring the transcriptomic characterization of the entire cell population, can holistically and comprehensively capture the complex effects of drugs on cellular state from the perspective of thousands of genes. It no longer views individual indicators in isolation, but rather considers the cell as a complete system. By calculating the distances between the entire distribution, it gains a deeper understanding of how drugs reshape the cellular gene expression network, thus more accurately reflecting the true mechanism of drug action.

[0020] Finally, this method establishes a novel and more biologically meaningful standard for evaluating drug efficacy. By calculating and comparing the "initial disease distance" (the distance between the tumor and normal tissue before intervention) and the "post-intervention disease distance" (the distance between the tumor and normal tissue after intervention), it shifts the goal of drug evaluation from simply "killing tumor cells" to "reversing the tumor cell state back to a healthy state." This evaluation logic based on "state restoration" can not only inhibit tumors but also uncover potential drugs that restore the function of tumor cells to normal cells, providing objective and quantitative decision-making basis for developing more precise treatment options with fewer side effects. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and therefore should not be considered as a limitation on the scope of protection of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of the hardware operating environment involved in an embodiment of the method for evaluating the effects of small molecules on tumor intervention according to the present invention; Figure 2This is a flowchart of Example 1 of the method for evaluating the effect of small molecules on tumor intervention according to the present invention; Figure 3 This is a detailed flowchart of step S400 in Example 2 of the method for evaluating the effect of small molecules on tumor intervention in this invention. Figure 4 This is a detailed flowchart of step S410 in Example 2 of the method for evaluating the effect of small molecules on tumor intervention in this invention. Figure 5 This is a flowchart illustrating steps S500 to S700 in Example 3 of the method for evaluating the effect of small molecules on tumor intervention in this invention. Figure 6 This is a detailed flowchart of step S510 in Example 3 of the method for evaluating the effect of small molecules on tumor intervention in this invention. Figure 7 This is a schematic diagram of the overall process of the deep learning method used in Example 4 of the evaluation method for the tumor intervention effect of small molecules of the present invention; Figure 8 This is a schematic diagram of the module connections of the evaluation device for the tumor intervention effect of small molecules according to the present invention. Detailed Implementation

[0023] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0024] The components of the embodiments of this application described and illustrated in the accompanying drawings can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0025] In the following, the terms “comprising,” “having,” and their cognates, which may be used in various embodiments of this application, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as excluding, firstly, the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more features, numbers, steps, operations, elements, components, or combinations thereof.

[0026] Furthermore, the terms "first," "second," and "third" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.

[0027] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of this application pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be construed as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of this application.

[0028] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0029] like Figure 1 The diagram shown is a structural schematic of the hardware operating environment of the terminal involved in an embodiment of the present invention.

[0030] The tumor intervention small molecule effect evaluation system of this invention can be a PC, or a mobile terminal device such as a smartphone, tablet, or portable computer. This tumor intervention small molecule effect evaluation system may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen, an input unit such as a keyboard, or a remote control; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed RAM memory or a stable memory, such as a disk storage device. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001. Optionally, the tumor intervention small molecule effect evaluation system may also include RF (Radio Frequency) circuitry, audio circuitry, a Wi-Fi module, etc.

[0031] Those skilled in the art will understand that Figure 1 The evaluation system for the effects of small molecules on tumor intervention shown is not intended to limit it and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. Figure 1 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a data interface control program, a network connection program, and an evaluation program for the effects of small molecules on tumor intervention.

[0032] Example 1 Reference Figure 2 This embodiment provides a method for evaluating the efficacy of small molecules in tumor intervention, including: Step S100: Obtain the first transcriptome characterization of the tumor cell population before intervention with the target small molecule, and the second transcriptome characterization of the normal tissue cell population corresponding to the tumor.

[0033] The goal of this step is to create a comprehensive and quantitative digital profile for both "disease state" and "health state." Instead of observing individual genes, it uses the entire transcriptome (the expression levels of all genes in a cell) to define the biological state of a cell population.

[0034] The processing involves taking raw single-cell gene expression data from tumor tissue and corresponding normal tissue and performing calculations using a deep learning-based model. This model can transform the high-dimensional, sparse gene expression data of each cell into a low-dimensional, dense numerical vector, which is the "transcriptome representation".

[0035] This step ultimately yields two results: a set of vectors representing the "tumor state before intervention," i.e., the first transcriptome representation; and a set of reference vectors representing the "normal healthy state," i.e., the second transcriptome representation.

[0036] This approach captures a holistic picture of the cellular state, offering a more comprehensive and robust view than relying on a few biomarkers. Furthermore, utilizing large-scale basic models helps to mitigate technical noise and biases introduced by different experiments, resulting in more biologically meaningful cellular state expressions.

[0037] Specifically, scRNA-seq data from publicly available sources or from sequencing tumor and adjacent normal tissues can be obtained. After preprocessing such as standardization and logarithmic transformation, the data is input into a pre-trained basic model (e.g., scGPT), and the model's output is the transcriptome representation vector for each cell.

[0038] Step S200: Input the first transcriptome characterization and the feature vector representing the target small molecule into the prediction model to obtain the third transcriptome characterization of the tumor cell population after intervention with the target small molecule.

[0039] The feature vector representing the target small molecule, as described above, refers to the process of converting the chemical information of a compound into a numerical vector in order to enable the neural network model to process and recognize different small molecule compounds. This feature vector is a digital, machine-readable representation of the small molecule, which can capture molecular properties that are crucial for predicting cellular transcriptomic responses, allowing the model to learn the complex relationship between molecular structure or properties and their biological effects.

[0040] Specifically, the generation of this feature vector can be achieved in several ways, including but not limited to the following: (1) Molecular fingerprinting: This method uses algorithms to encode the chemical structure information of small molecules into a fixed-length binary vector (or counting vector). Each bit in the vector represents the presence of a specific substructure, atom pair, or chemical feature in the molecule. For example, algorithms such as Extended Connectivity Fingerprinting (ECFP) or Molecular Access System (MACCS) keys can be used to generate it. Similar molecules will produce similar fingerprint vectors.

[0041] (2) Physicochemical property descriptor: This method calculates a series of numerical values ​​that can describe the physicochemical properties of small molecules and combines these values ​​into a vector. These properties may include molecular weight, lipid-water partition coefficient (LogP), polar surface area (TPSA), number of hydrogen bond donors and acceptors, etc.

[0042] (3) Learned embedding vectors: In some implementations, deep learning models (e.g., graph neural networks, GNNs) can be used to learn low-dimensional, dense vector representations directly from the graph structure of the molecule (atoms as nodes, chemical bonds as edges). This approach can automatically learn and capture more complex and abstract structural and chemical information in the molecule.

[0043] (4) One-hot encoding: For a known library of small molecule compounds with a fixed number of molecules, a unique index can be assigned to each molecule. Its feature vector is a vector with a length equal to the total number of molecules in the library, with 1 at the corresponding index position and 0 at the other positions.

[0044] In practice, one or more of the above methods can be selected and combined to construct the final feature vector, which can then be input into the prediction model, depending on the needs of the prediction task.

[0045] This step involves conducting a "virtual drug intervention experiment" on a computer. It uses a trained predictive model to infer how a specific small molecule drug will affect a population of tumor cells.

[0046] The process can involve taking the "first transcriptome representation" (tumor state before intervention) obtained in the previous step and a "conditional vector" representing a specific small molecule as inputs, and feeding them into a prediction model for computation. This prediction model (e.g., a neural network based on the Transformer architecture) will output a new set of transcriptome representation vectors, which is the result of this step: the third transcriptome representation, which represents the "tumor state after drug intervention" predicted by the model.

[0047] Its greatest advantages are high efficiency and low cost. It can quickly predict the effects of thousands of candidate small molecules without the need for time-consuming and expensive real biological experiments, greatly accelerating the early drug screening process.

[0048] This step requires a pre-trained prediction model. The input "conditional vector" can take many forms; for example, it can use a unique number to represent different small molecules (one-hot encoding) or a molecular fingerprint that describes its chemical structure as a vector.

[0049] Step S300: Calculate the first distance between the distribution of the first transcriptome representation and the distribution of the second transcriptome representation, and the second distance between the distribution of the third transcriptome representation and the distribution of the second transcriptome representation.

[0050] This step quantifies "how severe the disease is" and "how much of the disease state remains after treatment." It uses a specific numerical value to represent the degree of difference between the states of different cell populations.

[0051] The processing can employ statistical algorithms capable of measuring the differences between the distributions of two sets of data, such as E-distance. This algorithm calculates the distance between the two pairs of distributions. This step yields two numerical results: (1) First distance (can be denoted as Dis0): The distance between the "tumor before intervention" and the "normal tissue", representing the initial disease distance.

[0052] (2) Second distance (which can be denoted as Dis2): The predicted distance between the "tumor after intervention" and the "normal tissue", representing the distance of residual disease after intervention.

[0053] This step transforms a complex biological problem into an objective, measurable mathematical problem. Specific numerical values ​​allow for an intuitive and unbiased comparison of differences between different states, providing a solid data foundation for subsequent evaluations.

[0054] Specifically, statistical software packages or algorithm libraries that include E-distance calculation, such as E-test, can be used. By inputting the cell transcriptome characterization vector sets of the two states, the E-distance value between them can be calculated.

[0055] Step S400: Evaluate the intervention effect of the target small molecule based on the comparison between the first distance and the second distance.

[0056] This step is the final decision-making step. It determines whether the target small molecule has achieved the expected therapeutic effect based on the previously calculated distance value.

[0057] This step compares the magnitudes of the "first distance" and the "second distance." The result provides an evaluation of the intervention effect of this small molecule. If the "second distance" is smaller than the "first distance," it means that the drug has brought tumor cells closer to a normal state, and is considered an effective intervention.

[0058] This step establishes a novel and more clinically meaningful evaluation standard. It no longer focuses solely on whether a drug kills tumor cells, but rather assesses whether it can "correct" or "reverse" abnormal cellular states back to normal. This evaluation logic helps to screen for candidate drugs with superior mechanisms of action that can truly restore the body's health.

[0059] Example 2 Reference Figure 3 This embodiment provides a method for evaluating the effect of small molecules on tumor intervention, including: step S400, evaluating the intervention effect of the target small molecule based on a comparison of the first distance and the second distance, including: Step S410: Determine whether the target small molecule is an effective small molecule.

[0060] This step prepares training data for model learning. The aim is to screen for molecules with genuine therapeutic potential from a large pool of candidate compounds. It establishes a rigorous set of criteria to define what constitutes an "effective intervention."

[0061] The processing can involve single-condition or multi-condition logical judgments. A target small molecule must meet certain conditions to be considered "effective".

[0062] Further reference Figure 4 Step S410, determining whether the target small molecule is an effective small molecule, includes: Step S411: Calculate the third distance between the distribution of the first transcriptome representation and the distribution of the third transcriptome representation, and calculate the p-value of the third distance.

[0063] Step S412: When the second distance is less than the first distance and the p value of the third distance is less than 0.01, the small molecule is determined to be an effective intervention small molecule.

[0064] Drug intervention must induce a sufficiently strong and statistically significant change in cell state. This requires calculating the distance between the "tumor before intervention" and the "tumor after intervention" (i.e., the third distance, denoted as Dis1) and performing a statistical test (E-test) to obtain a p-value. This p-value must be less than a pre-defined significance threshold (e.g., 0.01).

[0065] The effect of the drug must be benign, that is, it makes the tumor state return to the normal state. This is achieved by comparing the first distance (Dis0) and the second distance (Dis2), and the condition Dis2 < Dis0 must be satisfied.

[0066] The final result of this step is to give a binary classification conclusion for each small molecule: "effective" or "ineffective".

[0067] The rigor of this dual - standard design lies in that it not only requires the drug to "toggle" the cell state, but also requires it to "toggle" in the "correct direction" and with sufficient "strength". This can effectively screen out those compounds with weak, random or even harmful effects, improving the reliability of the screening results.

[0068] Step S420, if so, calculate the intervention effect score of the target small molecule.

[0069] After screening out all "effective" small molecules in the first step, the purpose of this step is to rank these effective molecules. It provides a quantitative score to measure how much "credit" (importance) each effective drug has in reversing the tumor state back to the normal state.

[0070] Specifically, for each small molecule determined to be "effective", the processing can be to use the existing Dis0 and Dis2 values and substitute them into a specific mathematical formula for calculation. The result of this step is to generate a specific intervention effect score (Score) for each effective small molecule.

[0071] This scoring system provides an intuitive and standardized index for horizontally comparing the advantages and disadvantages of different effective drugs. The higher the score, the better the effect of the drug. This provides clear and objective data support for subsequent drug development decisions (for example, which candidate drug should be prioritized to enter the next stage). The definition of this score is also very ingenious. It represents the "percentage of the disease state distance eliminated by drug intervention", which is easy to understand and interpret.

[0072] In addition, if not, step S410 can be returned.

[0073] Furthermore, the calculation expression (Formula 1) of the intervention effect score is: ; where S 干预The formula represents the intervention effect score; Dis0 represents the first distance, which can be understood as the total distance that needs to be overcome from the "tumor state" to the "healthy state"; Dis2 represents the second distance, which is the distance "remaining" after drug intervention. Dis0 - Dis2 can represent how much distance the drug helps to "eliminate" or "close"; the meaning of the whole formula is to calculate the percentage of "eliminated distance" to "total distance".

[0074] For example, suppose there are two effective drugs, A and B. The initial disease distance Dis0 of the tumor is 10.0. After drug A intervention, the remaining distance Dis2 is 3.0. Its score is (10.0 - 3.0) / 10.0 = 0.7. After drug B intervention, the remaining distance Dis2 is 2.0. Its score is (10.0 - 2.0) / 10.0 = 0.8. Therefore, the conclusion is that drug B is more effective than drug A because it eliminates a larger proportion of the disease distance.

[0075] Example 3 Reference Figure 5 This embodiment provides a method for evaluating the efficacy of small molecules in tumor intervention, wherein the method for constructing the prediction model includes: Step S500: Obtain the training dataset; wherein the training dataset includes the first transcriptome characterization distribution of the cell population before various small molecule interventions, and the corresponding real transcriptome characterization distribution after intervention as observed in the experiment.

[0076] This step prepares the "teaching materials" (training data) for the model's learning. These materials need to contain numerous "questions" and corresponding "standard answers." Here, the "questions" are the state of the cells before drug administration, and the "standard answers" are the actual observed state of these cells after drug administration.

[0077] The process begins by collecting a large-scale, high-quality single-cell perturbation experimental dataset, such as gene expression data from various cancer cell lines treated with multiple small-molecule drugs. Subsequently, this raw gene expression data (typically high-dimensional, sparse counts) is transformed into low-dimensional, dense "transcriptome representation" vectors using a large-scale foundational model. The final result of this step is a structured training dataset. Each record in the dataset contains two parts: a set of transcriptome representation vectors representing the "pre-intervention" state, and a corresponding set of transcriptome representation vectors representing the "post-intervention" state.

[0078] Training a model using a large-scale, diverse dataset ensures that the final model possesses good robustness and generalization ability, enabling it to adapt to different cellular environments and drug types. Converting raw data into transcriptomic representations helps mitigate technical biases such as batch effects, allowing the model to learn more fundamental biological states.

[0079] Specifically, a large, publicly available single-cell perturbation map (such as the Tahoe-100M dataset) can be selected as the source of raw data. After standardizing and preprocessing the raw data, a pre-trained base model (such as scGPT) is used to calculate the embedding vector for each cell, thereby obtaining the transcriptome representation distribution required for training.

[0080] In some implementations, reference Figure 6 The method for constructing the training dataset in step S500 includes: Step S510: Preprocess the raw single-cell gene expression data; the preprocessing includes at least one of screening protein-coding genes, normalizing cell sequencing depth, and logarithmic transformation.

[0081] Step S520: Divide the preprocessed data into a training subset and a test subset to form the training dataset; the training dataset includes all the data in the training subset and some perturbation data divided from the test subset.

[0082] In some implementations, the first transcriptome representation distribution and the real transcriptome representation distribution included in the training dataset are calculated by inputting single-cell gene expression data into a pre-trained base model based on the Transformer architecture.

[0083] Step S600: Construct a neural network model based on the Transformer architecture; wherein the model is configured to receive the transcriptome representations of a set of cells as input and output the corresponding predicted transcriptome representations.

[0084] This step involves designing and building the learning tool itself, that is, defining the skeleton of the predictive model. The process can involve using a deep learning framework to design a neural network based on the Transformer architecture. The core design idea is that the model's input and output are not data from individual cells, but rather a collection of representations of the entire cell population. The result of this step is a neural network model with a specific structure that has not yet been trained.

[0085] Employing the Transformer architecture, particularly its internal self-attention mechanism, allows the model to capture the state differences and interactions between different cells within a cell population when processing cell population data. This design, which directly manipulates the cell ensemble, more realistically simulates the population response of heterogeneous tumors to drugs compared to models that predict the average population response.

[0086] Specifically, this can be implemented using Python's deep learning libraries (such as PyTorch). The model will consist of multiple stacked Transformer encoder layers, each containing a multi-head self-attention module and a feedforward network module to process and transform the input set of cellular representations.

[0087] Step S700: Using a loss function based on the maximum mean difference, the model is trained by minimizing the statistical distance between the transcriptome representation distribution predicted by the model based on the first transcriptome representation distribution and the corresponding real transcriptome representation distribution in the training dataset, thereby obtaining the trained prediction model.

[0088] This step is the model's "learning" process. It continuously adjusts and optimizes the model's internal parameters by repeatedly comparing the gap between the "model's predictions" and the "standard answer," making its predictive ability stronger and stronger.

[0089] During training, the model receives the "pre-intervention" representations from the training data and generates a "predicted post-intervention" representation distribution. Then, a loss function based on maximum mean difference (MMD) calculates the statistical distance, or dissimilarity, between this "predicted distribution" and the "true distribution" in the dataset. This dissimilarity guides an optimization algorithm (such as gradient descent) to update the model's parameters, aiming to minimize this dissimilarity (i.e., the value of the loss function). This process iterates until the model's predictions are sufficiently accurate. The final result of this step is a well-trained predictive model.

[0090] Using MMD as the loss function is a core advantage of this method. It doesn't just require the model to predict the "average effect" accurately, but rather that the model's prediction of the "distribution of the entire cell population" be as consistent as possible with reality. This allows the model to learn the complex effects of the drug on the distribution of the entire cell population, not just the average effect.

[0091] Furthermore, the calculation expression (Formula 2) for the loss function based on the maximum mean difference is as follows: ; in, This represents the maximum mean difference loss; S represents the number of cells in the cell set. Transcriptome characterization representing cells from real, experimentally observed cell populations; represents the transcriptomic characterization of cells from the cell population generated by the model prediction; i and j represent the indices used in the summation calculation; , and Both represent kernel functions.

[0092] This step is implemented through an iterative training loop. In each loop, a batch of data is taken from the training dataset, and the model's forward propagation (prediction), MMD loss calculation, backpropagation (gradient calculation), and parameter updates are performed.

[0093] The MMD loss function used is calculated using the formula shown in Formula 3. In this formula, This is the loss value that needs to be minimized. X={x1,...,x S} represents the set of true representations containing S cells, while ={ 1,..., S} represents the set of representations predicted by the model. k(u,v) is a kernel function used to measure the relationship between any two cell representations. The entire formula quantifies the differences in the distributions of the two populations by comparing the "predicted population" within itself, the "real population" within itself, and the relationship between the "predicted and real populations".

[0094] Furthermore, the loss function based on the maximum mean difference is calculated using the energy distance kernel function; wherein, the calculation expression of the kernel function (Formula 3) is: ; Where k(u,v) represents the kernel function; u represents the transcriptome representation vector of the first cell; and v represents the transcriptome representation vector of the second cell.

[0095] Example 4 Reference Figure 7 To better illustrate the evaluation method for the tumor intervention effect of small molecules provided in the foregoing embodiments, this embodiment provides the following specific implementation method: 1. Collect scRNA-seq data of cancer cell lines from the Tahoe 100M database: A solid data foundation is crucial for model training. Training a model using a large-scale, high-quality single-cell perturbation dataset can yield an effective model for evaluating small-molecule drugs used in tumor intervention. Therefore, this embodiment selects the Tahoe-100M dataset, a large single-cell perturbation atlas containing transcriptomic response data from 50 different cancer cell lines treated under 1,138 conditions (involving 380 different small-molecule drugs). This dataset contains over 100 million perturbed cells, and its scale and diversity provide a solid foundation for training a model that can generalize to different cellular environments.

[0096] 2. Construct datasets for model training and testing: To ensure the robustness and generalization ability of the model, this embodiment performs rigorous preprocessing and dataset partitioning on the raw data. The data preprocessing steps include: screening expression measurements of approximately 19,790 human protein-coding genes, normalizing them using standardized cell sequencing depth (normalized to 10,000 UMIs per cell), and then performing logarithmic transformation.

[0097] This logarithmic transformation (log1p) is calculated using the following formula (Formula 4), where x is the normalized gene expression count. This is the converted value: .

[0098] This transformation helps stabilize data variance and makes the data more suitable for processing by deep learning models. In terms of dataset partitioning, this embodiment selects 5 cell lines from 50 cell lines as the reserved test set using Principal Component Analysis (PCA). During training, the model uses all data from the remaining cell lines and additionally adds 30% perturbation data from the test cell lines.

[0099] 3. Calculate transcriptome characterization of cancer cell lines using the scGPT basic model: To address the technical noise introduced by different experimental platforms and to learn a transferable, universal cell state representation, this method employs the scGPT basic model to compute the transcriptome characterization of cells.

[0100] scGPT is a generative pre-trained model based on the Transformer architecture. It has been pre-trained on data from over 33 million cells and is designed to build a foundational model for single-cell omics data. By leveraging its powerful learning capabilities, scGPT can transform high-dimensional, sparse single-cell gene expression profile data into a low-dimensional, dense vector representation (i.e., transcriptome representation). This representation can capture key biological states of cells while mitigating technical biases such as batch effects to some extent.

[0101] 4. Establish a neural network computational model based on the transformer architecture to achieve transcriptome characterization prediction for small molecule interventions: The core of the method presented in this embodiment is to establish a novel neural network computational model based on the Transformer architecture to predict cellular transcriptome characterization after small molecule intervention. The key innovation of the model lies in the fact that it does not process individual cells, but rather operates directly on cell ensembles. This design allows the model to capture the state differences and interactions within the cell population through its internal self-attention mechanism, thereby more realistically simulating the collective response of heterogeneous tumors to drugs.

[0102] The model takes as input a set of transcriptomic representations (generated by scGPT) representing "untreated" control cells and a condition encoding information about a specific small molecule intervention, and then predicts the corresponding set of representations for these cells after intervention. To achieve this, the model training process aims to minimize the statistical distance between the predicted cell population distribution and the experimentally observed true distribution. In this embodiment, the Maximum Mean Discrepancy (MMD) is used as the loss function to guide model optimization, and its calculation formula is shown in Equation 2 above. This loss function ensures that the model not only learns the average effect of the drug, but also accurately reproduces its complex impact on the entire cell population distribution.

[0103] 5. Validate the model's predictive performance on an external test set: This framework evaluates the model's predictions across multiple dimensions, primarily focusing on the three core outputs of single-cell perturbation experiments: gene expression counts, differential expression statistics, and the strength of the perturbation effect. Evaluation metrics include: perturbation discrimination score, Pearson correlation between predicted and actual gene expression changes, overlap accuracy of differentially expressed genes, area under the precision-recall curve (AUPRC), and Spearman correlation for logarithmic fold changes.

[0104] 6. Collect corresponding scRNA-seq data from normal tissues for specific tumors: In the application phase, to evaluate whether small molecules can "reverse" the state of tumor cells to a healthy state, it is necessary to collect single-cell transcriptome (scRNA-seq) data from normal tissues corresponding to the target tumor. This data can be obtained from large, publicly available databases, which are also used to train the base model to ensure consistency in data sources and processing methods.

[0105] 7. Predict the transcriptomic characterization of tumors before small molecule intervention (Embedding_pre), after small molecule intervention (Embedding_aft), and the transcriptomic characterization of normal tissues without intervention (Embedding_normal): Using the pre-trained scGPT and the neural network model established in step 4, a virtual intervention experiment was conducted on a specific tumor sample: (1) Embedding_pre (Tumor characterization before intervention, first transcriptome characterization): Input scRNA-seq data from the patient's tumor tissue into the scGPT basic model, calculate the embedding vector of each cell, and form the tumor characterization distribution before intervention.

[0106] (2) Embedding_normal (normal tissue characterization, second transcriptome characterization): Input the scRNA-seq data of the corresponding normal tissue collected in step 6 into the scGPT basic model, calculate the embedding vector of normal cells, and form the reference characterization distribution of normal tissue.

[0107] (3) Embedding_aft (Tumor characterization after intervention, third transcriptome characterization): The tumor cell characterization before intervention (Embedding_pre) is used as input, along with the information of the target small molecule, and input into the neural network model established in step 4 to predict the distribution of the transcriptome characterization of tumor cells after receiving the small molecule intervention.

[0108] 8. Calculate the E-distance Dis0 (first distance) between Embedding_pre and Embedding_normal, calculate the E-distance Dis1 between Embedding_pre and Embedding_aft, calculate the E-distance Dis2 (second distance) between Embedding_normal and Embedding_aft, and calculate the P-value for each distance using E-test: To quantify the differences between different cell states, this embodiment uses E-distance for calculation. E-distance is mathematically equivalent to the MMD loss used during model training. This embodiment calculates the following three distances: (1) Dis0 (first distance): E-distance between Embedding_pre (tumor before intervention) and Embedding_normal (normal tissue), representing the initial "disease distance".

[0109] (2)Dis1 (Third Distance): The E-distance between Embedding_pre (tumor before intervention) and Embedding_aft (tumor after intervention), representing the intensity of the drug effect.

[0110] (3)Dis2 (Second Distance): The E-distance between Embedding_aft (tumor after intervention) and Embedding_normal (normal tissue), representing the "residual disease distance" after intervention.

[0111] Meanwhile, the p-value of each distance is calculated using the E-test to determine whether the observed distribution difference is statistically significant.

[0112] 9. Definition: Small molecules with a p-value of Dis1 < 0.01 and Dis2 < Dis0 are effective intervention small molecules, and the intervention small molecule effect score = (Dis0 - Dis2) / Dis0.

[0113] Based on the above calculation results, a set of evaluation criteria is proposed in this embodiment to screen effective small molecule drugs and rank their effects. Definition of effective intervention: A small molecule is considered "effective intervention" if it meets the following two conditions simultaneously: significant drug effect (p-value of Dis1 < 0.01) and state recovery towards normal (Dis2 < Dis0).

[0114] Intervention effect score: For effective small molecules, their effect score (Score) is calculated by Formula 1. This score represents "the percentage of the disease state distance eliminated by drug intervention", and the higher the score, the better the effect of the small molecule in reversing the tumor state back to the normal state.

[0115] Reference Figure 8 , this application embodiment provides a device for evaluating the effect of tumor intervention small molecules, including: An acquisition module 10 for acquiring the first transcriptomic characterization of the tumor cell population before the target small molecule intervention, and the second transcriptomic characterization of the normal tissue cell population corresponding to the tumor; A prediction module 20 for inputting the first transcriptomic characterization and the feature vector representing the target small molecule into a prediction model to obtain the third transcriptomic characterization of the tumor cell population after the target small molecule intervention; A calculation module 30 for calculating the first distance between the distribution of the first transcriptomic characterization and the distribution of the second transcriptomic characterization, and the second distance between the distribution of the third transcriptomic characterization and the distribution of the second transcriptomic characterization; Evaluation module 40 is used to evaluate the intervention effect of the target small molecule based on the comparison between the first distance and the second distance.

[0116] It is understood that the device in this embodiment corresponds to the evaluation method of the tumor intervention small molecule effect in the above embodiments, and the options in the above embodiments are also applicable to this embodiment, so they will not be described again here.

[0117] This application provides a computer system comprising a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the method for evaluating the effect of small molecules on tumor intervention as described in any of the foregoing embodiments.

[0118] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Network Processor (NP), Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0119] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The memory is used to store computer programs, and the processor can execute the computer programs accordingly after receiving execution instructions.

[0120] This application provides a computer storage medium storing a computer program, which, when executed on a processor, implements a method for evaluating the effect of small molecules on tumor intervention according to any one of the foregoing embodiments.

[0121] The computer storage medium can be a readable storage medium, a non-volatile storage medium, or a volatile storage medium. For example, the computer storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0122] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that, in alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0123] In addition, the functional modules or units in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0124] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a smartphone, personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0125] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A method for evaluating the efficacy of small molecules in tumor intervention, characterized in that, include: Obtain the first transcriptome characterization of the tumor cell population before intervention with the target small molecule, and the second transcriptome characterization of the normal tissue cell population corresponding to the tumor. The first transcriptome characterization and the feature vector representing the target small molecule are input into the prediction model to obtain the third transcriptome characterization of the tumor cell population after intervention with the target small molecule. Calculate a first distance between the distribution of the first transcriptome representation and the distribution of the second transcriptome representation, and a second distance between the distribution of the third transcriptome representation and the distribution of the second transcriptome representation; The intervention effect of the target small molecule is evaluated based on the comparison between the first distance and the second distance.

2. The method for evaluating the efficacy of small molecules in tumor intervention as described in claim 1, characterized in that, The evaluation of the intervention effect of the target small molecule based on the comparison of the first distance and the second distance includes: Determine whether the target small molecule is an effective small molecule; If so, then calculate the intervention effect score of the target small molecule.

3. The method for evaluating the efficacy of small molecules in tumor intervention as described in claim 2, characterized in that, The determination of whether the target small molecule is a valid small molecule includes: Calculate the third distance between the distribution of the first transcriptome representation and the distribution of the third transcriptome representation, and calculate the p-value of the third distance; When the second distance is less than the first distance and the p-value of the third distance is less than 0.01, the small molecule is determined to be an effective interventional small molecule.

4. The method for evaluating the efficacy of small molecules in tumor intervention as described in claim 2, characterized in that, The formula for calculating the intervention effect score is: ; Among them, S 干预 Dis represents the intervention effect score; Dis0 represents the first distance; Dis2 represents the second distance.

5. The method for evaluating the efficacy of small molecules in tumor intervention as described in claim 1, characterized in that, The method for constructing the prediction model includes: Obtain a training dataset; wherein the training dataset contains the first transcriptome characterization distribution of cell populations before various small molecule interventions, and the corresponding real transcriptome characterization distribution after interventions as observed in the experiment; A neural network model based on the Transformer architecture is constructed; wherein the model is configured to receive the transcriptome representations of a set of cells as input and output the corresponding predicted transcriptome representations; The model is trained by using a loss function based on the maximum mean difference, which minimizes the statistical distance between the transcriptome representation distribution predicted by the model based on the first transcriptome representation distribution and the corresponding real transcriptome representation distribution in the training dataset, thus obtaining the trained prediction model.

6. The method for evaluating the efficacy of small molecules in tumor intervention as described in claim 5, characterized in that, The method for constructing the training dataset includes: The raw single-cell gene expression data are preprocessed; the preprocessing includes at least one of screening protein-coding genes, normalizing cell sequencing depth, and logarithmic transformation. The preprocessed data is divided into a training subset and a test subset to form the training dataset; the training dataset includes all the data in the training subset and some perturbation data separated from the test subset.

7. The method for evaluating the efficacy of small molecules in tumor intervention as described in claim 5, characterized in that, The first transcriptome representation distribution and the true transcriptome representation distribution included in the training dataset are calculated by inputting single-cell gene expression data into a pre-trained base model based on the Transformer architecture; and / or, The loss function based on the maximum mean difference is calculated using an energy distance kernel function; wherein, the calculation expression of the kernel function is: ; Where k(u,v) represents the kernel function; u represents the transcriptome representation vector of the first cell; v represents the transcriptome representation vector of the second cell; and / or, The calculation expression for the loss function based on the maximum mean difference is as follows: ; in, This represents the maximum mean difference loss; S represents the number of cells in the cell set. Transcriptome characterization representing cells from real, experimentally observed cell populations; represents the transcriptomic characterization of cells from the cell population generated by the model prediction; i and j represent the indices used in the summation calculation; , and Both represent kernel functions.

8. A device for evaluating the efficacy of small molecules in tumor intervention, characterized in that, include: The acquisition module is used to acquire the first transcriptome characterization of the tumor cell population before intervention with the target small molecule, and the second transcriptome characterization of the normal tissue cell population corresponding to the tumor. The prediction module is used to input the first transcriptome characterization and the feature vector representing the target small molecule into the prediction model to obtain the third transcriptome characterization of the tumor cell population after intervention with the target small molecule. The calculation module is used to calculate a first distance between the distribution of the first transcriptome representation and the distribution of the second transcriptome representation, and a second distance between the distribution of the third transcriptome representation and the distribution of the second transcriptome representation; An evaluation module is used to evaluate the intervention effect of the target small molecule based on a comparison of the first distance and the second distance.

9. A computer system, characterized in that, The computer system includes a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the method for evaluating the effect of small molecules on tumor intervention as described in any one of claims 1-7.

10. A computer storage medium, characterized in that, It stores a computer program, which, when executed on a processor, implements the method for evaluating the effect of small molecules on tumor intervention according to any one of claims 1-7.

Citation Information

Patent Citations

  • Poultry epidemic disease treatment effect quantitative evaluation method and system

    CN117789989A

  • Compound perturbation cell line gene transcription profile prediction method, device, equipment and medium

    CN120015119A

  • Method and system for evaluating treatment effect of traditional Chinese medicine based on single cell and space transcriptome data

    CN120452838A

  • Drug relocation method and system

    CN120564891A

Cited By

  • Target intervention object determination method and device, equipment, medium and product

    CN122135781A