Nerve cell regeneration drug target delineation method based on large model
By constructing dynamic knowledge graphs and digital twin models, and integrating multimodal data for virtual intervention and experimental feedback, the problem of identifying drug targets for neural cell regeneration has been solved, achieving high efficiency and accuracy in target screening and improving clinical translation efficiency.
Patent Information
- Application Number
- CN202511428064.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-02-24
AI Technical Summary
Existing technologies face significant challenges in accurately identifying drug targets for neural cell regeneration. Target validation is a lengthy, costly, and low-success-rate process. Traditional methods cannot effectively integrate multimodal data and lack the ability to dynamically simulate and virtually intervene in the neural regeneration process, resulting in low clinical translation efficiency.
By constructing a dynamic knowledge graph and digital twin model based on a large model, and integrating multimodal biomedical data, virtual intervention and experimental feedback are conducted to optimize the target screening process, including acquiring multimodal data, constructing a spatiotemporal dynamic knowledge graph, digital twin model, virtual intervention, and experimental verification.
It improves the accuracy and efficiency of target screening, shortens the preclinical screening cycle, reduces experimental resource consumption, and increases the success rate of clinical translation of drug targets.
Smart Images

Figure CN121565236A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of drug target delineation technology, and in particular to a method for delineating drug targets for neural cell regeneration based on a large model. Background Technology
[0002] In the clinical treatment of neurodegenerative diseases and central nervous system injuries, nerve cell regeneration is a core breakthrough for repairing damaged nerve function and reversing disease progression. However, the development of nerve cell regeneration drugs currently faces two major bottlenecks: first, the precise delineation of drug targets is extremely difficult; and second, the target validation process is lengthy, costly, and has a low success rate, severely restricting the efficiency of clinical translation.
[0003] Currently, drug target discovery primarily relies on traditional biological experimental methods, such as high-throughput screening based on genomics and proteomics, and phenotypic validation using animal models. However, these methods are typically based on single types of data, making it difficult to comprehensively capture the dynamic panorama of multi-factor and multi-pathway synergistic effects in the complex physiological process of neurogenesis. Furthermore, traditional in vitro cell models or animal models differ significantly from the human physiological environment, resulting in a high failure rate for candidate targets screened in the preclinical stage in subsequent human clinical trials, leading to substantial time and resource consumption.
[0004] In recent years, computational biology and artificial intelligence technologies have utilized static knowledge graphs to integrate some biomedical knowledge or used machine learning models to mine public databases. However, most models cannot effectively integrate multimodal data and reveal their inherent spatiotemporal dynamic relationships; on the other hand, they are essentially static predictive models, lacking the ability to dynamically simulate and virtually intervene in neural regeneration processes, and unable to conduct in-depth dry experimental evaluations of the effectiveness and potential side effects of targets before wet experimental validation. Summary of the Invention
[0005] The purpose of this invention is to provide a method for identifying drug targets for neural cell regeneration based on a large model. By constructing a digital twin model to dynamically simulate and optimize the target, and continuously correcting the model through experimental data, the accuracy of target screening and R&D efficiency can be improved.
[0006] To achieve the above objectives, this invention provides a method for identifying drug targets for neural cell regeneration based on a large model, the method comprising: S11. Acquire and fuse at least two types of multimodal biomedical data reflecting the process of nerve cell regeneration to construct a dynamic knowledge graph; S12. Based on dynamic knowledge graphs, construct a digital twin model that simulates the regeneration of target nerve cells; S13. Identify potential drug targets based on digital twin models, conduct virtual interventions, and generate prediction results of nerve cell regeneration responses after intervention; S14. Based on the prediction results, experimental interventions are applied to potential drug targets in observable biological models, and quantifiable signals related to nerve cell regeneration are collected to obtain experimental feedback data. S15. Feed the experimental feedback data back to the digital twin model for iteration.
[0007] Furthermore, a spatiotemporal dynamic knowledge graph for the field of neural cell regeneration is constructed based on multimodal biomedical data, specifically including: S21. Acquire multimodal biomedical data, wherein the multimodal biomedical data includes at least proteomic data, metabolomic data, molecular interaction data, single-cell transcriptomic data, and spatial transcriptomic data; S22. Preprocess and align the multimodal biomedical data with entities, extract biomedical entities and their relationships, and add corresponding time dimension labels. S23. Construct an initial static knowledge graph using biomedical entities as nodes, relationships between entities as edges, and data under time dimension labels as attributes. S24. A spatiotemporal graph neural network is used to perform dynamic representation learning on the initial static knowledge graph to generate a spatiotemporal dynamic knowledge graph.
[0008] Furthermore, a digital twin model of the target nerve cell is constructed based on a dynamic knowledge graph, specifically including: S31. Extract the biological entities and corresponding dynamic attributes related to the functional state of the target nerve cell from the spatiotemporal dynamic knowledge graph, and convert them into state vectors; S32. Based on the interaction relationships between entities in the spatiotemporal dynamic knowledge graph, construct a state transition function that controls the evolution of the state vector over time. S33. Couple the state vector with the state transition function to generate a digital twin model of the target nerve cell.
[0009] Furthermore, potential drug targets are identified based on digital twin models, and virtual interventions are conducted, specifically including: S41. When the digital twin model reaches the target time step, extract the sub-vector corresponding to the potential drug target from the global state vector to obtain a snapshot of the target state. S42. Generate an intervention predictor according to a preset intervention type, wherein the intervention predictor includes at least one of inhibition, activation, knockout or reversible blocking; S43. Apply the interference vector to the subvector to obtain the instantaneous state after target intervention, and write it back to the global state vector. S44. Using the global state vector after intervention as the initial value, call the state transition function to roll forward simulation to the end of the experiment and generate the regeneration trajectory after intervention. S45. Compare the pre-intervention trajectory with the post-intervention trajectory, calculate the counterfactual effect value, and generate a prediction result of the nerve cell regeneration response after intervention.
[0010] Furthermore, the generation of prediction results specifically includes: S51. Input the pre-intervention trajectory and post-intervention trajectory into the time-series graph neural network to obtain the corresponding first embedding vector and second embedding vector, and perform signal purification on the first embedding vector and the second embedding vector. S52. After concatenating the purified first and second embedding vectors, input them into the counterfactual decoder, output the prediction vector, and simultaneously estimate the uncertainty covariance matrix of the prediction vector. S53. Reverse decode the predicted vector to generate node contribution scores between target points, pathways and phenotypes. S54. Based on the prediction vector, uncertainty covariance matrix, and node contribution score, a counterfactual prediction package is generated as the prediction result of the regenerative response after intervention.
[0011] Furthermore, signal purification is performed on the first and second embedding vectors, specifically including: By comparing the loss functions, the first embedding vector is forced to move away from the second embedding vector and closer to the prior embedding of successful nerve cell regeneration, while moving away from the prior embedding of failed nerve cell regeneration.
[0012] Furthermore, the generation of experimental feedback data specifically includes: S61. Analyze the prediction results to obtain target identification, intervention methods, and intervention parameters; S62. Based on an observable biological model, experimental interventions are applied to the target through intervention parameters, wherein the experimental interventions include at least one of chemical intervention, genetic intervention, protein degradation intervention, optogenetic stimulation, or chemogenetic stimulation. S63. Simultaneously / sequentially acquire at least one quantifiable neural cell regeneration-related signal, and generate a regeneration phenotype vector after preprocessing; S64. Compare the regenerated phenotypic vector with the predicted vector of the digital twin model to obtain experimental feedback data.
[0013] Compared with the prior art, the beneficial effects of the present invention are: This invention provides a method for identifying drug targets for neural cell regeneration based on a large-scale model. By fusing multimodal biomedical data from transcriptomics, proteomics, and radiomics, a dynamic knowledge graph is constructed. This graph integrates multi-level regulatory information from genes, proteins, and the cellular microenvironment, fully reconstructing the dynamic process of neural cell regeneration. This allows for screening potential targets that more closely aligns with real physiological mechanisms, improving the accuracy of target prediction. A digital twin model is built based on the dynamic knowledge graph to rapidly simulate the neural cell regeneration process and virtually intervene in potential targets. The regenerative response of neural cells after intervention is predicted, allowing for the early elimination of ineffective or high-risk targets, shortening the preclinical screening cycle, and reducing experimental resource consumption. By feeding experimental data from observable biological models back to the digital twin model, the simulation accuracy of the neural regeneration process is continuously optimized, increasing the success rate of target-to-clinical drug translation. This invention improves the accuracy of target screening and research efficiency by constructing a digital twin model for dynamic simulation and optimization of targets, and by continuously calibrating the model using experimental data. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort. Figure 1 A schematic diagram of a method for identifying drug targets for neural cell regeneration based on a large model, provided in an embodiment of the present invention; Figure 2 A schematic diagram illustrating the process of constructing a spatiotemporal dynamic knowledge graph according to an embodiment of the present invention; Figure 3 A schematic diagram illustrating the process of constructing a digital twin model according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the process for identifying potential drug targets and performing virtual intervention based on a digital twin model, provided in an embodiment of the present invention. Figure 5 This is a schematic diagram of the process for generating prediction results provided in an embodiment of the present invention; Figure 6 This is a schematic diagram illustrating the process of generating experimental feedback data for an embodiment of the present invention. Detailed Implementation
[0015] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0016] Reference Figure 1 This embodiment provides a method for identifying drug targets for neural cell regeneration based on a large model, the method comprising: S11. Acquire and fuse at least two types of multimodal biomedical data reflecting the process of nerve cell regeneration to construct a dynamic knowledge graph.
[0017] S12. Based on dynamic knowledge graphs, construct a digital twin model to simulate the regeneration of target nerve cells.
[0018] S13. Identify potential drug targets based on a digital twin model, conduct virtual intervention, and generate prediction results of nerve cell regeneration response after intervention.
[0019] S14. Based on the prediction results, experimental interventions are applied to potential drug targets in an observable biological model, and quantifiable signals related to nerve cell regeneration are collected to obtain experimental feedback data.
[0020] S15. Feed the experimental feedback data back to the digital twin model for iteration.
[0021] In this embodiment, biomedical data from multiple sources and of multiple types, including genomics, proteomics, and metabolomics, are collected and integrated to construct a dynamic knowledge graph reflecting the complex internal relationships within the neural cell regeneration process. This integrates scattered and heterogeneous data into a unified knowledge network, avoiding the limitations of traditional single-data perspectives. Simultaneously, it reveals deep, non-linear relationships between potential targets, signaling pathways, and cellular functions, providing a knowledge foundation for subsequent analysis. Based on the dynamic knowledge graph, large-scale modeling techniques are used to train a digital twin model simulating the dynamic process of real neural cell regeneration, thereby simulating the intrinsic dynamics of biological processes. This helps in understanding how targets function and enhances research insights.
[0022] In digital twin models, virtual interventions are performed on a large number of potential targets, such as simulating the inhibition or activation of a gene / protein, and the regenerative response of nerve cells after intervention is predicted, along with potential off-target effects or ineffective results. The candidate targets with the best predictive performance are experimentally validated in observable biological models, i.e., wet experiments are conducted in real cell or tissue models, and new, high-quality biological data generated by the experiments are collected as feedback. The results from the wet experiments are fed back into the digital twin model for calibration and retraining, making the digital twin model more accurate and initiating a new round of more intelligent predictions. This method transforms a one-off, unidirectional study into a continuously optimized closed loop, accelerating target identification and deepening our understanding of the mechanisms of neural regeneration.
[0023] As a preferred embodiment, a spatiotemporal dynamic knowledge graph in the field of neural cell regeneration is constructed based on multimodal biomedical data, specifically including: S21. Acquire multimodal biomedical data, wherein the multimodal biomedical data includes at least proteomic data, metabolomic data, molecular interaction data, single-cell transcriptomic data, and spatial transcriptomic data.
[0024] S22. Preprocess and align the multimodal biomedical data with entities, extract biomedical entities and their relationships, and add corresponding time dimension labels.
[0025] S23. Construct an initial static knowledge graph using biomedical entities as nodes, relationships between entities as edges, and data under time dimension labels as attributes.
[0026] S24. A spatiotemporal graph neural network is used to perform dynamic representation learning on the initial static knowledge graph to generate a spatiotemporal dynamic knowledge graph.
[0027] In this embodiment, neural cell regeneration is a complex biological process that is highly coordinated in time and space, and traditional static knowledge graphs cannot meet its research needs. First, the neural cell regeneration process follows a strict temporal logic; for example, an inflammatory response is initiated immediately after injury, followed by cell proliferation and migration, and finally myelination and the formation of functional connections. The expression of key genes and proteins varies significantly at different stages, meaning that static graphs cannot capture the crucial information of when it occurs. Second, the brain or neural tissue is not a homogeneous structure. Cell types, microenvironments, and molecular activities vary greatly in different brain regions, and even at different locations within the same injury area. Single-cell technology and spatial transcriptomics have revealed this complexity, meaning that static graphs cannot reflect the spatial context of where it occurs. Therefore, integrating multimodal data in the spatiotemporal dimensions is a prerequisite for constructing a digital twin model that realistically simulates the neural regeneration process.
[0028] Specifically, multimodal biomedical data is acquired from public databases including GEO, TCGA, and Allen Brain Atlas, published literature, and proprietary experiments. Data from different sources and batches is normalized to filter out low-quality cell, gene, or protein data. The processed data undergoes dimensionality reduction and feature extraction; for example, PCA or UMAP is used to reduce the dimensionality of single-cell transcriptome data to identify key gene features, and spatial domains are identified for spatial transcriptome data. Natural language processing tools are used to extract entities and relationships from the processed multimodal biomedical data. Entities include the gene EGFR, cell type oligodendrocyte precursor cells, and biological process myelination; relationships include, for example, EGFR inhibits oligodendrocyte precursor cell differentiation.
[0029] Molecules and cell clusters identified in omics data are mapped to standard biomedical ontologies, such as GeneOntology and CellOntology, to achieve entity alignment. Each data point is labeled with a specific timestamp; for example, in a spinal cord injury model, it is labeled "1 day after injury," "3 days," "1 week," etc., transforming static relationships into dynamic events; such as "TNF gene expression is upregulated in microglia 3 days after injury." Node types and relationship types are determined. Node types include genes, proteins, cell types, biological processes, anatomical regions, etc.; relationship types include upregulation, inhibition, location, participation, etc. Using a graph database or storing the graph structure in a matrix, an initial static knowledge graph is constructed with aligned entities as nodes and relationships as edges. At this stage, time information is only one of the node attributes.
[0030] An initial static knowledge graph and timestamped multimodal attribute data are input into a spatiotemporal graph neural network. A graph convolutional network or graph attention network is used to learn the topological relationships between nodes in the graph; for example, learning the interaction patterns between the "oligodendrocyte precursor cell" node and its neighboring "neuron" and "BDNF" protein nodes. A recurrent neural network or temporal convolutional network is used to learn the sequential patterns of each node's attributes changing over time; for example, the model learns how the expression level of the "BDNF" protein changes over time after injury. ST-GNN combines the above two approaches, capturing the spatial structure of the graph at each time step and transmitting and updating the state of nodes along the time axis, thereby learning spatiotemporally coordinated evolutionary patterns; for example, the model learns that "on day 3 after injury, macrophages located at the injury edge inhibit the differentiation process of their neighboring OPC cells on day 5 by secreting IL-1β." Finally, a spatiotemporal dynamic knowledge graph is output to characterize how biological entities interact and evolve in a specific temporal and spatial context.
[0031] As a preferred embodiment, constructing a digital twin model of the target nerve cell based on a dynamic knowledge graph specifically includes: S31. Extract the biological entities and corresponding dynamic attributes related to the functional state of the target nerve cell from the spatiotemporal dynamic knowledge graph, and transform them into state vectors.
[0032] S32. Based on the interaction relationships between entities in the spatiotemporal dynamic knowledge graph, construct a state transition function that controls the evolution of the state vector over time.
[0033] S33. Couple the state vector with the state transition function to generate a digital twin model of the target nerve cell.
[0034] In this embodiment, although the spatiotemporal dynamic knowledge graph contains rich associations and temporal information, it is a descriptive knowledge base rather than a predictive simulator. That is, the knowledge graph represents relationships between entities, such as "protein A inhibits process B," but it cannot quantify the strength of these relationships, nor can it predict the final result under the simultaneous influence of multiple factors. A digital twin model is constructed by introducing mathematical relationships through state transition functions, aiming to simulate the causal state changes caused by these interactions. The regeneration of nerve cells is the result of the evolution of their internal molecular and cellular states over time. Only by combining the cell's state (state vector) with the evolutionary rules (state transition function) can we predict how the state will change in the next moment if the current state is as it is.
[0035] Specifically, from the spatiotemporal dynamic knowledge graph, core biological entities most relevant to the functional state of the target nerve cell are selected, such as the expression level of key genes, the phosphorylation state of proteins, and the concentration of metabolites. The dynamic attributes of these entities, such as their values or levels at specific time points, are extracted to form a multidimensional vector. Using domain knowledge or the attention mechanism in graph neural networks, it is determined which nodes and attributes have the greatest impact on the cell state. After normalizing the selected attribute values, such as mRNA expression level and protein abundance, they are arranged in order into a vector representing a snapshot of the digital state of the nerve cell at time t.
[0036] Define the evolutionary rules of cells, i.e., construct a state transition function to represent that the state at the next moment is jointly determined by the current state and the external environment. The state transition function needs to be learned from the interaction relationships in the knowledge graph. Transform the relationships in the knowledge graph into mathematical forms, such as activation and inhibition; for example, use a machine learning model to fit the state transition function. Use the state vector as the current state of the model and the state transition function F as the evolutionary rule of the model, and couple them through numerical integration or iterative loops to simulate the continuous evolution of the neuron state over time, thus constructing a digital twin model.
[0037] As a preferred embodiment, identifying potential drug targets based on a digital twin model and performing virtual intervention specifically includes: S41. When the digital twin model reaches the target time step, extract the sub-vectors corresponding to the potential drug targets from the global state vector to obtain a snapshot of the target state.
[0038] S42. Generate an intervention predictor based on a preset intervention type, wherein the intervention predictor includes at least one of inhibition, activation, knockout, or reversible blocking.
[0039] S43. Apply the interference vector to the subvector to obtain the instantaneous state after target intervention, and write it back to the global state vector.
[0040] S44. Using the global state vector after intervention as the initial value, call the state transition function to roll forward and simulate to the end of the experiment to generate the regeneration trajectory after intervention.
[0041] S45. Compare the pre-intervention trajectory with the post-intervention trajectory, calculate the counterfactual effect value, and generate a prediction result of the nerve cell regeneration response after intervention.
[0042] In this embodiment, after constructing a digital twin model of neural cell regeneration, the core objective is to leverage it for efficient drug target screening. Traditional methods involve direct intervention on real organisms, which suffers from significant drawbacks such as high cost, long cycle, and difficulty in parallelization. Digital twins, however, allow for large-scale, highly efficient, and risk-free pre-experiments in a virtual environment before investing in real experimental resources, significantly reducing early trial-and-error costs. Considering that drug intervention in reality is not a simple binary relationship of presence or absence, but rather involves differences in intensity and mode of action, these differences can be simulated using intervention quantum computing. Furthermore, drug intervention is a dynamic process whose effects change over time and trigger chain reactions; by forward-rolling simulation, dynamic and systematic long-term effects can be captured. Combining counterfactual reasoning with causal inference, the simulated trajectory after intervention (factual) is compared with the original trajectory without intervention (counterfactual), quantifying the net effect of the intervention and predicting "what kind of result will be produced after taking a certain type of intervention."
[0043] Specifically, when the digital twin model simulates key time points in nerve cell regeneration, the simulation is paused, such as the peak of inflammation after injury. Then, from the vector reflecting the global state of the cellular system, dimensions or sub-vectors related to candidate targets, such as the specific receptor protein EGFR, are located. These sub-vectors contain core information such as the target's current activity level and conformational state. Based on the intended mechanism of action of the drug design, intervention methods such as inhibition, activation, knockout, and reversible blockade are defined as mathematical intervention operators. The selected operators are then applied to the target-related sub-vectors to generate the instantaneous state after intervention, which is then written back into the global state vector to obtain the new initial state after intervention.
[0044] The model is restarted starting from the new initial state. The state transition function is invoked, and the model is integrated forward step by step according to the time step Δt to simulate the entire process of cell evolution from the intervention point to the experimental endpoint. Simultaneously, changes in key state variables such as neuron survival rate, axon length, and degree of myelination are recorded to form the regeneration trajectory after intervention. Finally, by quantitatively comparing the intervention trajectory with the original uninterrupted trajectory, the difference in the area under the regeneration index curve is calculated. This difference comprehensively reflects the degree of improvement in neural regeneration during the entire time process. The final output effect value is the quantitative prediction result of the intervention effect on this target.
[0045] As a preferred embodiment, the generation of the prediction result specifically includes: S51. Input the pre-intervention trajectory and post-intervention trajectory into the time-series graph neural network to obtain the corresponding first embedding vector and second embedding vector, and perform signal purification on the first embedding vector and the second embedding vector.
[0046] S52. After concatenating the purified first and second embedding vectors, input them into the counterfactual decoder, output the prediction vector, and simultaneously estimate the uncertainty covariance matrix of the prediction vector.
[0047] S53. Perform reverse decoding on the predicted vector to generate node contribution scores between targets, pathways, and phenotypes.
[0048] S54. Based on the prediction vector, uncertainty covariance matrix, and node contribution score, a counterfactual prediction package is generated as the prediction result of the regenerative response after intervention.
[0049] In this embodiment, after obtaining the simulated trajectories before and after intervention, the original time-series trajectories are curves. By extracting the most critical features from the trajectories, they are transformed into structured predictions of the regeneration response. Then, by synchronously estimating the uncertainty covariance matrix, credible values are provided to reveal which specific biological entities are driving the prediction results. Finally, a counterfactual prediction package containing predictions, confidence indices, and mechanism explanations is output, providing a comprehensive and solid basis for whether to conduct wet experiment verification.
[0050] Specifically, two time-series trajectories, one before and one after the intervention, are input into a temporal graph neural network (TMN). Each time point in the time-series trajectory is a high-dimensional state vector. The TMN simultaneously captures the temporal dynamics of each biological entity and the spatial relationships between entities, encoding the entire complex spatiotemporal trajectory into a fixed-length, low-dimensional embedding vector, namely the first embedding vector and the second embedding vector. By refining the first and second embedding vectors—namely, denoising and enhancing key signals—the model focuses more on the key stages and features in the trajectory most relevant to the regenerative phenotype, suppressing the influence of irrelevant fluctuations. The refined first and second embedding vectors are concatenated and input into a counterfactual decoder to generate the final prediction vector, while simultaneously estimating its uncertainty.
[0051] The counterfactual decoder uses a neural network to learn the mapping from the differences in trajectories before and after intervention to the final regenerative response. The generated prediction vector contains predictions across multiple dimensions, such as the percentage increase in neuron survival rate, axonal regeneration speed, and reduction in inflammation score. Uncertainty covariance matrix estimation uses techniques like Bayesian neural networks or Monte Carlo Dropout to output the mean and covariance matrix of the prediction vector. The diagonal elements of the covariance matrix represent the uncertainty (variance) of each prediction dimension, while the off-diagonal elements represent the uncertainty correlation (covariance) between different prediction dimensions. The prediction vector is then back-decoded to calculate the contribution of each node in the knowledge graph to the final prediction result. Attribution methods or attention weights are used for back-calculation. By perturbing the input of a node in the knowledge graph, the degree of change in the prediction vector is observed; the greater the change, the higher the contribution score of that node to the prediction result. All the above outputs are integrated to form the counterfactual prediction package.
[0052] As a preferred embodiment, signal purification is performed on the first embedding vector and the second embedding vector, specifically including: By comparing the loss functions, the first embedding vector is forced to move away from the second embedding vector and closer to the prior embedding of successful nerve cell regeneration, while moving away from the prior embedding of failed nerve cell regeneration.
[0053] In this embodiment, the prior embedding for successful neural cell regeneration is obtained by extracting the embedding vector from a digital twin trajectory generated from data of a known positive control that successfully promotes regeneration, using the same temporal graph neural network. This vector represents the paradigm of success. Positive controls include experiments using neurotrophic factors such as BDNF and CNTF. The prior embedding for failed neural cell regeneration is obtained by extracting the embedding vector from a trajectory generated from data of a known negative control that leads to regeneration failure. This vector represents the paradigm of failure. Negative controls include experiments applying the inflammatory factor TNF-α or some kind of inhibitor. By anchoring the first and second embedding vectors to the successful and failed priors with clear biological significance, a more accurate discrimination and correlation that conforms to the biological laws of neural cell regeneration is formed, thereby reducing overfitting to specific noise in the training data.
[0054] As a preferred embodiment, the generation of experimental feedback data specifically includes: S61. Analyze the prediction results to obtain target identification, intervention method and intervention parameters.
[0055] S62. Based on an observable biological model, experimental interventions are applied to the target through intervention parameters, wherein the experimental interventions include at least one of chemical intervention, genetic intervention, protein degradation intervention, optogenetic stimulation, or chemogenetic stimulation.
[0056] S63. Simultaneously / sequentially acquire at least one quantifiable neural cell regeneration-related signal, and generate a regeneration phenotype vector after preprocessing.
[0057] S64. Compare the regenerated phenotypic vector with the predicted vector of the digital twin model to obtain experimental feedback data.
[0058] In this embodiment, the target, intervention method, and parameters are clearly identified by analyzing the prediction results, ensuring that the experimental intervention aligns with the virtual screening direction of the digital twin and avoiding a disconnect between real experiments and virtual predictions. Based on observable biological models, diverse experimental interventions are employed to verify the effectiveness of target interventions in real biological environments. Quantitative regeneration-related signals are collected synchronously / sequentially, and phenotypic vectors are generated, transforming the abstract regeneration process into concrete and comparable numerical data, providing an objective basis for subsequent comparisons. Feedback data is obtained by comparing the regeneration phenotypic vector with the prediction vector, providing real biological data support for model parameter correction and dynamic characterization optimization. Finally, through multiple rounds of feedback iteration, the realism and reliability of the digital twin model's simulation of the neural cell regeneration process are improved, thereby enhancing the clinical translation potential of drug targets screened based on the model.
[0059] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for identifying drug targets for neural cell regeneration based on a large model, characterized in that, The method includes: S11. Acquire and fuse at least two types of multimodal biomedical data reflecting the process of nerve cell regeneration to construct a dynamic knowledge graph; S12. Based on dynamic knowledge graphs, construct a digital twin model that simulates the regeneration of target nerve cells; S13. Identify potential drug targets based on digital twin models, conduct virtual interventions, and generate prediction results of nerve cell regeneration responses after intervention; S14. Based on the prediction results, experimental interventions are applied to potential drug targets in observable biological models, and quantifiable signals related to nerve cell regeneration are collected to obtain experimental feedback data. S15. Feed the experimental feedback data back to the digital twin model for iteration.
2. The method for identifying drug targets for neural cell regeneration based on a large model according to claim 1, characterized in that, A spatiotemporal dynamic knowledge graph for the field of neural cell regeneration is constructed based on multimodal biomedical data, specifically including: S21. Acquire multimodal biomedical data, wherein the multimodal biomedical data includes at least proteomic data, metabolomic data, molecular interaction data, single-cell transcriptomic data, and spatial transcriptomic data; S22. Preprocess and align the multimodal biomedical data with entities, extract biomedical entities and their relationships, and add corresponding time dimension labels. S23. Construct an initial static knowledge graph using biomedical entities as nodes, relationships between entities as edges, and data under time dimension labels as attributes. S24. A spatiotemporal graph neural network is used to perform dynamic representation learning on the initial static knowledge graph to generate a spatiotemporal dynamic knowledge graph.
3. The method for identifying drug targets for neural cell regeneration based on a large model according to claim 1, characterized in that, The construction of a digital twin model of a target nerve cell based on a dynamic knowledge graph specifically includes: S31. Extract the biological entities and corresponding dynamic attributes related to the functional state of the target nerve cell from the spatiotemporal dynamic knowledge graph, and convert them into state vectors; S32. Based on the interaction relationships between entities in the spatiotemporal dynamic knowledge graph, construct a state transition function that controls the evolution of the state vector over time. S33. Couple the state vector with the state transition function to generate a digital twin model of the target nerve cell.
4. The method for identifying drug targets for neural cell regeneration based on a large model according to claim 1, characterized in that, Identifying potential drug targets based on digital twin models and conducting virtual interventions, specifically including: S41. When the digital twin model reaches the target time step, extract the sub-vector corresponding to the potential drug target from the global state vector to obtain a snapshot of the target state. S42. Generate an intervention predictor according to a preset intervention type, wherein the intervention predictor includes at least one of inhibition, activation, knockout or reversible blocking; S43. Apply the interference vector to the subvector to obtain the instantaneous state after target intervention, and write it back to the global state vector. S44. Using the global state vector after intervention as the initial value, call the state transition function to roll forward simulation to the end of the experiment and generate the regeneration trajectory after intervention. S45. Compare the pre-intervention trajectory with the post-intervention trajectory, calculate the counterfactual effect value, and generate a prediction result of the nerve cell regeneration response after intervention.
5. The method for identifying drug targets for neural cell regeneration based on a large model according to claim 4, characterized in that, The generation of prediction results specifically includes: S51. Input the pre-intervention trajectory and post-intervention trajectory into the time-series graph neural network to obtain the corresponding first embedding vector and second embedding vector, and perform signal purification on the first embedding vector and the second embedding vector. S52. After concatenating the purified first and second embedding vectors, input them into the counterfactual decoder, output the prediction vector, and simultaneously estimate the uncertainty covariance matrix of the prediction vector. S53. Reverse decode the predicted vector to generate node contribution scores between target points, pathways and phenotypes. S54. Based on the prediction vector, uncertainty covariance matrix, and node contribution score, a counterfactual prediction package is generated as the prediction result of the regenerative response after intervention.
6. The method for identifying drug targets for neural cell regeneration based on a large model according to claim 5, characterized in that, Signal purification is performed on the first and second embedding vectors, specifically including: By comparing the loss functions, the first embedding vector is forced to move away from the second embedding vector and closer to the prior embedding of successful nerve cell regeneration, while moving away from the prior embedding of failed nerve cell regeneration.
7. The method for identifying drug targets for neural cell regeneration based on a large model according to claim 6, characterized in that, The generation of experimental feedback data specifically includes: S61. Analyze the prediction results to obtain target identification, intervention methods, and intervention parameters; S62. Based on an observable biological model, experimental interventions are applied to the target through intervention parameters, wherein the experimental interventions include at least one of chemical intervention, genetic intervention, protein degradation intervention, optogenetic stimulation, or chemogenetic stimulation. S63. Simultaneously / sequentially acquire at least one quantifiable neural cell regeneration-related signal, and generate a regeneration phenotype vector after preprocessing; S64. Compare the regenerated phenotypic vector with the predicted vector of the digital twin model to obtain experimental feedback data.