Single cell data analysis method, device, equipment and storage medium
By sorting single-cell data and analyzing pre-trained models, transcriptome features are constructed, solving the problem of capturing sequence patterns and long-distance dependencies in existing technologies, and improving the efficiency and accuracy of single-cell data analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-10
- Publication Date
- 2026-03-10
AI Technical Summary
Existing single-cell data analysis methods are ineffective at capturing sequence patterns and long-distance dependencies, and have high computational resource requirements, resulting in low efficiency when processing large-scale single-cell data.
By acquiring gene expression data from target cells, sorting the data, and then using a pre-trained single-cell basic model for gene expression analysis, transcriptome features are constructed for downstream analysis. A single-cell basic model based on the Mamba architecture is used to capture the complex relationship between genes and cells.
It improves the performance of downstream biological analysis tasks, enhances the ability to sort and process gene expression data, effectively captures long-distance dependencies, and reduces the demand for computing resources.
Smart Images

Figure CN121641189A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of single-cell analysis technology, and in particular to a single-cell data processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Technology
[0002] With the rapid development of single-cell sequencing technology, the field of single-cell basic model training schemes has emerged rapidly, becoming a key tool for understanding individual cell heterogeneity, cell lineage relationships, and cell state transitions. Through single-cell basic model training, massive amounts of single-cell sequencing data can be deeply mined, revealing the complex mechanisms of cells under different physiological and pathological states. Through large-scale self-supervised learning, models can acquire knowledge in relevant domains, thereby improving the performance of downstream tasks and analyses. In single-cell application scenarios, models are pre-trained on large-scale scRNA data to learn the complex relationships between genes and cells from gene expression profiles. However, due to the special characteristics and complexity of single-cell RNA sequencing data (scRNA-seq), existing generative models struggle to capture sequence patterns. Furthermore, existing technologies face high computational resource requirements and memory consumption when processing large-scale single-cell data. In addition, existing models have limited performance in capturing long-range dependencies. Therefore, there is an urgent need for a single-cell data analysis method that can effectively capture sequence patterns while also effectively capturing long-range dependencies in single-cell data. Summary of the Invention
[0003] Therefore, it is necessary to provide a single-cell data analysis method, device, computer equipment, computer-readable storage medium, and computer program product that can effectively capture sequence patterns and long-distance dependencies in single-cell data, thereby addressing the aforementioned technical problems.
[0004] In a first aspect, this application provides a single-cell data analysis method, including:
[0005] Obtain gene expression data from target cells;
[0006] The gene expression data is sorted to obtain ordered gene expression data.
[0007] Gene expression analysis is performed on the ordered gene expression data based on the pre-trained single-cell basic model to obtain transcriptome features, which are then used for downstream analysis tasks.
[0008] In one embodiment, the process of sorting the gene expression data to obtain ordered gene expression data includes:
[0009] Obtain the gene regulation map of the target cell and preprocess the gene regulation map to obtain a directed acyclic graph;
[0010] The gene expression data is filtered to obtain non-zero gene expression data, and a subgraph is constructed based on the non-zero gene expression data.
[0011] The subgraphs are sorted according to the directed acyclic graph to obtain the ordered gene expression data.
[0012] In one embodiment, the single-cell basic model includes a first prediction network and a second prediction network; wherein the first prediction network and the second prediction network are each constructed from multi-layer Mamba modules, and the first prediction network and the second prediction network share model parameters.
[0013] In one embodiment, the single-cell basic model includes a first prediction network and a second prediction network; the gene expression analysis of the ordered gene expression data based on the pre-trained single-cell basic model to obtain transcriptome features includes:
[0014] The expression values of the ordered gene expression data are analyzed based on the first prediction network to obtain the predicted expression values.
[0015] Based on the second prediction network, gene regulation relationship analysis is performed on the ordered gene expression data to obtain predicted regulatory relationships;
[0016] The transcriptome features are constructed based on the predicted expression values and the predicted regulatory relationships.
[0017] In one embodiment, the downstream analysis task is a drug response prediction task; the method further includes:
[0018] Obtain the chemical structure information of the drug to be analyzed from a pre-defined knowledge base;
[0019] Drug response prediction is performed based on the transcriptomic features and chemical structure information to obtain analytical prediction values; wherein, the analytical prediction values are used to describe the inhibitory activity of the target drug on the target cells.
[0020] In one embodiment, the pre-trained single-cell base model includes:
[0021] Obtain training samples; the training samples include single-cell sequencing data.
[0022] The single-cell sequencing data is preprocessed to obtain the first training data;
[0023] The first training data is sorted to obtain the second training data;
[0024] The single-cell basic model is trained based on the second training data.
[0025] In one embodiment, the data preprocessing of the single-cell sequencing data to obtain the first training data includes:
[0026] The single-cell sequencing data is matrix-transformed to obtain a cell gene matrix; the cell gene expression matrix includes the gene expression values of each cell.
[0027] The cell gene matrix is standardized to obtain the first training data.
[0028] Secondly, this application also provides a single-cell data analysis device, comprising:
[0029] The acquisition module is used to acquire gene expression data of the target cells.
[0030] The sorting module is used to sort the gene expression data to obtain ordered gene expression data;
[0031] The gene expression analysis module is used to perform gene expression analysis on the ordered gene expression data based on the pre-trained single-cell basic model to obtain transcriptome features, and to perform downstream analysis tasks based on the transcriptome features.
[0032] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0033] Obtain gene expression data from target cells;
[0034] The gene expression data is sorted to obtain ordered gene expression data.
[0035] Gene expression analysis is performed on the ordered gene expression data based on the pre-trained single-cell basic model to obtain transcriptome features, which are then used for downstream analysis tasks.
[0036] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0037] Obtain gene expression data from target cells;
[0038] The gene expression data is sorted to obtain ordered gene expression data.
[0039] Gene expression analysis is performed on the ordered gene expression data based on the pre-trained single-cell basic model to obtain transcriptome features, which are then used for downstream analysis tasks.
[0040] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0041] Obtain gene expression data from target cells;
[0042] The gene expression data is sorted to obtain ordered gene expression data.
[0043] Gene expression analysis is performed on the ordered gene expression data based on the pre-trained single-cell basic model to obtain transcriptome features, which are then used for downstream analysis tasks.
[0044] The aforementioned single-cell data analysis methods, devices, computer equipment, computer-readable storage media, and computer program products acquire gene expression data from target cells; sort the gene expression data to obtain ordered gene expression data; and perform gene expression analysis on the ordered gene expression data based on a pre-trained single-cell basic model to obtain transcriptome features, which are then used for downstream analysis tasks. Therefore, sorting the gene expression data effectively solves the problem of disorder in gene expression data, and performing gene expression analysis on the ordered gene expression data based on a pre-trained single-cell basic model yields transcriptome features that reflect gene expression data and the regulatory relationships between genes. Downstream analysis tasks are then performed based on these transcriptome features, thereby enhancing the performance and effectiveness of these methods in downstream biological analysis tasks. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a diagram illustrating the application environment of a single-cell data analysis method in one embodiment;
[0047] Figure 2 This is a flowchart illustrating a single-cell data analysis method in one embodiment;
[0048] Figure 3 This is a flowchart illustrating a single-cell data analysis method in one embodiment;
[0049] Figure 4 This is a schematic diagram comparing the model performance of a single-cell data analysis method in one embodiment;
[0050] Figure 5 This is a schematic diagram comparing the predicted values of various drugs in a single-cell data analysis method in one embodiment;
[0051] Figure 6 This is a flowchart illustrating a single-cell data analysis method in one embodiment;
[0052] Figure 7 This is a schematic diagram of the model training process for a single-cell data analysis method in one embodiment;
[0053] Figure 8 This is a structural block diagram of a single-cell data analysis device in one embodiment;
[0054] Figure 9 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0056] The single-cell data analysis method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. Server 104 acquires gene expression data from target cells; sorts the gene expression data to obtain ordered gene expression data; and performs gene expression analysis on the ordered gene expression data based on a pre-trained single-cell basic model to obtain transcriptome features, which are then used for downstream analysis tasks. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. Server 104 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides cloud computing services.
[0057] In one exemplary embodiment, such as Figure 2 As shown, a single-cell data analysis method is provided, which can be applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps 202 to 206. Wherein:
[0058] Step 202: Obtain gene expression data of the target cells;
[0059] The target cell can be a cancer cell or other single cell. In the downstream analysis task of drug response prediction, the target cell can be a cell line. For example, in the drug response prediction task, the target cell can be a cell line with DepMap ID ACH-000636.
[0060] In some embodiments, gene expression data of target cells can be obtained from a preset knowledge base, which is a data index library obtained by integrating drug sensitivity databases, genome databases and cell line databases; the corresponding gene expression data can be obtained from the data index library according to the index identifier corresponding to the target cell (the index may refer to the index identifier corresponding to the cell or cell line, which is not limited here).
[0061] Among them, the drug sensitivity database is used to record the sensitivity and response of cancer cells to drugs (GDSC), the genome database is used to provide gene expression data of cancer cell lines (CCLE), and the cell line database is used to provide cancer type of cell lines (TCGA). The cell line identifier can be a DepMap ID.
[0062] Step 204: Sort the gene expression data to obtain ordered gene expression data.
[0063] The sorting process can be topological sorting or other sorting methods, such as machine learning-based sorting algorithms or statistical methods. Different sorting strategies can optimize the processing effect of gene expression data and improve the accuracy and applicability of the model, and are not limited to these. Since gene expression data usually cannot be directly sorted, it is necessary to preprocess the gene expression data before sorting to obtain ordered gene expression data.
[0064] In some embodiments, sorting gene expression data to obtain ordered gene expression data includes: acquiring a gene regulation map of the target cell and preprocessing the gene regulation map to obtain a directed acyclic graph; filtering the gene expression data to obtain non-zero gene expression data and constructing a subgraph based on the non-zero gene expression data; and sorting the subgraph based on the directed acyclic graph to obtain ordered gene expression data.
[0065] In some embodiments, the gene regulation map of the target cell can be obtained by retrieving the regulatory map of the gene expressed in the target cell from the preset knowledge base. Alternatively, the regulatory map of the gene expressed in the target cell can be obtained by other means, and is not limited thereto. The preset knowledge base is a data index library obtained in advance by integrating the three databases GDSC, CCLE and TCGA.
[0066] In some embodiments, the gene regulation graph can be constructed into a directed acyclic graph (DAG) by removing all cycles from the gene regulation graph and deleting the edges pointing to nodes with the same head and tail nodes. This DAG can then be considered a large graph. ,in This represents the set of nodes containing all genes, while express All the regulatory relationships between them, thus It is a directed acyclic graph.
[0067] Among them, all loops account for about 2% of the edges in the gene regulation graph, and the directed acyclic graph can be represented by a DAG graph.
[0068] In some embodiments, gene expression data is filtered to remove genes with zero expression values, retaining only genes with non-zero expression values, resulting in non-zero gene expression data. A subgraph is then constructed based on this non-zero gene expression data. The specific representation is as follows:
[0069] For all ;
[0070] in, For subgraph, It represents a sequence of values.
[0071] In some embodiments, the subgraphs are sorted according to the directed acyclic graph to generate sorted gene sequences and corresponding expression value sequences. For each cell (cell index omitted here) This maps non-zero gene expression data to a set of nodes in a directed acyclic graph. and extract subgraphs The sorting result is represented as :
[0072]
[0073] Accordingly, the expression value sequence will be reordered as follows: ;
[0074] The reordered expression value sequences and subgraphs are used as ordered gene expression data.
[0075] Step 206: Perform gene expression analysis on ordered gene expression data based on the pre-trained single-cell basic model to obtain transcriptome features, and then perform downstream analysis tasks based on the transcriptome features.
[0076] Among them, the single-cell basic model is obtained through large-scale pre-training using the Mamba architecture and a pre-training scheme that can convey biological meaning, in order to capture richer and more complex relationships between genes and cells and improve performance in various single-cell analysis tasks.
[0077] In some embodiments, the single-cell basic model includes a first prediction network and a second prediction network; wherein the first prediction network and the second prediction network are each constructed by a multi-layer Mamba module, and the first prediction network and the second prediction network share model parameters.
[0078] In some embodiments, it should be noted that the sharing of model parameters between the first prediction network and the second prediction network means that the first prediction network and the second prediction network share the same model parameters during the pre-training stage. This design can reduce the number of parameters in the single-cell basic model and improve the training efficiency of the single-cell basic model during the pre-training stage.
[0079] In some embodiments, gene expression analysis can be performed on ordered gene expression data using 12 layers of Mamba blocks in a single-cell basic model to obtain transcriptome features that reflect gene expression values and regulatory relationships between genes in single-cell sequencing data.
[0080] In some embodiments, the single-cell basic model includes a first prediction network and a second prediction network; gene expression analysis of ordered gene expression data based on the pre-trained single-cell basic model to obtain transcriptome features includes: analyzing the expression values of ordered gene expression data based on the first prediction network to obtain predicted expression values; analyzing the gene regulatory relationships of ordered gene expression data based on the second prediction network to obtain predicted regulatory relationships; and constructing transcriptome features based on the predicted expression values and predicted regulatory relationships.
[0081] The first prediction network can be a gene expression prediction model, and the second prediction network can be a gene regulation relationship prediction model.
[0082] In some embodiments, the first prediction network first transforms ordered gene expression data through an embedding layer, then extracts features through a Mamba module, and finally predicts expression values for each gene.
[0083] In some embodiments, the second prediction network performs gene regulatory role analysis on the gene regulatory network of ordered gene expression data to predict the role of each gene in the gene regulatory network, uses the Mamba module for feature extraction, and finally infers the predicted regulatory relationship through sequence prediction.
[0084] In some embodiments, the downstream analysis task is a drug response prediction task; the method further includes: obtaining the chemical structure information of the drug to be analyzed from a preset knowledge base; performing drug response prediction based on transcriptome features and chemical structure information to obtain an analysis prediction value; wherein the analysis prediction value is used to describe the inhibitory activity of the target drug on the target cells.
[0085] The drug to be analyzed can be any known existing drug or a chemical substance that has not been used as a drug. The drug to be analyzed can use PubChem ID as a unique identifier. The target cell in the drug response prediction task can be the target cell line.
[0086] In some embodiments, please refer to Figure 3 , Figure 3 The document specifically demonstrates the workflow for predictive analysis of whether an inhibitory activity relationship exists between the drug to be analyzed and the target cells in practical applications.
[0087] Specifically, by integrating the GDSC, CCLE, and TCGA databases, a data index library—a pre-defined knowledge base—can be obtained that can quickly index information from these three databases. The drug response model deepCDR trains itself by integrating corresponding training data from this pre-defined knowledge base. For example, the integration result for the drug Fauldiscipla and cell line ACH-000636 is: ('ACH-000636', '84691', 1.991343, 'ALL'). From left to right, the results represent the DepMap ID of the target cell line, the PubChem ID of the drug, the natural logarithm of the drug-cell IC50, and the cancer type of the target cell line. Gene expression data of each cell in the target cell line can be indexed based on the cell line identifier (DepMap ID), and the chemical structure information of the drug can be indexed based on the drug identifier (PubChem ID).
[0088] In this embodiment, since the drug response model input includes two parts—the first part being the cell's gene expression data and the second part being the chemical structure information of the drug to be analyzed—the original single-cell branch is replaced by assembling the single-cell basic model with the drug response model. Specifically, the transcriptome features output by the single-cell basic model replace the first part of the input in the original drug response model (e.g., ...). Figure 3 As shown in the figure, the inhibitory effect of the drug to be analyzed on the target cells can be predicted by using transcriptome features and chemical structure information of the drug.
[0089] like Figure 3 In the deepCDR structure of the drug response model shown, the chemical structure information of the drug to be analyzed is extracted using a graph neural network to obtain the drug molecule vector. Gene expression analysis is performed on the ordered gene expression data of the target cell using a single-cell basic model to obtain transcriptome features. Based on the drug molecule vector and transcriptome features, the analysis prediction value is output.
[0090] Among them, the drug molecule vector can be a drug embedding, the transcriptome feature can be a single-cell vector (Bulkembedding), and the analysis prediction value can be the IC50 value of the cell pair of the drug to be analyzed and the target cell.
[0091] It should be noted that the deepCDR drug response model is a predictive model used to predict the sensitivity of target cell lines to drugs. For example, for the target drug Fauldiscipla and the target cell line ACH-000636, the predictive model predicted an analytical value of 1.994977 (the actual value was 1.991343), where the analytical prediction value is the logarithm of IC50.
[0092] In some embodiments, Figure 4 The pre-trained single-cell baseline model in this case was applied to the drug response prediction model, replacing the cell gene expression input used in the drug response prediction model. The Pearson correlation coefficient scatter plots of the scMamba architecture in the single-cell baseline model compared to the original drug response prediction models deepCDR and scFoundation show that, for most drugs... Figure 4 The Pearson correlation coefficient between the predicted values of the target analysis objects and the actual IC50 values of the target cell lines is high, meaning that most of the points of the target analysis objects in the scatter plot are located above y=x. This indicates that the prediction performance of the pre-trained scMamba architecture single-cell basic model used in this case is better than other models.
[0093] In some embodiments, Figure 5 The violin plot is used to show the distribution of predicted values (IC50 values) under different drugs. Figure 5 The image shows the distribution of sensitivity (sens_resi) to several drugs, including the five drugs with the lowest and highest IC50 values. The bottom right corner shows the five drugs predicted to have the lowest IC50, corresponding to... Figure 5 The five violin pictures in the bottom right corner. Figure 5 The top left corner shows the top 5 drugs with the highest predicted IC50. Figure 5 The five violin images in the upper left corner are explained in detail below:
[0094] The X-axis (horizontal axis) displays different drugs, each represented by a violin plot. Drug identifiers are displayed on the X-axis, such as "387447" and "448013". The Y-axis (vertical axis) represents the IC50 value after negative logarithmic transformation (-lnIC50), a commonly used indicator for analyzing drug efficacy. IC50 is the half-maximal concentration at which a drug inhibits cell proliferation; a lower value indicates higher drug efficacy. Negative logarithmic transformation is typically used to better visualize the data and avoid values that are too large or too small. The width of the violin plot indicates the density of the data distribution. Wider plots indicate more concentrated data points, while narrower plots indicate sparser data points. Each violin plot also includes a box plot showing the median, interquartile range (IQR), and possible outliers.
[0095] Figure 5 In the middle, the bottom right corner shows low-efficacy drugs, which have lower drug potency, corresponding to... Figure 5 The five violin images in the bottom right corner. Figure 5 The top left corner displays high-efficacy drugs, which have a higher potency. Figure 5 The top left corner shows five violin pictures.
[0096] from Figure 5 As can be seen from the graph, the drugs with the lowest predicted IC50 values (highest efficacy) are mainly concentrated on the right side of the graph, such as "9910224" and "85668777". These drugs have relatively high negative logarithmic IC50 values, indicating that they can effectively inhibit cell proliferation at low concentrations. The drugs with the highest predicted IC50 values (lowest efficacy) are concentrated on the left side of the graph, such as "387447" and "448013". Their low negative logarithmic IC50 values indicate that higher concentrations are required to achieve the same inhibitory effect. The distribution of highly effective drugs is relatively more concentrated, while the distribution of less effective drugs is more dispersed.
[0097] In the aforementioned single-cell data analysis method, gene expression data of target cells is acquired; this data is then sorted to obtain ordered gene expression data; and gene expression analysis is performed on this ordered gene expression data using a pre-trained single-cell basic model to obtain transcriptomic features, which are then used for downstream analysis tasks. Therefore, sorting the gene expression data effectively solves the problem of disorder in gene expression data. Furthermore, analyzing the ordered gene expression data using a pre-trained single-cell basic model yields transcriptomic features that reflect gene expression data and the regulatory relationships between genes. Downstream analysis tasks are then performed based on these transcriptomic features, thereby enhancing the method's performance in downstream biological analysis tasks.
[0098] In one exemplary embodiment, such as Figure 6 As shown, the single-cell data analysis method also includes a pre-training single-cell basic model step, which includes steps 602 to 608. Wherein:
[0099] Step 602: Obtain training samples; training samples include single-cell sequencing data.
[0100] Among them, single-cell sequencing data is single-cell RNA sequencing data (scRNA-seq).
[0101] In some embodiments, training samples can be obtained from a database or through preparation methods, and are not limited thereto.
[0102] Step 604: Perform data preprocessing on the single-cell sequencing data to obtain the first training data;
[0103] In some embodiments, data preprocessing is performed on single-cell sequencing data to obtain first training data, including: performing matrix transformation on the single-cell sequencing data to obtain a cell gene matrix; the cell gene expression matrix includes the gene expression values of each cell; and standardizing the cell gene matrix to obtain the first training data.
[0104] Matrix transformation is a preprocessing technique. It should be noted that single-cell sequencing data is typically preprocessed into a cell-gene matrix. Each element Indicated in cells Zhonggen Gene expression values.
[0105] In some embodiments, the standardization process for the cell gene matrix is achieved through binning, where the expression value of each gene in the cell gene matrix is binned to... consecutive intervals Among them To facilitate operation, the same binning pipeline can be used. Then, each element x_{ij} will be reformatted as a binned value:
[0106]
[0107]
[0108] It should be noted that the binning values obtained after reformatting can be used to classify cells. The expression values are considered as a sequence, and this sequence can be used as the expression value sequence of the first training data: Accordingly, cells Gene expression These can be represented as sequences of the same length, and this sequence can be used as the gene expression sequence for the first training data: .
[0109] Step 606: Sort the first training data to obtain the second training data.
[0110] The second training data consists of ordered gene expression data, which can be sorted using topological sorting or other sorting methods, and is not limited to these.
[0111] In some embodiments, a gene regulation map of the training samples is obtained, and the gene regulation map is preprocessed to obtain a training directed acyclic graph; the first training data is filtered to obtain non-zero training data, and a training subgraph is constructed based on the non-zero training data; the training subgraph is sorted according to the training directed acyclic graph to obtain ordered second training data.
[0112] Step 608: Train the single-cell basic model based on the second training data.
[0113] In some embodiments, the training of the single-cell basic model requires training the original single-cell basic model based on the second training data to obtain the single-cell basic model.
[0114] The original single-cell basic model includes a first original prediction network and a second original prediction network. The first original prediction network can be an original expression value prediction model, which is an untrained model containing an embedding layer and a Mamba module. The second original prediction network can be an original gene regulation relationship prediction model, which is an untrained model containing an embedding layer and a Mamba module.
[0115] In some embodiments, the original expression value prediction model uses a "mask value prediction" strategy to train the model by randomly masking a portion of the second training data, thereby obtaining the gene expression prediction model and enhancing the generalization ability of the gene expression prediction model.
[0116] In some embodiments, please refer to Figure 7 The pre-training process of the original expression value prediction model corresponds to Figure 7The pre-training scheme 1 in the model embeds the expression value sequence in the second training data as the value of the embedding layer, and obtains the mask value by randomly masking part of the expression value sequence. The mask value is trained, and the weights of the gene expression sequence (i.e. word embedding) of the second training data obtained through pre-training are frozen using the Frozen module to prevent the weights corresponding to these pre-trained embedding layer data from changing in subsequent training. Then, the Mamba module is used to extract features and complete the mask value prediction. Finally, the prediction loss is calculated and the parameters of the Mamba module are tuned to obtain the gene expression prediction model.
[0117] It should be noted that a shared module is used during the Mamba module training phase. The shared module indicates that the first prediction network and the second prediction network share the same model parameters. Since the training of the original expression value prediction model and the original gene regulation relationship prediction model are carried out simultaneously, sharing the Mamba module parameters of the two can effectively reduce the number of model parameters and improve the training efficiency of the model.
[0118] In some embodiments, please refer to Figure 7 The pre-training process of the original gene regulation relationship prediction model corresponds to Figure 7 Pre-training scheme 2 in the model uses gene expression sequences from the second training data as word embeddings in the embedding layer. After training the gene expression sequences, features are extracted from them using the Mamba module. Finally, in order to infer regulatory relationships through sequence prediction, the training objective of the original gene regulation relationship prediction model is set to predict the next gene. The loss value of the prediction result is calculated to fine-tune the parameters of the Mamba module, thereby obtaining the gene regulation relationship prediction model. By predicting the regulatory relationships between genes, the gene regulation relationship prediction model can understand the structure and function of the gene regulation network.
[0119] It should be further explained that the single-cell basic model includes a pre-trained gene expression prediction model and a gene regulation relationship prediction model. The deployment of the single-cell basic model incorporates the Mamba architecture, a novel state-space model capable of simultaneously addressing the problems of excessively long or sparse sequences. Unlike traditional transformer models, Mamba processes single-cell sequence data using a state-space model.
[0120] The state-space model is SSM, while traditional transformer models can be scGPT or scBert.
[0121] It should be noted that the SSM model processes information in the following ways:
[0122] 1. State update: through evolution parameters and projection parameters , input sequence Mapping to hidden state .
[0123] 2. Output generation: via projection parameters Hidden state Mapping to output sequence .
[0124] The specific mathematical representation is as follows:
[0125]
[0126] ;
[0127] To accommodate the sequence-based nature of single-cell sequencing data, the SSM model employs zero-order hold for discretization, enabling it to operate effectively in the discrete environment of single-cell data. The discretized state update and output generation process is as follows:
[0128]
[0129] ;
[0130] in, and From and It comes from discretization.
[0131] In some embodiments, a single-cell base model with an scMamba architecture is trained on large-scale single-cell sequencing data for fine-tuning in other downstream tasks. Through large-scale pre-training, the model can capture richer and more complex relationships between genes and cells, improving performance in various biological analysis tasks and effectively addressing the computational efficiency, long-distance dependency capture, and sequence integrity issues faced by existing technologies in processing single-cell RNA sequencing data.
[0132] In some embodiments, it should be noted that, Figure 7Median embedding represents Value Embedding, Token embedding represents Token Embedding, Pretrain Scheme 1 represents Pretrain Scheme 1, Expression value prediction represents Expression value prediction, Pretrain Scheme 2 represents Pretrain Scheme 2, Regulatory role prediction represents Regulatory role prediction, Mamba Block represents Mamba module, MLP represents Multilayer Perceptron, Conv represents Convolutional Layer, and SSM represents State Space Model.
[0133] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0134] Based on the same inventive concept, this application also provides a single-cell data analysis device for implementing the single-cell data analysis method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the single-cell data analysis device provided below can be found in the limitations of the single-cell data analysis method described above, and will not be repeated here.
[0135] In one exemplary embodiment, such as Figure 8 As shown, a single-cell data analysis device is provided, including: an acquisition module 801, a sorting module 802, and a gene expression analysis module 803, wherein:
[0136] Acquisition module 801 is used to acquire gene expression data of target cells;
[0137] The sorting module 802 is used to sort the gene expression data to obtain ordered gene expression data;
[0138] The gene expression analysis module 803 is used to perform gene expression analysis on ordered gene expression data based on a pre-trained single-cell basic model to obtain transcriptome features, which are then used for downstream analysis tasks.
[0139] In some embodiments, the sorting module 802 is further configured to acquire the gene regulation map of the target cell, preprocess the gene regulation map to obtain a directed acyclic graph; filter the gene expression data to obtain non-zero gene expression data, and construct a subgraph based on the non-zero gene expression data; and sort the subgraph based on the directed acyclic graph to obtain ordered gene expression data.
[0140] In some embodiments, the single-cell basic model includes a first prediction network and a second prediction network; wherein the first prediction network and the second prediction network are each constructed by a multi-layer Mamba module, and the first prediction network and the second prediction network share model parameters.
[0141] In some embodiments, the single-cell basic model includes a first prediction network and a second prediction network; the gene expression analysis module 803 is further configured to perform expression value analysis on ordered gene expression data according to the first prediction network to obtain predicted expression values; perform gene regulatory relationship analysis on ordered gene expression data according to the second prediction network to obtain predicted regulatory relationships; and construct transcriptome features based on predicted expression values and predicted regulatory relationships.
[0142] In some embodiments, the downstream analysis task is a drug response prediction task; the single-cell data analysis device further includes: a drug response prediction module, used to obtain the chemical structure information of the drug to be analyzed from a preset knowledge base; to perform drug response prediction based on transcriptome features and chemical structure information, and to obtain an analysis prediction value; wherein the analysis prediction value is used to describe the inhibitory activity of the target drug on the target cell.
[0143] In some embodiments, the single-cell data analysis device further includes: a model pre-training module for acquiring training samples; the training samples include single-cell sequencing data; performing data preprocessing on the single-cell sequencing data to obtain first training data; sorting the first training data to obtain second training data; and training a single-cell basic model based on the second training data.
[0144] In some embodiments, the model pre-training module is further used to perform matrix transformation on single-cell sequencing data to obtain a cell gene matrix; the cell gene expression matrix includes the gene expression values of each cell; and the cell gene matrix is standardized to obtain the first training data.
[0145] In the aforementioned single-cell data analysis device, gene expression data of target cells is acquired; the gene expression data is sorted to obtain ordered gene expression data; gene expression analysis is performed on the ordered gene expression data based on a pre-trained single-cell basic model to obtain transcriptomic features, which are then used for downstream analysis tasks. Therefore, sorting the gene expression data effectively solves the problem of disorder in gene expression data, and performing gene expression analysis on the ordered gene expression data based on a pre-trained single-cell basic model yields transcriptomic features that reflect gene expression data and the regulatory relationships between genes. Downstream analysis tasks are then performed based on these transcriptomic features, thereby enhancing the device's performance in downstream biological analysis tasks.
[0146] Each module in the aforementioned single-cell data analysis device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0147] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores single-cell sequencing data of the target single cell. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a single-cell data analysis method.
[0148] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0149] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the single-cell data analysis method described above.
[0150] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the single-cell data analysis method described above.
[0151] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the single-cell data analysis method described above.
[0152] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0153] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0154] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0155] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A single-cell data analysis method, characterized in that, The method comprises: obtaining gene expression data of a target cell; sorting the gene expression data to obtain ordered gene expression data; performing gene expression analysis on the ordered gene expression data based on a pre-trained single-cell base model to obtain transcriptome features for downstream analysis based on the transcriptome features.
2. The method of claim 1, wherein, The sorting of the gene expression data to obtain ordered gene expression data comprises: obtaining a gene regulatory graph of the target cell and preprocessing the gene regulatory graph to obtain a directed acyclic graph; filtering the gene expression data to obtain non-zero gene expression data and constructing a subgraph based on the non-zero gene expression data; sorting the subgraph based on the directed acyclic graph to obtain the ordered gene expression data.
3. The method of claim 1, wherein, The single-cell base model comprises a first prediction network and a second prediction network; wherein the first prediction network and the second prediction network are respectively constructed by a plurality of layers of mamba modules, and the first prediction network and the second prediction network share model parameters.
4. The method of claim 1, wherein, The single-cell base model comprises a first prediction network and a second prediction network; and the gene expression analysis on the ordered gene expression data based on the pre-trained single-cell base model to obtain transcriptome features comprises: performing expression value analysis on the ordered gene expression data based on the first prediction network to obtain predicted expression values; performing gene regulation relationship analysis on the ordered gene expression data based on the second prediction network to obtain predicted regulation relationships; constructing the transcriptome features based on the predicted expression values and the predicted regulation relationships.
5. The method of claim 1, wherein, The downstream analysis task is a drug response prediction task; and the method further comprises: obtaining chemical structure information of a drug to be analyzed from a preset knowledge base; performing drug response prediction based on the transcriptome features and the chemical structure information to obtain analysis prediction values; wherein the analysis prediction values are used to describe the inhibitory activity of the target drug on the target cell.
6. The method of claim 1, wherein, The pre-trained single-cell base model comprises: obtaining training samples; the training samples comprise single-cell sequencing data; performing data preprocessing on the single-cell sequencing data to obtain first training data; performing sorting on the first training data to obtain second training data; training the single-cell base model based on the second training data.
7. The method of claim 6, wherein, The data preprocessing on the single-cell sequencing data to obtain first training data comprises: performing matrix conversion on the single-cell sequencing data to obtain a cell gene matrix; the cell gene expression matrix comprises gene expression values of each cell; performing standardization processing on the cell gene matrix to obtain the first training data.
8. A single cell data analysis apparatus, characterized by, The device comprises: an obtaining module configured to obtain gene expression data of a target cell; a sorting module configured to sort the gene expression data to obtain ordered gene expression data; a gene expression analysis module configured to perform gene expression analysis on the ordered gene expression data based on a pre-trained single-cell base model to obtain transcriptome features for downstream analysis based on the transcriptome features. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.