Single cell sequence generation system based on gene network graph representation
Through a single-cell sequence generation system based on gene network graph representation, the problems of limited generalization ability and high computational cost in different data sets are solved, and more efficient single-cell model performance and better generalization ability are achieved.
Patent Information
- Application Number
- CN202510049937.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-09
AI Technical Summary
The existing technology has limited generalization capabilities on different data sets, and dense attention networks have high requirements for computing power, and it is difficult to inject structured knowledge.
A single-cell sequence generation system based on gene network diagram representation is proposed. The single-cell gene sequence is converted into a gene network diagram through a sequence-graph converter. The iterative module learns the parameters of the edges and deduces the multi-hop relationship through the graph diffusion algorithm. Finally, the graph-sequence converter rearranges the node vector to generate a single-cell gene sequence.
Improved the performance of single-cell models, reduced computational costs, significantly improved generalization capabilities on small sample data sets, and achieved optimal results in single-cell annotation and perturbation prediction tasks.
Smart Images

Figure CN119964644A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of genetic engineering, specifically a single-cell sequence generation system based on gene network graph representation. Background Art
[0002] The remarkable success of large language models has stimulated a growing interest in their application to single-cell biology. Models like scBERT and Geneformer are expected to become versatile tools in this specialized field. However, representing cells as sentences composed of genes remains an open problem because the order of genes is interchangeable. The introduction of gene network graphs provides the relative positions of genes and provides a compact data representation that can better model the interaction relationships between genes. Summary of the invention
[0003] In view of the shortcomings of the prior art, such as the limited generalization ability of the dense attention network on different data sets, the high requirements of the computing power, and the difficulty of injecting structured knowledge into the network, the present invention proposes a single-cell sequence generation system based on the gene network graph representation, and processes the single-cell sequencing data based on the computational framework of the single-cell basic model represented by the gene network graph. By using the gene network knowledge to optimize and lightweight the representation of the single-cell basic model, various single-cell basic models are integrated as pillars and their performance is improved.
[0004] The present invention is achieved through the following technical solutions:
[0005] The present invention relates to a single-cell sequence generation system based on gene network graph representation, comprising: a sequence-graph converter, an iteration module, a graph-sequence converter and an iteration module, wherein: the sequence-graph converter converts a single-cell gene sequence into a single-cell gene network graph according to genetic and regulatory relationships; the iteration module imposes attenuation constraints and learns one-hop parameters, i.e., parameters of edges in the single-cell gene network graph, and iteratively derives multi-hop relationships through a graph diffusion algorithm; the graph-sequence converter separates all node vectors on the single-cell gene network graph and rearranges them according to the gene order in the sequence before conversion to obtain a single-cell gene sequence.
[0006] The sequence-graph converter comprises: an initialization gene network graph unit and a single-cell sequence index graph structure unit, wherein: the initialization gene network graph unit constructs a directed graph between all genes as a base graph through the gene regulatory relationships obtained from a public data set; the single-cell sequence index graph structure unit uses the type of genes in the single-cell sequence as the node type, and indexes the regulatory relationships of the genes in the single-cell sequence from the base graph to form a single-cell gene network graph.
[0007] The base map includes: intracellular co-expression patterns and gene regulation relationship network diagrams.
[0008] The single-cell gene network diagram includes:
[0009] The node vector is (batch number, feature number), where each gene corresponds to a node vector.
[0010] The intracellular co-expression pattern is generated by calculating the Pearson correlation coefficient ρ(x, y) of the expression levels between genes, and connecting each gene to the top K genes with the highest expression level above a specific threshold ρ, specifically: Among them: x, y represent the expression levels of two genes respectively, μ x ,μ y is the mean of the expression levels of these two genes, σ x ,σ y is the standard deviation of the expression levels of these two genes, Represents the mean.
[0011] The gene type K is preferably 20, and the specific threshold ρ is preferably 0.4.
[0012] The gene regulation network diagram is obtained from public datasets, including but not limited to: KEGG, RegNetwork, TRRUST, HTRIdb, EVEX, JASPAR, CHEA, TRANSFAC, MOTIFMAP and ENCODE.
[0013] The iterative module supports personalized PageRank diffusion mode and heat core diffusion mode, specifically:
[0014] When scGPT is used as the base model, the PageRank diffusion model is adopted, that is, , where: V is the value vector after mapping of the single-cell large model, V k is the value vector after the kth iteration, A is the attention matrix, and α is the transfer coefficient.
[0015] When Geneformer is used as the base model, the thermonuclear diffusion model is adopted, that is, Where: V k is the value vector at the kth iteration and t is the temperature coefficient.
[0016] The calculation method of the attention matrix A is: Among them: The function of softmax is Indicates that each feature, that is, dimension value z i Mapped to the interval [0,1], each dimension consists of C values, a total of d kdimensions, Q, K, and V are the query / key / value obtained from the encoding layer parameters and framework of the base model.
[0017] The iterative module includes: a mapping unit, an information diffusion unit and an information aggregation unit, wherein: the mapping unit uses the framework and parameters of the base model to calculate the attention matrix as a fully connected graph, and then performs sparse processing on the attention matrix according to the single-cell gene network graph; the information diffusion unit performs iterative sparse processing on the attention matrix A through the library DGL; the information aggregation unit updates the query (Query) Q according to the iterated attention matrix A 更新 =AQ.
[0018] The sparse processing means that the value in A is set to 0 where there is no edge between two nodes, thus completing the mapping of the graph structure.
[0019] The library DGL refers to: Wang, M. et al. Deep Graph Library: A Graph-Centric, Highly-Performant Package for Graph Neural Networks (2019). URL https: / / arxiv.org / abs / 1909.617 01315v2. User documentation URL: https: / / www.dgl.ai /
[0020] The rearrangement refers to: taking down all node vectors on the single-cell gene network diagram and rearranging them into a sequence form of size (number of batches, number of genes, number of features) according to the gene order in the sequence before conversion. Technical Effects
[0021] The present invention uses a highly compatible plug-in to inject structural knowledge into the existing single-cell large model, that is, each single cell is used as a graph and the genes in the cell are used as nodes, so that the regulatory relationship between genes can be better simulated, and the original dense attention mechanism can be sparsely distributed by using the graph structure to reduce the computational cost. Compared with the prior art, the present invention achieves the best results in the two basic tasks of single-cell annotation and single-cell perturbation prediction. In addition, the present invention also achieves the best results in the single-cell annotation task with samples of different data volumes, reflecting that it has more potential for generalization on small sample data sets than other base models, which is very valuable in the field of bioinformatics that lacks high-quality labeled data. The present invention is also significantly higher than the traditional transformer model in terms of time and space efficiency of the architecture. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic diagram of the system of the present invention;
[0023] Figure 2 Schematic diagram of the comparison experiment for the single-cell annotation task;
[0024] Figure 3 Schematic diagram of the single-cell perturbation prediction comparison experiment;
[0025] Figure 4 Schematic diagram of the single cell annotation task comparison experiment under different training set ratios;
[0026] Figure 5 This is a schematic diagram for comparing model efficiency;
[0027] Figure 6 Schematic diagram for selecting the diffusion step number K of the model;
[0028] Figure 7 Schematic diagram for selecting the transfer parameters α and t of the model;
[0029] Figure 8 The following is a flow chart of an embodiment. DETAILED DESCRIPTION
[0030] like Figure 6 As shown, this embodiment relates to a single cell sequence generation method based on the above system, including:
[0031] Step 1. Build a neural network: Choose any existing single-cell large language model as the base model to generate query / key / value. According to the formula described in the iteration module, use the open source DGL library and PyTorch library to build a dense attention network that replaces the base model with a graph-based attention calculation mechanism.
[0032] Step 2: Construct the dataset: Use the same dataset as scGPT and perform the same preprocessing, divide the training set and validation set into 75% of each other, and collect a series of public datasets to build a directed graph based on the regulatory direction between genes for the initialized gene network graph unit of the sequence-to-graph converter.
[0033] The test set is preferably pre-divided by scGPT.
[0034] Step 3: Construct a training objective function: The training objective of the cell annotation task is to minimize the cross entropy loss function between cell types: Where: y i is the label represented by onehot, that is, a C-dimensional vector with 1 only in the i-th dimension; C is the number of categories; p i It is a probability distribution and is also C-dimensional, representing the probability of each category. The training goal of the perturbation prediction task is to minimize the mean square error between the predicted perturbation value and the true perturbation value: Represents the mean.
[0035] Step 4: Select the model's hyperparameters based on the results of the cell annotation task in the Pancreas dataset. The final parameter selection results are as follows:
[0036] Cell type annotation was performed on 3 benchmark datasets, and perturbation prediction was performed on 2 datasets to demonstrate that the present invention improves the performance of existing single-cell large language models. Each experiment was repeated 5 times, and the average and standard deviation were reported, running on a machine equipped with two Intel Xeon Gold 6138 processors, four RTX 3090 graphics cards and 503GB RAM.
[0037] For the cell type annotation datasets, this example uses Multiple Sclerosis (MS.), Myeloid, and Pancreas - these have been split into training and test datasets, keeping the test dataset unchanged and separating 25% of the training set as a validation split. The dataset was further divided into four different parts in the few-shot study. The other two datasets, Adamson and Norman, were used for perturbation prediction and were generated from GEARS. Similar to the previous datasets, this example allocates 75% of the samples for training. All test splits contain a group of cells or perturbation types that are different from those in the training and validation sets.
[0038] For a comprehensive comparison, the accuracy and F1 score of the cell type annotation are reported simultaneously. In addition, for the prediction of the unperturbed images, the Pearson correlation and mean square error of the non-zero expressed genes are evaluated, as well as the evaluation on the top 20 differentially expressed genes (Δ). Figure 1 As shown, the comparison of the effect on the cell type annotation task, in all three data sets, significantly improves the accuracy and F1 score of the baseline. This embodiment also exceeds other effective variants, such as +Longformer. These improvements may be attributed to the use of gene structure features in this embodiment, making it more suitable for cell type annotation than previous sparse methods that mostly rely on local window features.
[0039] This example is benchmarked against standard attention in Geneformer, Performer in scBERT, and FlashAttention in scGPT: batches of 10 are processed, each batch takes 32 inputs, and the sequence length is variable. Since the performer only supports processing sequences per GPU, the performance of the performer with a batch size of 32 is inferred using the ratio of the performer and standard attention with a batch size of 1.
[0040] like Figure 3 As shown, the comparison results of perturbation prediction are shown. Since Geneformer does not support the prediction of specific perturbation values, only scGPT is benchmarked. Except for the mean square error metric, this embodiment outperforms scGPT on all datasets. This can be attributed to the application of gene network knowledge, which enables the model to capture perturbations and genetic structure patterns more accurately.
[0041] Since this embodiment can narrow the search space, it may have a more significant improvement in a few-sample environment. By using prior network knowledge to exclude certain gene interactions, this narrowing excludes gene interactions that would otherwise require a large amount of training data to discover and guides the model to a specific, decay-restricted area rather than the entire possible space. Figure 4 As shown, this assumption is consistent with the results on pancreas datasets of different scales.
[0042] like Figure 5 As shown, a comparative evaluation of resource utilization of different models is shown, where the present embodiment shows competitive performance compared to other frameworks, and in some cases better memory efficiency than other frameworks. Due to the requirement of graph diffusion at each layer, it shows lower efficiency in time efficiency compared to FlashAttention and Performer. Time utilization can be further optimized by increasing the batch size, an option that seems impossible for performers within the limited number of GPUs.
[0043] This embodiment provides a graphical representation for each individual cell, which is a feature not supported by FlashAttention, so each method has its own unique advantages. Overall, this embodiment can be flexibly used as a pre-training or fine-tuning framework for various applications.
[0044] Compared with the prior art, the present invention improves the accuracy and F1-score in single-cell annotation tasks: Assume that type C cells, Improve the Pearson correlation coefficient and mean square error between the predicted value and the true value in the single-cell perturbation prediction task. In particular, the performance of the top 20 differentially expressed genes, i.e., the 20 genes with the most significant changes after single-cell perturbation, in these two indicators was improved.
[0045] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principle and purpose of the present invention. The protection scope of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. Each implementation scheme within its scope shall be subject to the constraints of the present invention.
Claims
1. A single-cell sequence generation system based on gene network graph representation, characterized in that: include: A sequence-graph converter, an iteration module, a graph-sequence converter and an iteration module, wherein: the sequence-graph converter converts the single-cell gene sequence into a single-cell gene network graph according to the genetic and regulatory relationship; the iteration module imposes attenuation constraints and learns one-hop parameters, i.e., the parameters of the edge in the single-cell gene network graph, and iteratively derives the multi-hop relationship through the graph diffusion algorithm; the graph-sequence converter separates all node vectors on the single-cell gene network graph and rearranges them according to the gene order in the sequence before conversion to obtain the single-cell gene sequence.
2. The single-cell sequence generation system based on gene network graph representation according to claim 1 is characterized in that: The sequence-graph converter comprises: an initialization gene network graph unit and a single-cell sequence index graph structure unit, wherein: the initialization gene network graph unit constructs a directed graph between all genes as a base graph through the gene regulatory relationships obtained from a public data set; the single-cell sequence index graph structure unit uses the type of genes in the single-cell sequence as the node type, and indexes the regulatory relationships of genes in the single-cell sequence from the base graph to form a single-cell gene network graph; The base map includes: intracellular co-expression patterns and gene regulation relationship network diagram; The single-cell gene network diagram includes: The node vector is (batch number, feature number), where: each gene corresponds to a node vector; The intracellular co-expression pattern is generated by calculating the Pearson correlation coefficient ρ(x, y) of the expression levels between genes, and connecting each gene to the top K genes with the highest expression level above a specific threshold ρ, specifically: Among them: x, y represent the expression levels of two genes respectively, μ x ,μ y is the mean expression level of these two genes, σ x ,σ y is the standard deviation of the expression levels of these two genes, Represents the mean.
3. The single-cell sequence generation system based on gene network graph representation according to claim 1 is characterized in that: The iterative module supports personalized PageRank diffusion mode and heat core diffusion mode, specifically: When scGPT is used as the base model, the PageRank diffusion model is adopted, that is, , where: V is the value vector after mapping of the single-cell large model, V k is the value vector after the kth iteration, A is the attention matrix, and α is the transfer coefficient; When Geneformer is used as the base model, the thermonuclear diffusion model is adopted, that is, Where: V k is the value vector at the kth iteration and t is the temperature coefficient.
4. The single-cell sequence generation system based on gene network graph representation according to claim 3 is characterized in that: The calculation method of the attention matrix A is: Among them: the softmax function is Indicates that each feature, that is, dimension value z i Mapped to the interval [0,1], each dimension consists of C values, a total of d k dimensions, Q, K, and V are the query / key / value obtained from the encoding layer parameters and framework of the base model.
5. The single-cell sequence generation system based on gene network graph representation according to claim 1 or 3, characterized in that: The iterative module includes: a mapping unit, an information diffusion unit and an information aggregation unit, wherein: the mapping unit uses the framework and parameters of the base model to calculate the attention matrix as a fully connected graph, and then performs sparse processing on the attention matrix according to the single-cell gene network graph; the information diffusion unit performs iterative sparse processing on the attention matrix A through the library DGL; the information aggregation unit updates the query (Query) Q according to the iterated attention matrix A 更新 =AQ.
6. The single-cell sequence generation system based on gene network graph representation according to claim 5 is characterized in that: The sparse processing means that the value in A is set to 0 where there is no edge between two nodes, thus completing the mapping of the graph structure.
7. The single-cell sequence generation system based on gene network graph representation according to claim 1 is characterized in that: The rearrangement refers to: taking down all node vectors on the single-cell gene network diagram and rearranging them into a sequence form of size (number of batches, number of genes, number of features) according to the gene order in the sequence before conversion.