An efficient multi-omics meta-cell inference method based on single-cell quantification
Through a neural network method based on single-cell quantification, efficient multi-omics meta-cell inference is achieved, which solves the problems of low accuracy and high time cost on large-scale data sets, supports multi-omics data analysis, and improves computational efficiency and biological understanding capabilities.
Patent Information
- Application Number
- CN202411789997.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-12-06
AI Technical Summary
Existing meta-cell inference methods have low inference accuracy, low efficiency, and high time cost on large-scale datasets, and are difficult to process multi-omics data.
A single-cell quantization-based method is adopted to use neural networks and quantization codebooks for feature extraction and quantization reconstruction to achieve efficient multi-omics meta-cell inference, including initializing the neural network encoder and decoder, and performing cell quantization and reconstruction through the quantization codebook.
It enables efficient inference on large-scale single-cell datasets, supports multi-omics data analysis, significantly reduces computational costs and improves the understanding of biological processes.
Smart Images

Figure CN119694411B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of meta-cell inference technology, and in particular to an efficient multi-omics meta-cell inference method based on single-cell quantification. Background Art
[0002] In recent years, single-cell sequencing technology has made rapid progress, enabling the capture of ever-increasing numbers of cells. It has demonstrated significant advantages in revealing cellular heterogeneity and reconstructing cell developmental trajectories. However, with the exponential growth in data size, the analysis of single-cell sequencing data also faces severe computational challenges. For example, a typical single-cell data analysis pipeline (including data integration, clustering, visualization, and differential expression analysis) takes approximately 16 hours to process 500,000 cells. However, when the number of cells increases to 600,000, insufficient memory can cause the program to crash, even on a dedicated computing platform equipped with 512 GB of memory.
[0003] To address the computational pressures posed by large-scale single-cell sequencing data, researchers have proposed a variety of efficient single-cell analysis tools, primarily for tasks such as data interpolation, integration, clustering, and cell type annotation. However, these tools are often designed for specific tasks and are difficult to directly integrate into existing single-cell data analysis frameworks. To achieve more general and efficient single-cell data processing, one solution is to compress the raw data, thereby reducing data redundancy and enabling traditional analysis tools to process large-scale data more efficiently. A representative approach to single-cell data compression is metacell inference, which aggregates biologically similar cell populations into representative metacells, effectively reducing the number of cells to be processed. Metacell inference offers significant advantages in large-scale data processing. First, the data compression provided by metacells reduces computational overhead, facilitating large-scale data processing. Second, by aggregating cells with similar characteristics, metacells mitigate data sparsity, making downstream analyses (such as cell type annotation and developmental trajectory inference) more efficient and robust. However, while metacell inference methods have shown promising results in some application scenarios, they still face shortcomings in terms of accuracy and efficiency on large-scale datasets. Specifically, the runtime required by existing metacell inference methods increases exponentially with the number of cells, limiting their application to larger datasets. Furthermore, existing metacell inference algorithms typically only support single-omics data (such as gene sequencing data) and cannot directly process multi-omics data (such as gene + protein sequencing data). Furthermore, in the field of single-cell data analysis, the applicability of such algorithms is limited by the increasing amount of multi-omics data. Therefore, how to reduce the enormous time overhead of metacell inference methods and achieve metacell inference on multi-omics data is an urgent problem to be solved in the field of single-cell analysis. Summary of the Invention
[0004] In response to the above-mentioned deficiencies in the prior art, the present invention provides an efficient multi-omics meta-cell inference method based on single-cell quantification, which is used to solve the problems of low inference accuracy, low efficiency and high time cost of existing meta-cell inference methods on large-scale data sets.
[0005] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:
[0006] An efficient multi-omics meta-cell inference method based on single-cell quantification, including the following steps:
[0007] S1. Obtain cell sequencing data and standardize it to obtain standard cell sequencing data;
[0008] The cell sequencing data includes a number of single-cell sequencing data or a number of multi-omics cell sequencing data;
[0009] S2. Initialize the quantization codebook, the neural network encoder, the first neural network decoder, and the second neural network decoder, and input the standard cell sequencing data for training. By performing feature extraction, cell quantization, and quantization reconstruction on the standard cell sequencing data, the trained neural network encoder, the first neural network decoder, the second neural network decoder, and the quantization codebook are obtained;
[0010] The quantization codebook includes several entries, each of which corresponds to the representation of a meta-cell;
[0011] S3. Input the original single-cell sequencing data or the original multi-omics cell sequencing data into the trained neural network encoder for feature extraction, and perform cell quantization based on the trained quantization codebook to obtain the meta-cell inference result.
[0012] The present invention has the following beneficial effects:
[0013] 1. This paper proposes an efficient multi-omics meta-cell inference method based on single-cell quantification, which can achieve efficient inference on large-scale single-cell datasets. Existing meta-cell inference methods have high computational overhead when processing large-scale data, which increases exponentially with the number of cells, limiting their scalability. The method proposed in this paper aims to achieve meta-cell inference with linear time complexity, effectively reducing computational cost while ensuring accuracy, enabling it to support processing millions of single-cell data.
[0014] 2. Supporting the joint analysis of multi-omics data; because existing methods are mostly limited to single-omics data, it is difficult to integrate multi-omics information and cannot fully explore the multi-dimensional characteristics of cells; the method proposed in this invention can support meta-cell inference of multi-omics single-cell data, improving the ability to understand biological processes. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a schematic diagram of the process of an efficient multi-omics meta-cell inference method based on single-cell quantification proposed in the present invention;
[0016] Figure 2 Schematic diagram of the working principles of the neural network encoder, the first neural network decoder, the second neural network decoder and the quantization codebook in the embodiment. DETAILED DESCRIPTION
[0017] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.
[0018] like Figure 1 As shown in FIG, an efficient multi-omics meta-cell inference method based on single-cell quantification includes the following steps S1-S3:
[0019] S1. Obtain and standardize cell sequencing data to obtain standard cell sequencing data; wherein the cell sequencing data includes a plurality of single-cell sequencing data or a plurality of multi-omics cell sequencing data.
[0020] In this embodiment, the purpose of normalizing the cell sequencing data is to unify the dimensions so that the same set of neural network encoders can be used for feature extraction for different cells.
[0021] Specifically, step S1 includes S11-S12:
[0022] S11. Acquire a number of cell sequencing data and generate a cell sequencing data matrix.
[0023] S12. Scale the total cell counts in each row of the cell sequencing data matrix by S times, take the natural logarithm, and normalize to obtain standard cell sequencing data.
[0024] In this embodiment, if the cell sequencing data is single-cell sequencing data, the single-cell counting matrix is input ,in is the total number of single cells, Represents the feature dimension, so the standardization of single-cell sequencing data is specifically as follows: the count of each cell (that is, the sum of each row in the single-cell count matrix) is scaled by 10,000 times, and the natural logarithm of the scaled single-cell count matrix is taken to smooth out excessive values. Finally, it is standardized to satisfy 0 mean and 1 variance to obtain standard cell sequencing data. In addition, the standardization process of multi-omics cell sequencing data is the same as that of single-cell sequencing data.
[0025] S2. Initialize the quantization codebook, the neural network encoder, the first neural network decoder, and the second neural network decoder, and input the standard cell sequencing data for training. By performing feature extraction, cell quantization, and quantization reconstruction on the standard cell sequencing data, a trained neural network encoder, a first neural network decoder, a second neural network decoder, and a quantization codebook are obtained; wherein the quantization codebook includes several entries, each entry corresponding to the representation of a meta-cell.
[0026] In this embodiment, the working principles of the quantization codebook, the neural network encoder, the first neural network decoder and the second neural network decoder are as follows: Figure 2 As shown, the neural network encoder is a fully connected neural network with a dimension of M-512-128-32; the first neural network decoder includes a first fully connected network, a first fully connected network layer, and a second fully connected network layer; the second neural network decoder includes a second fully connected network, a third fully connected network layer, and a fourth fully connected network layer; the dimensions of the first fully connected network and the second fully connected network are 32-128-512, and the dimensions of the first fully connected network layer, the second fully connected network layer, the second fully connected network, and the third fully connected network layer are 512-M; the dimension of the input layer corresponds to the feature dimension M of the input single-cell matrix (normalized count matrix), and the quantization codebook has entries, each with a dimension of 32; its working principle is: feature extraction is performed in the neural network encoder, and then the extracted low-dimensional features of the cells are input into the first neural network decoder of the upper layer for reconstruction, and at the same time, the extracted low-dimensional features of the cells are input into the quantization codebook for cell quantization. After obtaining the cell quantization representation, the cell quantization representation is input into the second neural network decoder of the lower layer for quantization reconstruction, thereby helping to optimize the quantization codebook; in addition, the training process of single-cell sequencing data is the same as that of multi-omics cell sequencing data, but because multi-omics cell sequencing data includes several omics, in the training process of multi-omics cell sequencing data, each omics corresponds to a neural network encoder, a first neural network decoder, and a second neural network decoder. The specific training process of single-cell sequencing data or multi-omics cell sequencing data is as follows:
[0027] Specifically, when the standard cell sequencing data is input for training in step S2, if the standard cell sequencing data is standard single-cell sequencing data, the standard single-cell sequencing data is input into the initialized neural network encoder, the first neural network decoder, the second neural network decoder, and the quantization codebook for training to obtain the trained neural network encoder, the first neural network decoder, the second neural network decoder, and the quantization codebook. The specific process is as follows:
[0028] Step 1: Input the standard single-cell sequencing data into the neural network encoder for feature extraction. By mapping the high-dimensional standard single-cell sequencing data features into low-dimensional representations, a low-dimensional representation of the single-cell sequencing data is obtained.
[0029] In this embodiment, a fully connected neural network encoder is used to extract a low-dimensional representation of standard single-cell sequencing data so that the low-dimensional representation can be fed into a neural network decoder for feature reconstruction in subsequent steps.
[0030] Step 2: Input the low-dimensional representation of the single-cell sequencing data into the first neural network decoder for feature reconstruction and calculate the standard reconstruction loss function; wherein, the first neural network decoder includes a first fully connected network, a first fully connected network layer, and a second fully connected network layer.
[0031] In this embodiment, during the cell reconstruction process, a negative binomial distribution is used to model the original single-cell sequencing data. The optimization objective of the neural network encoder and the first neural network decoder is to maximize the following log-likelihood (standard reconstruction loss function) to help extract a low-dimensional representation of the cell from the standard single-cell sequencing data, namely:
[0032]
[0033] in, represents the standard reconstruction loss function, represents the total number of single cells, represents the logarithmic function, represents the probability density function of the negative binomial distribution, Indicates the The original count of single cells, 、 Represent the first fully connected network layer and the second fully connected network layer respectively, represents the mean of the negative binomial distribution calculated by the first fully connected network layer, represents the variance of the negative binomial distribution calculated by the second fully connected network layer, represents the gamma function, Indicates the The scaling factor of the original count of a single cell in the normalization process, represents a diagonal matrix, represents the exponential function, Indicates the The representation of a single cell is calculated by the first fully connected network, Indicates the The low-dimensional representation of a single cell extracted by the neural network encoder, represents the first fully connected network, Represents a neural network encoder.
[0034] Step 3: Calculate the cosine similarity between the low-dimensional representation of the single-cell sequencing data and each entry in the quantization codebook and sort them, obtain the entry with the most similar low-dimensional representation of the single-cell sequencing data, assign each single cell in the single-cell sequencing data to the entry with which it is most similar, and calculate the quantization codebook loss function.
[0035] In this embodiment, this step is a cell quantization operation, that is, in the low-dimensional representation space of the neural network, the cosine similarity between the cell representation and each entry in the quantization codebook is calculated, so as to assign each cell to the entry that is most similar to it.
[0036] Specifically, the most similar entry for obtaining the low-dimensional representation of single-cell sequencing data in Step 3 is:
[0037]
[0038] in, represents the quantization codebook, Indicates the Quantitative characterization of single cells, represents the cell quantization operation, Represents the quantization codebook No. entries, It means to find the maximum value. represents cosine similarity.
[0039] In this embodiment, since each entry of the quantization codebook corresponds to the representation of a meta-cell, the total number of entries contained therein corresponds to the total number of target meta-cells given by the user.
[0040] Specifically, the quantization codebook loss function in step 3 is:
[0041]
[0042] in, represents the quantization codebook loss function, Indicates a gradient stop operation.
[0043] In this embodiment, since the quantization codebook is initialized randomly, in order to help each entry in the quantization codebook be closer to the low-dimensional representation of the cell, the quantization codebook loss function is used to optimize the quantization codebook. In addition, the gradient stop operation , indicating that the loss function is only used to optimize the quantization codebook without affecting the low-dimensional representation of the cell. This design can avoid the relatively random quantization codebook from interfering with the low-dimensional representation of the cell in the early stage of training.
[0044] Step 4: Count the number of times each entry is used to assign a single cell, adjust entries with excessive or insufficient usage, update the quantization codebook, and input each entry of the updated quantization codebook into the second neural network decoder to reconstruct the low-dimensional representation of all single-cell sequencing data assigned to the entry and calculate the quantization reconstruction loss function; wherein, the second neural network decoder includes a second fully connected network, a third fully connected network layer, and a fourth fully connected network layer.
[0045] In this embodiment, since the above-mentioned loss function only optimizes the entries assigned to cells, these optimized entries will become closer and closer to the low-dimensional representation of the cells, and thus will be used more and more. However, those entries that are less used will always correspond to fewer cells due to insufficient optimization, resulting in large differences in the number of original cells corresponding to different meta-cells. However, when a meta-cell corresponds to too many original cells, it may cover different cell types and states at the same time, resulting in some characteristics of the original cells being diluted in the meta-cell; and when a meta-cell corresponds to too few original cells, it is difficult to effectively alleviate the sparsity of the data and play a role in data compression; therefore, this step counts the number of times each entry is assigned to a single cell, adjusts the entries with too much or too little usage, and updates the quantization codebook to ensure that the original cells are relatively evenly quantized to each entry.
[0046] Specifically, in step 4, the number of times each entry is used to allocate a single cell is counted, entries with excessive or insufficient usage are adjusted, and the specific process of updating the quantization codebook is as follows:
[0047] First, count the number of times each entry allocates a single cell, that is:
[0048]
[0049] in, Indicates the The quantization codebook in the historical iteration process of the statistics No. Items The number of times it has been used, Indicates the The quantization codebook in the historical iteration process of the statistics No. Items The number of times it has been used, Indicates the The quantization codebook in the historical iteration process of the statistics No. Items The number of times it has been used, Indicates the smoothing momentum between the control history and the current usage times. Indicates the Quantization codebook for iteration No. Items The number of times it has been used.
[0050] In this embodiment, The value is 0.9.
[0051] Secondly, we adjust the entries that are underused and push them to the cell representation farthest from the quantization codebook. This converts the farthest distance between the cell representation and the quantization codebook into the probability that the low-dimensional representation of the single-cell sequencing data will be selected as the update target, i.e.:
[0052]
[0053] in, Indicates that it is used to adjust the entries that are too little used. The probability that the low-dimensional representation of a single cell is selected as the update target, Indicates taking the maximum value, Represents the first entries, Indicates the total number of entries in the quantization codebook, that is, the number of target metacells.
[0054] by The probability of Select from the low-dimensional representation of single-cell sequencing data Characterization of individual cells , and update the entries of the quantization codebook to obtain the updated quantization codebook, that is:
[0055]
[0056] in, It represents the probability that the low-dimensional representation of the first single cell is selected as the update target when adjusting the entries that are too rarely used. Indicates that it is used to adjust the entries that are too little used. The probability that the low-dimensional representation of a single cell is selected as the update target, Indicates the cell representation selected to adjust the first entry that is too little used, Indicates the selected Cellular representation of items, Indicates the Update the first iteration after the less used entry entries, Indicates the Update the first iteration after the less used entry entries, Indicates the weight of controlling the update of entries that are rarely used. Indicates the selected Cellular representation of items, Indicates a minimum value.
[0057] In this embodiment, since the weights corresponding to the less used entries are higher, they will be pushed more to the farthest cell representation of the cluster quantization codebook. Therefore, the parameter is set Controls the weight of the update of the rarely used entries. At the same time, the minimum Used to avoid excessive update weights, the value is 0.001.
[0058] Then, we adjust the overused entries and push them to the cell representation with a medium distance from them. That is, we convert the medium distance between the cell representation and the quantization codebook into the probability that the low-dimensional representation of the single-cell sequencing data is selected as the update target, that is:
[0059]
[0060] in, Indicates that it is used to adjust the entries that are used too much. The low-dimensional features of a single cell are selected as The probability of updating the target of entries, represents the median, Indicates the The median distance of the entries to the cell representation, Indicates the Low-dimensional representation of a single cell.
[0061] by The probability of Select from the low-dimensional representation of single-cell sequencing data Characterization of individual cells , and update the entries of the quantization codebook to obtain the updated quantization codebook, that is:
[0062]
[0063] in, When the low-dimensional representation of the first single cell is selected as the The probability of updating the target of entries, Indicates that it is used to adjust the entries that are used too much. The low-dimensional representation of a single cell is selected as the The probability of updating the target of entries, Indicates the cell representation selected to adjust the first item that is overused. Indicates the selected one for adjusting the overused Cellular representation of items, Indicates the In the next iteration, update the first entry after the overused entry entries, Indicates the weight of controlling the update of overused entries. Indicates the selected one used to adjust the overused Cellular representation of the items.
[0064] In this embodiment, since the weights corresponding to the more frequently used entries are higher, they will be pushed more towards the cell representations in the medium range of the distance quantization codebook. Therefore, the parameter Controls the weight of updates for overused entries.
[0065] Specifically, the quantized reconstruction loss function in Step 4 is:
[0066]
[0067] in, represents the quantized reconstruction loss function, Indicates the The quantitative representation of a single cell is calculated by the second fully connected network, 、 Respectively represent the third fully connected network layer and the fourth fully connected network layer, represents the mean of the negative binomial distribution calculated by the third fully connected network layer, represents the variance of the negative binomial distribution calculated by the fourth fully connected network layer, represents the second fully connected network.
[0068] In this embodiment, the meta-cell, as the compression result of the original cell, should fully aggregate the information in the original cell and have the common characteristics of the cell group it represents. In other words, ideally, a meta-cell should be able to effectively reconstruct several original cells it represents. To this end, in the above quantization codebook loss function In addition to the update strategy, each entry in the quantization codebook is required to reconstruct the original counts of all cells assigned to the entry, that is, the quantization reconstruction loss function is used to optimize the entry of each quantization codebook to reconstruct the original counts of all cells assigned to the entry.
[0069] Step 5: Sum the standard reconstruction loss function, the quantization codebook loss function, and the quantization reconstruction loss function to obtain the total loss function, which is:
[0070]
[0071] in, represents the total loss function, represents the standard reconstruction loss function, represents the quantization codebook loss function, represents the quantized reconstruction loss function.
[0072] In this embodiment, the total loss function is used to optimize the neural network encoder, the first neural network decoder, the second neural network decoder and the quantization codebook .
[0073] Step 6: Based on the total loss function, the gradient descent method is used to optimize the neural network encoder, the first neural network decoder, the second neural network decoder and the quantization codebook, and it is determined whether the current number of iterations reaches the maximum number of iterations. If so, the trained neural network encoder, the first neural network decoder, the second neural network decoder and the quantization codebook are obtained. Otherwise, continue to execute steps 1-5.
[0074] Specifically, when inputting standard cell sequencing data for training in step S2, if the standard sequencing data is standard multi-omics cell sequencing data, the training process is the same as the training process of standard single-cell sequencing data. Since the standard multi-omics cell sequencing data includes several omics, in the cell quantification process, the low-dimensional representation of each cell is the splicing of the representations of each omics, that is:
[0075]
[0076] in, Represents the first A low-dimensional representation of a single cell, represents the total number of omics, Represents a splicing operation, Indicates the The first low-dimensional representation of single-cell omics, Indicates the Low-dimensional representation of single-cell 2-omics, Indicates the Single cell Low-dimensional representation of omics, Indicates the Single cell Low-dimensional representation of omics, Indicates the The neural network encoder corresponding to each omics, Indicates the Single cell The count matrix after omics standardization.
[0077] In this embodiment, each omics in the standard multi-omics cell sequencing data needs to correspond to a neural network encoder.
[0078] Therefore, the total loss function for standard multi-omics sequencing data training is:
[0079]
[0080] in, represents the total loss function when training standard multi-omics sequencing data, Indicates the The standard reconstruction loss function for omics, Indicates the Quantitative reconstruction loss function for omics.
[0081] S3. Input the original single-cell sequencing data or the original multi-omics cell sequencing data into the trained neural network encoder for feature extraction, and perform cell quantization based on the trained quantization codebook to obtain the meta-cell inference result.
[0082] Specifically, step S3 includes S31-S32:
[0083] S31. Input the original single-cell sequencing data or the original multi-omics cell sequencing data into the trained neural network encoder for feature extraction, and assign each cell to the entry that is most similar to it based on the similarity between the cell representation and the trained quantization codebook entry.
[0084] S32. For each entry in the trained quantization codebook, calculate the mean of all single-cell sequencing data or original multi-omics sequencing data assigned to the entry to obtain the meta-cell inference result, that is:
[0085]
[0086] in, Indicates the The inference results of individual cells, Indicates the first The number of cells per entry, represents the first entries.
[0087] To verify the effectiveness of the proposed method for efficient multi-omics meta-cell inference based on single-cell quantification, we conducted comparative experiments on a large-scale human embryonic single-cell dataset and a multi-omics human bone marrow dataset using the proposed method and the most advanced meta-cell inference methods (including SEACell, MetaCell V2, and SuperCell methods). The results are as follows:
[0088] The meta-cell-based class-balanced classification accuracy is used to measure the effectiveness of meta-cell inference. After predicting all the original cells corresponding to the same meta-cell to the type with the largest number of cells in the meta-cell, the class-balanced classification accuracy is calculated as follows:
[0089]
[0090] in, represents the class-balanced classification accuracy, represents the total number of original single cells, Indicates the The weight of a cell is calculated based on its cell type, represents the Dirichlet function, The predicted cell types, Indicates the The actual cell type of each cell, Indicates the The actual cell type of each cell. The formula uses weighting to emphasize the importance of rare cell types.
[0091] Experiment 1: We used a large-scale human embryonic single-cell dataset containing 433,695 cells of 54 cell types. The distribution of cell numbers of each type is shown in Table 1:
[0092] Table 1. Distribution of cell numbers of various types in large-scale human embryo single-cell dataset
[0093]
[0094] 500 meta-cells are inferred from the above dataset, and the experimental results are shown in Table 2:
[0095] Table 2 Comparison results of large-scale human embryo single-cell datasets
[0096]
[0097] As can be seen from Table 2, the existing SEACell and MetaCell V2 methods will experience memory overflow on a 512GB computing server due to huge memory overhead, and are unable to process this large-scale dataset. SuperCell is currently the only method that can process this dataset, but its effect is poor. The method proposed in the present invention can not only process millions of single-cell data, but also achieves exponential performance improvement compared to SuperCell, indicating that it can effectively compress large-scale single-cell data.
[0098] Experiment 2: We used a multi-omics human bone marrow dataset containing 30,672 cells, covering both gene and protein omics, and encompassing 27 cell types. The distribution of cell numbers for each type is shown in Table 3:
[0099] Table 3 Distribution of cell numbers of various types in the multi-omics human bone marrow dataset
[0100]
[0101] 613 metacells were inferred from the above dataset (corresponding to 50 times data compression). The experimental results are shown in Table 4:
[0102] Table 4 Comparison results of multi-omics human bone marrow datasets
[0103]
[0104] Because existing methods do not support meta-cell inference from multi-omics data, for performance comparison, the data from the two omics categories were concatenated at the input level and treated as single-omics data for the existing methods. Therefore, as shown in Table 4, the proposed method achieved higher class-balanced classification accuracy than existing methods, demonstrating its ability to effectively integrate information from different omics categories to achieve accurate meta-cell inference. Furthermore, while the proposed method was validated on single-omics gene single-cell sequencing data and multi-omics gene + protein sequencing data, the proposed efficient multi-omics meta-cell inference method based on single-cell quantification is also applicable to other omics categories (such as transcriptomics). Specifically, during the extraction of low-dimensional cellular features, the proposed method used a negative binomial distribution to model single-cell data, but other distributions (such as Gaussian and Poisson distributions) can also be used for modeling.
[0105] In summary, the present invention achieves efficient multi-omics meta-cell inference through a single-cell quantization strategy. On the one hand, since both the neural network and the quantization codebook are optimized through batch data, the time overhead of the method proposed in the present invention is linear with the number of input single cells. Compared with the exponential time complexity of the existing technology, it significantly reduces the computational overhead and improves the usability of the meta-cell inference method. On the other hand, the method proposed in the present invention is the first solution that supports the inference of meta-cells from multi-omics data. At the current time when multi-omics sequencing technology is rapidly developing, it effectively fills the gap in multi-omics meta-cell inference methods.
[0106] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.
[0107] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.
Claims
1. An efficient multi-omics meta-cell inference method based on single-cell quantification, characterized by: The following steps are involved: S1. Obtain cell sequencing data and standardize it to obtain standard cell sequencing data; The cell sequencing data includes a number of single-cell sequencing data or a number of multi-omics cell sequencing data; Wherein, step S1 specifically includes: S11, obtaining a number of cell sequencing data and generating a cell sequencing data matrix; S12. Scale the total cell counts in each row of the cell sequencing data matrix by S times, take the natural logarithm, and normalize to obtain standard cell sequencing data; S2. Initialize the quantization codebook, the neural network encoder, the first neural network decoder, and the second neural network decoder, and input the standard cell sequencing data for training. By performing feature extraction, cell quantization, and quantization reconstruction on the standard cell sequencing data, the trained neural network encoder, the first neural network decoder, the second neural network decoder, and the quantization codebook are obtained; The quantization codebook includes several entries, each of which corresponds to the representation of a meta-cell; In step S2, when the standard cell sequencing data is input for training, if the standard cell sequencing data is standard single-cell sequencing data, the standard single-cell sequencing data is input into the initialized neural network encoder, the first neural network decoder, the second neural network decoder, and the quantization codebook for training to obtain the trained neural network encoder, the first neural network decoder, the second neural network decoder, and the quantization codebook. The specific process is as follows: Step 1: Input the standard single-cell sequencing data into the neural network encoder for feature extraction. The low-dimensional representation of the single-cell sequencing data is obtained by mapping the high-dimensional standard single-cell sequencing data features into low-dimensional representations. Step 2: Input the low-dimensional representation of the single-cell sequencing data into the first neural network decoder for feature reconstruction and calculate the standard reconstruction loss function; The first neural network decoder includes a first fully connected network, a first fully connected network layer, and a second fully connected network layer; Among them, the standard reconstruction loss function in step 2 is: in, represents the standard reconstruction loss function, represents the total number of single cells, represents the logarithmic function, represents the probability density function of the negative binomial distribution, Indicates the The original count of single cells, 、 Represent the first fully connected network layer and the second fully connected network layer respectively, represents the mean of the negative binomial distribution calculated by the first fully connected network layer, represents the variance of the negative binomial distribution calculated by the second fully connected network layer, represents the gamma function, Indicates the The scaling factor of the original count of a single cell in the normalization process, represents a diagonal matrix, represents the exponential function, Indicates the The representation of a single cell is calculated by the first fully connected network, Indicates the The low-dimensional representation of a single cell extracted by the neural network encoder, represents the first fully connected network, represents the neural network encoder; Step 3: Calculate the cosine similarity between the low-dimensional representation of the single-cell sequencing data and each entry in the quantization codebook and sort them, obtain the entry with the most similar low-dimensional representation of the single-cell sequencing data, assign each single cell in the single-cell sequencing data to the entry with which it is most similar, and calculate the quantization codebook loss function; Step 4: Count the number of times each entry is used to assign a single cell, adjust entries that are over- or under-used, update the quantization codebook, and input each entry of the updated quantization codebook into the second neural network decoder to reconstruct the low-dimensional representation of all single-cell sequencing data assigned to the entry and calculate the quantization reconstruction loss function. The second neural network decoder includes a second fully connected network, a third fully connected network layer, and a fourth fully connected network layer; Step 5: Sum the standard reconstruction loss function, the quantization codebook loss function, and the quantization reconstruction loss function to obtain the total loss function, which is: in, represents the total loss function, represents the standard reconstruction loss function, represents the quantization codebook loss function, represents the quantized reconstruction loss function; Step 6: Based on the total loss function, the gradient descent method is used to optimize the neural network encoder, the first neural network decoder, the second neural network decoder, and the quantization codebook, and determine whether the current number of iterations has reached the maximum number of iterations. If so, the trained neural network encoder, the first neural network decoder, the second neural network decoder, and the quantization codebook are obtained. Otherwise, continue to execute steps 1-5; S3. Input the original single-cell sequencing data or the original multi-omics cell sequencing data into the trained neural network encoder for feature extraction, and perform cell quantization based on the trained quantization codebook to obtain the meta-cell inference result.
2. The efficient multi-omics meta-cell inference method based on single-cell quantification according to claim 1, characterized in that The most similar entry to the low-dimensional representation of single-cell sequencing data obtained in Step 3 is: in, represents the quantization codebook, Indicates the Quantitative characterization of single cells, represents the cell quantization operation, Represents the quantization codebook No. entries, It means to find the maximum value. represents cosine similarity.
3. The efficient multi-omics meta-cell inference method based on single-cell quantification according to claim 2, characterized in that The quantization codebook loss function in step 3 is: in, represents the quantization codebook loss function, Indicates a gradient stop operation.
4. The efficient multi-omics meta-cell inference method based on single-cell quantification according to claim 3, characterized in that In step 4, the number of times each entry is used to allocate a single cell is counted, and entries with excessive or insufficient usage are adjusted. The specific process of updating the quantization codebook is as follows: First, count the number of times each entry allocates a single cell, that is: in, Indicates the The quantization codebook in the historical iteration process of the statistics No. Items The number of times it has been used, Indicates the The quantization codebook in the historical iteration process of the statistics No. Items The number of times it has been used, Indicates the The quantization codebook in the historical iteration process of the statistics No. Items The number of times it has been used, Indicates the smoothing momentum between the control history and the current usage times. Indicates the Quantization codebook for iteration No. Items the number of times it has been used; Secondly, we adjust the entries that are underused and push them to the cell representation farthest from the quantization codebook. This converts the farthest distance between the cell representation and the quantization codebook into the probability that the low-dimensional representation of the single-cell sequencing data will be selected as the update target, i.e.: in, Indicates that it is used to adjust the entries that are too little used. The probability that the low-dimensional representation of a single cell is selected as the update target, Indicates taking the maximum value, Represents the first entries, represents the total number of entries in the quantization codebook, i.e., the number of target metacells; by The probability of Select from the low-dimensional representation of single-cell sequencing data Characterization of individual cells , and update the entries of the quantization codebook to obtain the updated quantization codebook, that is: in, It represents the probability that the low-dimensional representation of the first single cell is selected as the update target when adjusting the entries that are too rarely used. Indicates that it is used to adjust the entries that are too little used. The probability that the low-dimensional representation of a single cell is selected as the update target, Indicates the cell representation selected to adjust the first entry that is too little used, Indicates the selected Cellular representation of items, Indicates the Update the first iteration after the less used entry entries, Indicates the Update the first iteration after the less used entry entries, Indicates the weight of controlling the update of entries that are rarely used. Indicates the selected Cellular representation of items, Indicates the minimum value; Then, we adjust the overused entries and push them to the cell representation with a medium distance from them. That is, we convert the medium distance between the cell representation and the quantization codebook into the probability that the low-dimensional representation of the single-cell sequencing data is selected as the update target, that is: in, Indicates that it is used to adjust the entries that are used too much. The low-dimensional features of a single cell are selected as The probability of updating the target of entries, represents the median, Indicates the The median distance of the entries to the cell representation, Indicates the Low-dimensional representation of single cells; by The probability of Select from the low-dimensional representation of single-cell sequencing data Characterization of individual cells , and update the entries of the quantization codebook to obtain the updated quantization codebook, that is: in, When the low-dimensional representation of the first single cell is selected as the The probability of updating the target of entries, Indicates that it is used to adjust the entries that are used too much. The low-dimensional representation of a single cell is selected as the The probability of updating the target of entries, Indicates the cell representation selected to adjust the first item that is overused. Indicates the selected one used to adjust the overused Cellular representation of items, Indicates the In the next iteration, update the first entry after the overused entry entries, Indicates the weight of controlling the update of overused entries. Indicates the selected one used to adjust the overused Cellular representation of the items.
5. The efficient multi-omics meta-cell inference method based on single-cell quantification according to claim 4, characterized in that The quantized reconstruction loss function in Step 4 is: in, represents the quantized reconstruction loss function, Indicates the The quantitative representation of a single cell is calculated by the second fully connected network, 、 Respectively represent the third fully connected network layer and the fourth fully connected network layer, represents the mean of the negative binomial distribution calculated by the third fully connected network layer, represents the variance of the negative binomial distribution calculated by the fourth fully connected network layer, represents the second fully connected network.
6. The efficient multi-omics meta-cell inference method based on single-cell quantification according to claim 5, characterized in that When inputting standard cell sequencing data for training in step S2, if the standard sequencing data is standard multi-omics cell sequencing data, the training process is the same as the training process of standard single-cell sequencing data. Since the standard multi-omics cell sequencing data includes several omics, in the cell quantification process, the low-dimensional representation of each cell is the splicing of the representations of each omics, that is: in, Represents the first A low-dimensional representation of a single cell, represents the total number of omics, Represents a splicing operation, Indicates the The first low-dimensional representation of single-cell omics, Indicates the Low-dimensional representation of single-cell 2-omics, Indicates the Single cell Low-dimensional representation of omics, Indicates the Single cell Low-dimensional representation of omics, Indicates the The neural network encoder corresponding to each omics, Indicates the Single cell The count matrix after omics standardization; Therefore, the total loss function for standard multi-omics sequencing data training is: in, represents the total loss function when training standard multi-omics sequencing data, Indicates the The standard reconstruction loss function for omics, Indicates the Quantitative reconstruction loss function for omics.
7. The efficient multi-omics meta-cell inference method based on single-cell quantification according to claim 6, characterized in that Step S3 specifically includes: S31, inputting the original single-cell sequencing data or the original multi-omics cell sequencing data into the trained neural network encoder for feature extraction, and assigning each cell to the entry that is most similar to it based on the similarity between the cell representation and the trained quantization codebook entry; S32. For each entry in the trained quantization codebook, calculate the mean of all single-cell sequencing data or original multi-omics sequencing data assigned to the entry to obtain the meta-cell inference result, that is: in, Indicates the The inference results of individual cells, Indicates the first The number of cells per entry, represents the first entries.