Multi-organ single-cell sequencing data integration method
By combining particle swarm optimization algorithm and nearest neighbor algorithm, the problem of batch effect in multi-organ single-cell RNA sequencing data is solved, more accurate data integration and cross-organ cell type analysis are achieved, and the recognition ability of rare cell types is improved.
Patent Information
- Application Number
- CN202510347236.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
When integrating multi-organ data, the existing single-cell RNA sequencing technology ignores the differences in cell population composition, resulting in serious batch effect and ignores single-cell clustering label information, affecting the accuracy of data analysis.
The particle swarm optimization algorithm is used to combine the nearest neighbor algorithm, and the batch effect is corrected by constructing the coincidence degree matrix and the particle swarm algorithm, and data dimensionality reduction and clustering are performed by combining PCA and Leiden methods, and highly variable genes are selected for data integration.
It effectively reduces the batch effect, improves the integration quality of multi-organ single-cell sequencing data, reveals the similarities and differences of cross-organ cell types and functions, and enhances the detection ability of rare cell types.
Smart Images

Figure CN120299516A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics technology, and particularly relates to a method for integrating multi-organ single-cell sequencing data. Background Art
[0002] The emergence of single-cell RNA sequencing (scRNA-seq) technology has made it possible to discover new cell types, understand dynamic biological processes, and spatially reconstruct tissues. The continuous development of scRNA-seq technology, the continuous reduction of experimental prices, and the continuous emergence of large-scale experiments have provided unprecedented opportunities for biological insights. Among many single-cell sequencing platforms, 10X Genomics is currently the most popular platform. It uses droplet-based microfluidics to classify cells, enabling true single-cell sequencing, ensuring ultra-high cell throughput and cell capture efficiency, and having advantages such as high cost performance, short time cycle, and no cell type limitation. Currently, the joint analysis of multi-organ data has become a hot topic in single-cell transcriptomics research. However, new problems have emerged. Single-cell data is usually compiled from multiple experiments, with different capture times, processors, reagent batches, and devices. These differences lead to huge variations in data or batch effects and may confound real biological changes.
[0003] Currently, all batch effect correction methods based on linear regression have a prerequisite assumption that the composition of cell populations in each batch of cells is the same. Any systematic differences in average gene expression between batches are attributed to technical differences that can be regressed. However, in practice, in scRNA-seq studies, the population compositions of different batches are usually not the same. Even assuming the same cell types exist in each batch of cells, the abundance of each cell type in the dataset will vary according to subtle differences in aspects such as cell culture or tissue extraction, separation, and sorting. In addition, current methods ignore single-cell clustering label information, but such information can improve the effectiveness of batch correction, especially in the realistic situation where biological differences and batch effects are not orthogonal.
[0004] Therefore, aiming at the non-linear realistic situation, that is, the biological premise that the composition of cell populations in each batch of cells is different, this project will explore the biological and technical factors causing batch effects through a particle swarm optimization algorithm to avoid overcorrection. Summary of the Invention
[0005] Single-cell RNA sequencing (scRNA-seq) technology has developed rapidly in the past few years and has made important contributions to the identification and characterization of cell types, states, and functions. The scRNA-seq technology can simultaneously detect the transcriptional states of multiple cells in a single experiment, so it has important significance in biological research. The scRNA-seq technology can measure gene expression within the transcriptome of a single cell, which is of great significance for identifying cell cluster types, describing cell heterogeneity in complex diseases, studying transcriptional dynamics, and studying the relationships between tissue composition or gene networks.
[0006] The present invention proposes a method for integrating multi-organ single-cell sequencing data. Aiming at the problem of batch effects in single-cell RNA sequence data, the particle swarm optimization algorithm and the nearest neighbor algorithm are combined to effectively reduce the batch effect phenomenon. The invention flow chart is as Figure 1 shown.
[0007] The technical solution of the present invention is: a method for integrating multi-organ single-cell sequencing data based on the particle swarm optimization algorithm, and the steps of the method are as follows:
[0008] Step1 Data preprocessing, including two parts: data collection and data cleaning, and performing normalization, scaling, and feature selection operations on each single-cell sequencing dataset;
[0009] Step2 Data dimensionality reduction and data clustering. For each organ dataset, use the PCA and Leiden methods in the scanpy package to select the first 40 principal components for dimensionality reduction and clustering;
[0010] Step3 Construct a coincidence matrix. For each organ, construct a coincidence matrix of itself and other organs, and input the coincidence matrix into the particle swarm algorithm;
[0011] Stpe4 Combine the particle swarm algorithm to correct the batch effect of data from multiple organs. The particle swarm algorithm flow chart is as Figure 2 shown.
[0012] The present invention uses the ARI value as a performance index, and the calculation formula of ARI is as follows:
[0013]
[0014] Where:
[0015]
[0016] TP and TN represent true positives and false positives, and FP and FN represent false positives and false negatives.
[0017] Further, the specific content of the step S1 is:
[0018] S1-1. In the data collection stage, mouse cell data from 12 organs are used. All the mouse cell data are from the dataset provided by Tabula Muris, and a single-cell sequencing dataset is constructed using the mouse data of 12 organs.
[0019] S1-2. In the data preprocessing stage, each single-cell sequencing dataset is successively normalized, scaled, and feature selected. Then, the expression values of each gene in all cells of each dataset are adjusted to make each gene have unit variance, and the top 2000 highly variable genes are selected. When integrating and analyzing datasets of three or more organs, highly variable genes in each organ dataset need to be selected.
[0020] Further, step S2 is specifically as follows:
[0021] S2-1. In the data dimensionality reduction stage, first, PCA dimensionality reduction is performed on each single-cell sequencing dataset, where the scaled data with highly variable genes is used. The number of PCs is a parameterized value, and by default, n_pcs = 15 is used.
[0022] S2-2. In the data clustering stage, for each single-cell sequencing dataset, the clustering method of the sc.pp.neighbors method in scanpy is used to obtain neighbor node scores, and then the sc.tl.leiden method is used to calculate the clustering results of each single-cell sequencing dataset.
[0023] Further, step S3 is specifically as follows:
[0024] S3-1. In the stage of constructing the coincidence degree matrix, calculate the proportion of each organ in the data to be integrated. Take the organ with the largest proportion as the benchmark, and calculate the remaining organs respectively and calculate the overlap degree of highly expressed phenotype genes between other organs and this organ respectively.
[0025] Further, step S4 is specifically as follows:
[0026] S4-1. In the stage of combining the particle swarm algorithm, the processed single-cell transcriptome sequencing data is combined with the particle swarm algorithm to achieve the batch effect correction function. Select the Pearson correlation coefficient as the fitness function of the particle swarm algorithm, use the first 40 principal components as the dimensions of the particle swarm algorithm, and set the coincidence degree matrix as the search range of the particle swarm algorithm.
[0027] The multi-organ single-cell sequencing data integration method provided by the present invention has the following beneficial effects:
[0028] Through the present invention, the similarities and differences in cell types and functions across organs are revealed. Single-cell sequencing data from different organs are effectively integrated. By integrating single-cell sequencing data from multiple organs, the characteristics of the same cell type in different organs can be systematically analyzed, the heterogeneity of these cells in different tissue microenvironments can be identified, and further, the functional changes of cells under different physiological conditions or disease states can be revealed. Secondly, the detection ability of rare cell types is improved. In the data of a single organ, some cell types may be difficult to detect due to their scarcity. The integration of multi-organ data can expand the sample size, improve the recognition rate of rare cell types, and reveal their potential roles in different organs, thus promoting in-depth research on rare cell populations. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 This is a flowchart of the present invention.
[0030] Figure 2 This is a flowchart of the particle swarm optimization algorithm. DETAILED DESCRIPTION OF THE INVENTION
[0031] 1. Data preprocessing
[0032] 1.1 Data collection
[0033] The mouse cell data from 12 organs used in this experiment are all from the dataset provided by Tabula Muris (http: / / tabula-muris.ds.czbiohub.org / ). This dataset also provides a large amount of unannotated cell data, which cannot be used in this experiment. Therefore, we selected 55,656 annotated cells for analysis.
[0034] 1.2 Data preprocessing
[0035] For all scRNA-seq datasets, we performed normalization, scaling, and feature selection. More specifically, for each dataset, we used the gene expression matrix X, where X ij is the unique molecular identifier (UMI, 10X) of gene i detected in cell j, and used logarithmic normalization, which is the default normalization method in Scanpy. Then, we adjusted the expression values of each gene in all cells in each dataset so that each gene has a unit variance. For each dataset, we implemented the sc.pp.highly_variable_genes function in Scanpy to select the top 2,000 highly variable genes. For the integrated analysis of more than two organ datasets, we selected the highly variable genes in each organ.
[0036] 2. Dimensionality reduction and clustering
[0037] We first performed PCA on each dataset, using scaled data with highly variable genes. The number of PCs is a parameter (we default to n_pcs = 15). Then, for each dataset, we used the clustering method of the sc.pp.neighbors function in Scanpy to cluster cells based on their PC scores. Then, we used sc.tl.leiden to cluster each organ, which is the default clustering method in Scanpy.
[0038] 3. Construct the coincidence matrix
[0039] The difference from traditional methods is that if the integration of multi-organ data fails to correctly account for the biological differences between different organs, it may obscure or confound signals specific to a particular organ, making it difficult to accurately identify organ-specific gene expression or biological processes. We calculate the proportion of each organ in the data to be integrated, take the organ with the largest proportion as the benchmark, calculate the remaining organs respectively, and calculate the overlap degree of high-phenotype genes between other organs and this organ. We believe that the higher the overlap degree, the stronger the shared signals between organs, and the more similar their gene expression matrix values should be. Organs with a lower overlap degree show significant biological differences and should exhibit obvious separation after clustering.
[0040] 4. Incorporate the Particle Swarm Optimization algorithm
[0041] The Particle Swarm Optimization (PSO) algorithm is a branch of evolutionary computation and a stochastic search algorithm that simulates the biological activities in nature.
[0042] The main process of the PSO algorithm is as follows:
[0043] S1 Initialize all particles, that is, assign values to their velocities and positions, and set the historical best pBest of each individual to the current position, and the best individual in the population as the current gBest.
[0044] S2 In each generation of evolution, calculate the fitness function value of each particle.
[0045] S3 If the current fitness function value is better than the historical best value, then update pBest.
[0046] S4 If the current fitness function value is better than the global historical best value, then update gBest.
[0047] S5 Update the velocity and position of the d-th dimension of each particle i according to formulas (1) and (2) respectively:
[0048]
[0049] where ω is the inertia weight with a default value of 0.9, and c1 and c2 are acceleration coefficients (also known as learning factors) with default values of 2.0, and are two random numbers in the range [0, 1].
Claims
1. A method for integrating multi-organ single-cell sequencing data, characterized in that The method includes: S1: Data preprocessing, including two parts: data collection and data cleaning, and performing normalization, scaling, and feature selection operations on each single-cell sequencing dataset; S2: Data dimensionality reduction and data clustering. For each organ dataset, use PCA and Leiden in the scanpy package method to select the first 40 principal components for dimensionality reduction clustering; S3: Construct a coincidence matrix. For each organ, construct a coincidence matrix of itself and other organs, and input the coincidence matrix into the particle swarm optimization algorithm; S4: Combine the particle swarm optimization algorithm to perform batch effect correction on the data of multiple organs.
2. The multi-organ single-cell sequencing data integration method according to claim 1, wherein: The specific steps of S1 are as follows: S1-1: In the data collection stage, use mouse cell data from 12 organs. The mouse cell data are all from the datasets provided by Tabula Muris, and use the mouse data of 12 organs to construct a single-cell sequencing dataset; S1-2: In the data preprocessing stage, perform normalization, scaling, and feature selection on each single-cell sequencing dataset in turn. Then, adjust the expression values of each gene in all cells in each dataset so that each gene has a unit variance, and select the top 2000 highly variable genes. When performing integrated analysis on three or more organ datasets, it is necessary to select the highly variable genes in each organ dataset.
3. The multi-organ single-cell sequencing data integration method according to claim 1, wherein: The specific steps of S2 are as follows: S2-1: In the data dimensionality reduction stage, first perform PCA dimensionality reduction on each single-cell sequencing dataset, where the scaled data with highly variable genes is used. The number of PCs is a parameterized one, and by default, n_pcs = 15 is used; S2-2: In the data clustering stage, for each single-cell sequencing dataset, use the clustering method of the sc.pp.neighbors method in scanpy to obtain neighbor node scores, and then use the sc.tl.leiden method to calculate the clustering results of each single-cell sequencing dataset.
4. The multi-organ single-cell sequencing data integration method according to claim 1, wherein: The specific steps of S3 are as follows: S3-1: In the stage of constructing the coincidence matrix, calculate the proportion of each organ in the data to be integrated. Take the organ with the largest proportion as the benchmark, and calculate the remaining organs respectively and calculate the overlap degree of the highly expressed phenotype genes between other organs and this organ.
5. The multi-organ single-cell sequencing data integration method according to claim 1, characterized in that: The specific steps of S4 are as follows: S4-1: In the stage of combining the particle swarm optimization algorithm, combine the processed single-cell transcriptome sequencing data with the particle swarm optimization algorithm to implement the batch effect correction function. Select the Pearson correlation coefficient as the fitness function of the particle swarm optimization algorithm, use the first 40 principal components as the dimensions of the particle swarm optimization algorithm, and set the coincidence matrix as the search range of the particle swarm optimization algorithm.