Comprehensive cell data analysis system
By using a comprehensive cell data analysis system to preprocess and analyze single-cell RNA sequencing data, the problem of limited functionality in existing platforms is solved, data analysis efficiency and accuracy are improved, and user experience is enhanced.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- JINFENG LAB
- Filing Date
- 2024-10-23
- Publication Date
- 2026-04-30
AI Technical Summary
Existing single-cell RNA sequencing analysis platforms are limited in function and lack flexibility, resulting in insufficient efficiency and accuracy in data analysis.
A comprehensive cell data analysis system is provided, including an acquisition module, a filtering module, a batch effect elimination module, a double cell elimination module, an intercellular communication analysis module, a clustering module, a cell annotation module, and a visualization module. These modules are used to preprocess and analyze single-cell omics data, supporting operations such as cell filtering, batch effect elimination, double cell elimination, cell communication analysis, dimensionality reduction clustering, and cell annotation.
It improves the efficiency and accuracy of data analysis from single-cell RNA sequencing, enhances the flexibility of data analysis, and improves the user experience through visualization.
Smart Images

Figure CN2024126747_30042026_PF_FP_ABST
Abstract
Description
Cell Data Comprehensive Analysis System Technical Field
[0001] This invention relates to the field of cell data processing technology, and more specifically, to a comprehensive cell data analysis system. Background Technology
[0002] In multicellular organisms, differences typically exist between cells, and these differences also vary between different cell populations. These differences are not only reflected in morphology but also in genetic information, such as genomic information and gene expression levels. With the deepening and refinement of single-cell RNA sequencing (scRNA-seq) applications, it is often necessary to perform single-cell sequencing on complex organs; simply sequencing a few cells no longer meets research needs. In other words, large-scale single-cell RNA sequencing has become a powerful way to break down the heterogeneity of individual cells. Currently, although large-scale single-cell RNA sequencing analysis platforms exist (such as GranatumX and Cellxgene), their functions are relatively limited, lacking preprocessing of single-cell transcriptome data and resulting in a lack of flexibility in data analysis, thus affecting the efficiency and accuracy of single-cell RNA sequencing.
[0003] Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a comprehensive cell data analysis system that can improve the lack of flexibility in data analysis and improve the efficiency and accuracy of single-cell RNA sequencing.
[0005] To achieve the above technical objectives, the technical solution adopted in this application is as follows:
[0006] This application provides a comprehensive cell data analysis system, which includes: an acquisition module, a filtering module, a batch effect elimination module, a twin-cell elimination module, an intercellular communication analysis module, a clustering module, a cell annotation module, and a visualization module;
[0007] The acquisition module acquires the first dataset generated based on single-cell omics uploaded by the user. The data formats in the first dataset include tsv, txt, csv, RDS and HDF5 formats.
[0008] When the filtering module receives the preprocessing instruction, it filters the first dataset based on the indicator data for filtering cells set by the user to obtain the second dataset. The indicator data includes a first threshold range for UMI count, a second threshold range for feature gene count, a third threshold range for mitochondrial gene percentage, and a fourth threshold range corresponding to β-action expression.
[0009] The batch effect elimination module calls the batch effect elimination tool to perform batch effect elimination on the second dataset to obtain the third dataset;
[0010] The twin removal module calls the DoubletFinder tool to remove twins from the third dataset, thus obtaining the fourth dataset.
[0011] When the intercellular communication analysis module receives the first instruction for intercellular communication analysis, it calls the pre-set CellPhoneDB tool, CellChat tool, and Cellcall tool from the tool library, and uses the CellPhoneDB tool, the CellChat tool, and the Cellcall tool to determine the strength of intercellular interactions and the difference pairs of ligand receptors in the fourth dataset, so as to obtain communication analysis results characterizing intercellular communication.
[0012] When the clustering module receives the second instruction for cell type annotation, it calls a pre-set dimensionality reduction clustering tool from the tool library to perform dimensionality reduction clustering on the fourth dataset to obtain the clustered fifth dataset. The dimensionality reduction clustering tool includes PCA, SC3, tSNE, and UMAP tools.
[0013] The cell annotation module annotates the cell types of the fifth dataset by calling the annotation tools in the tool library to obtain the annotation results. The annotation tools include any one of the SingleR tool and a graph annotation tool based on the weighted nearest neighbor network algorithm.
[0014] The visualization module displays the communication analysis results and the annotation results in a visual format based on a preset visualization strategy.
[0015] In some alternative implementations, the system further includes:
[0016] The selection module is used to select cell data of the cell type of interest from the fifth dataset as cells to be tested.
[0017] The evaluation module is used to determine the F score of the test cells after correction for a specified pathway using a preset algorithm.
[0018] The cluster acquisition module is used to acquire cell subclass cluster information from the cells to be tested based on the corrected F score.
[0019] The input module is used to input the cell subclass data obtained from the cell to be tested into a pre-created single-cell pathway enrichment module. The cell subclass data includes the original gene expression matrix, the list of hypervariable genes, and the cluster information of the cell subclass.
[0020] The enrichment analysis module is used to call the clusterProfiler tool in the single-cell pathway enrichment module, take the list of hypervariable genes in the cell subclass data as input data, and perform GO enrichment analysis to obtain the enrichment analysis results. The enrichment analysis results include GO enrichment results, FDR p-value, and gene information contained in the pathway.
[0021] In some optional implementations, the preset algorithm is as follows:
[0022] Among them, bg ratio Refers to the background value, N genes in a cell N refers to the number of genes detected in the cells being tested. total genes The total number of genes detected in the sample, fg ratio Foreground value; Ngenes from a term in a cell refers to the number of genes in the specified pathway detected in the cell being tested; Ntotal genes of a term refers to the total number of genes in the specified pathway; F score The F score of the specified pathway in the test cells, adjF score The F-score, Q, refers to the F-score of the test cells after correction for the specified pathway. term The p-value, Q, is the value corrected by FDR. term Calculated using the clusterProfiler tool.
[0023] In some alternative implementations, the system further includes:
[0024] The developmental trajectory prediction module is used to, upon receiving a third instruction for trajectory prediction, call a pre-set trajectory prediction tool from the tool library and determine the developmental trajectory of cells in the fourth dataset using the trajectory prediction tool, wherein the trajectory prediction tool includes the Monocle2 tool or the SPRING tool.
[0025] In some optional implementations, the clustering module is further configured to call the PCA tool from the tool library and perform dimensionality reduction on the fourth dataset using the PCA tool to obtain a dimensionality-reduced fourth dataset, and to call the SC3 tool from the tool library and perform clustering on the dimensionality-reduced fourth dataset using the SC3 tool to obtain the fifth dataset;
[0026] The visualization module is also used to call the tSNE tool or the UMAP tool from the tool library to visualize the fifth dataset;
[0027] The cell annotation module is also used to annotate the cells in the fifth dataset with cell type based on the transcriptome data of the reference cell type by calling the SingleR tool, and obtain the annotation result; or, by calling the atlas annotation tool, based on the weighted nearest neighbor network algorithm, to map the scRNA-seq data corresponding to the cells in the fifth dataset onto the pre-annotated single-cell transcriptome atlas, and obtain the annotation result.
[0028] In some alternative implementations, the filtering module is further configured to:
[0029] Based on the first threshold range of UMI counts set by the user via the first bar slide button, filter out data in the first dataset whose UMI counts are not within the first threshold range;
[0030] Based on the second threshold range of the feature gene count set by the user via the second bar slider button, data in the first dataset whose feature gene count is not within the second threshold range are filtered out;
[0031] Based on the third threshold range of the percentage of mitochondrial genes set by the user via the third bar slider button, data in the first dataset whose percentage of mitochondrial genes is not within the third threshold range are filtered out;
[0032] Based on the fourth threshold range corresponding to the β-action expression set by the user through the fourth bar slide button, data in the first dataset whose β-action expression is not within the fourth threshold range are filtered out;
[0033] The first dataset, filtered by UMI count, feature gene count, mitochondrial gene percentage, and β-action expression, is used as the second dataset.
[0034] In some optional implementations, the twin-cell elimination module is further configured to:
[0035] The DoubletFinder tool is used to randomly fuse artificial twin cells from single-cell data uploaded by the user in advance.
[0036] The simulated twin cells and the cells in the third dataset are mixed to obtain mixed cell data;
[0037] Using the PCA dimensionality reduction or PCA distance matrix in the DoubletFinder tool, find the proportion of artificial k nearest neighbor pANNs for each unit, where each unit is a bin divided from the mixed cell data, and each bin includes multiple feature genes;
[0038] The mixed cell data is sorted based on a preset number of doublets, and a threshold for the pANN value is determined.
[0039] Based on the threshold of the pANN value, twin-cell data are determined from the mixed cell data, and the twin-cell data are filtered out.
[0040] In some alternative implementations, the system further includes:
[0041] The statistics module is used to classify and statistically analyze the fourth dataset according to preset categories, and to display the classification results on a page. The preset categories include the number of cells and the number of feature genes.
[0042] In some alternative implementations, the system further includes:
[0043] A storage module is used to store the first dataset, the second dataset, the third dataset, the fourth dataset, and the classification results based on a pre-created temporary project repository;
[0044] The query module is used to retrieve the second dataset from the temporary project repository and display the second dataset in a violin diagram when a query instruction for querying the second dataset is received.
[0045] In some alternative implementations, the system further includes:
[0046] The data upload module is used to upload personalized genomes to a temporary project repository based on a pre-created upload interface;
[0047] The data deletion module is used to delete a specified genome from the temporary project repository based on a pre-created deletion interface.
[0048] The invention employing the above technical solution has the following advantages:
[0049] In the technical solution provided in this application, various preprocessing operations, such as cell filtering, batch effect elimination, and double-cell elimination, are performed on the first dataset generated based on single-cell omics through an acquisition module, a filtering module, a batch effect elimination module, and a double-cell elimination module. These preprocessing operations can improve the validity and reliability of the data. Furthermore, the intercellular communication analysis module, clustering module, cell annotation module, and visualization module support cell communication analysis, dimensionality reduction clustering, and cell annotation, enriching the analysis functions, improving the flexibility of data analysis, and thus enhancing the efficiency of data analysis. In addition, visualizing the communication analysis results and annotation results allows users to intuitively view the analysis results, improving the user experience. Attached Figure Description
[0050] This application can be further illustrated by the non-limiting embodiments given in the accompanying drawings. It should be understood that the following drawings only illustrate some embodiments of this application and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained from these drawings without any inventive effort.
[0051] Figure 1 is a block diagram of the cell data comprehensive analysis system provided in an embodiment of this application.
[0052] Figure 2 is a schematic diagram of the bar-shaped sliding button provided in an embodiment of this application.
[0053] Figure 3 is a functional block diagram of the tool library in the cell data comprehensive analysis system provided in the embodiments of this application.
[0054] Icons: 100 - Comprehensive Cell Data Analysis System; 110 - Acquisition Module; 120 - Filtering Module; 130 - Batch Effect Elimination Module; 140 - Twin Cell Removal Module; 150 - Intercellular Communication Analysis Module; 160 - Clustering Module; 170 - Cell Annotation Module; 180 - Visualization Module. Detailed Implementation
[0055] The present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that similar or identical parts are referred to by the same reference numerals in the drawings or description. Implementations not shown or described in the drawings are forms known to those skilled in the art. In the description of this application, terms such as "first" and "second" are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0056] Referring to Figure 1, this application provides a cell data comprehensive analysis system 100, which can also be called a GRACE (GRaphical Analyzing Cell Explorer) system. The cell data comprehensive analysis system 100 can be deployed on personal computers, servers, and other devices. It can be used to preprocess datasets from single-cell RNA sequencing (scRNA-seq) data, filtering out interfering data to improve the effectiveness of the dataset. Single-cell transcriptome data is a dataset generated based on single-cell omics and can be called scRNA-seq data.
[0057] The cell data comprehensive analysis system 100 can also perform downstream analysis based on the preprocessed dataset, such as cell type annotation and communication between homotype cells. This is beneficial to improving the accuracy and reliability of subsequent cell analysis.
[0058] In this embodiment, the cell data comprehensive analysis system 100 may include: an acquisition module 110, a filtering module 120, a batch effect elimination module 130, a twin-cell elimination module 140, an inter-cell communication analysis module 150, a clustering module 160, a cell annotation module 170, and a visualization module 180.
[0059] The acquisition module 110 acquires the first dataset generated based on single-cell omics uploaded by the user. The data formats in the first dataset include tsv, txt, csv, RDS and HDF5 formats.
[0060] In this embodiment, the user can upload a first dataset containing data in multiple formats to the cell data comprehensive analysis system 100. Understandably, regarding data upload, the user can customize the creation of project folders and the relevant information for entering data, loading various data and formats generated by single-cell omics. For example, data formats may include tab-delimited values / text (tsv / txt) format and comma-delimited values (csv) format downloaded directly from the GEO database, the normal RDS format from the Seurat package as input, and the Hd5 format generated from the 10x Cellranger package. That is, the cell data comprehensive analysis system 100 can accommodate the upload and preprocessing of single-cell transcriptome data in multiple data formats.
[0061] When the filtering module 120 receives the preprocessing instruction, it filters the first dataset based on the indicator data for filtering cells set by the user to obtain the second dataset. The indicator data includes a first threshold range for UMI count, a second threshold range for feature gene count, a third threshold range for mitochondrial gene percentage, and a fourth threshold range corresponding to β-action expression.
[0062] In this embodiment, the filtering module 120 is further used for:
[0063] Based on the first threshold range of UMI counts set by the user via the first bar slide button, filter out data in the first dataset whose UMI counts are not within the first threshold range;
[0064] Based on the second threshold range of the feature gene count set by the user via the second bar slider button, data in the first dataset whose feature gene count is not within the second threshold range are filtered out;
[0065] Based on the third threshold range of the percentage of mitochondrial genes set by the user via the third bar slider button, data in the first dataset whose percentage of mitochondrial genes is not within the third threshold range are filtered out;
[0066] Based on the fourth threshold range corresponding to the β-action expression set by the user through the fourth bar slide button, data in the first dataset whose β-action expression is not within the fourth threshold range are filtered out;
[0067] The first dataset, filtered by UMI count, feature gene count, mitochondrial gene percentage, and β-action expression, is used as the second dataset.
[0068] Referring to Figure 2, the first slider button (UMI Count), the second slider button (Gene Count), the third slider button [MT-Genes(%)], and the fourth slider button (ACTB Activity) can be shown in Figure 2. Different lengths of a single slider button correspond to different values. Users can drag one or both ends of the slider button to flexibly select the corresponding threshold range. Each threshold data point (such as the first threshold range, second threshold range, third threshold range, and fourth threshold range) can be flexibly set by the user by dragging the corresponding slider button. That is, users can independently set the corresponding threshold range according to their individual needs, thus improving the flexibility of threshold setting and cell data filtering.
[0069] In this embodiment, the first threshold range can be set based on the total UMI count in each cell, and the second threshold range can be set based on the number of characteristic gene counts. Using a slider button, the user can select an appropriate range for the maximum values of "UMI Count" and "Gene Count" for each cell, serving as the corresponding threshold range. By utilizing the first and second threshold ranges, low-quality cells can be filtered out.
[0070] In this embodiment, by using the corresponding bar slider button, the user can select an appropriate percentage range of mitochondrial genes among the total detected genes, and the relative expression level of β-actin, which is typically highly expressed in living cells, as the third and fourth threshold ranges, respectively. Using the third and fourth threshold ranges, potentially dead cells in the first dataset can be filtered out.
[0071] Once the user sets the appropriate threshold range, the filtering module 120 can automatically filter out data in the first dataset that is not within the threshold range, thus obtaining the second dataset.
[0072] Understandably, raw single-cell transcriptome data acquired at different times, by different operators, using different reagents, and with different instruments (i.e., the first dataset may contain data from multiple batches) may lead to experimental errors, which are reflected in gene expression levels as batch effects. Single-cell transcriptome data from different batches typically contain some error; if this data is used directly for downstream analysis without any processing, it may affect the accuracy of the analysis results. In this embodiment, the batch effect elimination module 130 calls a batch effect elimination tool to eliminate batch effects from the first dataset, which can improve the validity of the data. An effective second dataset is beneficial for improving the accuracy and reliability of subsequent cell analyses.
[0073] For example, the batch effect elimination module 130 calls a batch effect elimination tool to perform batch effect elimination on the second dataset to obtain a third dataset. Referring to Figure 3, the batch effect elimination tool is a software toolkit in a tool library, which may include multiple tools for implementing batch effect elimination. For example, the batch effect elimination tool may include any one of the following tools in Seurat v4: RPCA, FastMNN, Harmony, scVI, and svANVI. When batch effect elimination is required, the user can flexibly select the appropriate batch effect elimination tool according to the actual situation; of course, the cell data comprehensive analysis system 100 can also automatically select the appropriate batch effect elimination tool. The batch effect elimination tool can be encapsulated in the Nextflow analysis workflow to achieve simplified and efficient data processing.
[0074] Understandably, single-cell RNA sequencing expects only one real cell under each barcode tag, but in actual data, there may be two or more cells sharing a single barcode. A twin cell refers to two or more cells contained within a single droplet or microwell. Cells in the same droplet carry the same Cell Barcode in subsequent analysis and are thus considered pseudo-cells. Therefore, twin cells are not only a mixture of two cells but may also be multiple cells confused as one. The main characteristic of these pseudo-cells is that the number of UMIs and genes detected is often twice or more than that of normal cells. The presence of twin cells can significantly impact cell analysis results, such as affecting cell type identification. In this embodiment, the twin removal module 140 calls the DoubletFinder tool to remove twins from the third dataset, obtaining a fourth dataset. This reduces the impact of twins on subsequent cell analysis results.
[0075] Specifically, the twin-cell elimination module 140 can be used for:
[0076] The DoubletFinder tool is used to randomly fuse artificial twin cells from single-cell data uploaded by the user in advance.
[0077] The simulated twin cells and the cells in the third dataset are mixed to obtain mixed cell data;
[0078] Using the PCA dimensionality reduction or PCA distance matrix in the DoubletFinder tool, find the proportion of artificial k nearest neighbor pANNs for each unit, where each unit is a bin divided from the mixed cell data, and each bin includes multiple feature genes;
[0079] The mixed cell data is sorted based on a preset number of doublets, and a threshold for the pANN value is determined.
[0080] Based on the threshold of the pANN value, twin-cell data are determined from the mixed cell data, and the twin-cell data are filtered out.
[0081] Understandably, DoubletFinder can generate artificially simulated doublets and incorporate them into the original single-cell expression data. In principle, these simulated doublets will be closer to the real doublets in the third dataset. Based on this, by calculating the proportion of artificially simulated doublets in each cell's K nearest neighbor cells (pANN), the probability of doublets in each sample of the third dataset can be ranked according to the pANN value. Furthermore, based on the statistical principle of the Poisson distribution, the number of doublets in each sample can be calculated. Combining this with the previous cell pANN value ranking, doublet filtering can be achieved.
[0082] The preset number of doublets and the threshold for pANN values can be flexibly set according to the actual situation. The feature genes included in each bin can be divided from the mixed cell data according to the actual situation.
[0083] Referring again to Figure 3, during the creation of the tool library, users can prepare the corresponding software packages for the CellPhoneDB, CellChat, and Cellcall tools in advance. Then, these packages are integrated to achieve the integration of multiple tools. When integrating the tools, Shiny is used to build the GUI (Graphical User Interface) and Plotly's graphical library to create an interactive interface for communication analysis functions, and the software packages are then integrated into the tool library.
[0084] In addition, after users have prepared PCA (Principal Component Analysis), SC3 (Single Cell Consensus Clustering), tSNE (t-Stochastic Neighbor Embedding), and UMAP (Uniform Manifold Approximation and Projection) tools in advance, when integrating the tools, a visualization framework for dimensionality reduction and clustering functions is built using Shiny's GUI and Plotly's graphics library for the pre-prepared software packages, and the software packages are integrated into the tool library for easy subsequent use.
[0085] In this embodiment, a visualization framework for the tools is constructed based on the GUI and Shiny, which facilitates the analysis of single-cell transcriptome datasets by non-programming experts. The functions of each tool (such as CellPhoneDB, CellChat, Cellcall, SC3, tSNE, and UMAP) are conventional techniques and will not be elaborated upon here.
[0086] When the intercellular communication analysis module 150 receives a first instruction for intercellular communication analysis, it calls the pre-set CellPhoneDB, CellChat, and Cellcall tools from the tool library, and uses the CellPhoneDB, CellChat, and Cellcall tools to determine the strength of intercellular interactions and the difference pairs of ligand-receptor in the fourth dataset to obtain communication analysis results characterizing intercellular communication.
[0087] When the data-driven cellular comprehensive analysis system is deployed on a server, users can control cell analysis through a web page on their personal computers. For example, clicking the "Start" button for intercellular communication analysis generates the first instruction for this analysis. Upon receiving the first instruction, the server can call tools such as CellPhoneDB, CellChat, and Cellcall from its tool library and calculate the strength of intercellular interactions and the differences in ligand-receptor pairs in the fourth dataset, thereby performing intercellular communication analysis and obtaining the results.
[0088] When the clustering module 160 receives the second instruction for cell type annotation, it calls a pre-set dimensionality reduction clustering tool from the tool library to perform dimensionality reduction clustering on the fourth dataset to obtain the clustered fifth dataset. The dimensionality reduction clustering tool includes PCA, SC3, tSNE and UMAP tools.
[0089] In this embodiment, the clustering module 160 is further configured to call the PCA tool from the tool library and use the PCA tool to reduce the dimensionality of the fourth dataset to obtain a dimensionality-reduced fourth dataset; and to call the SC3 tool from the tool library and use the SC3 tool to cluster the dimensionality-reduced fourth dataset to obtain the fifth dataset. The visualization module 180 can be configured to call the tSNE tool or the UMAP tool from the tool library to visualize the fifth dataset.
[0090] Understandably, users can customize the number of Highly Variable Genes (HVGs) used to construct PCA. For example, the current default number of HVG genes is 2000. Furthermore, users can use sklearn, LinearSVC models, and ExtraTreesClassifier to select the feature gene set using prior classification information. Dimensionality reduction of HVGs can be achieved using PCA tools. The functions of the SC3, tSNE, and UMAP tools are standard techniques and will not be elaborated upon here.
[0091] The cell annotation module 170 annotates the cell types of the fifth dataset by calling the annotation tools in the tool library to obtain annotation results. The annotation tools include any one of the SingleR tool and a graph annotation tool based on the weighted nearest neighbor network algorithm.
[0092] In this embodiment, the cell annotation module 170 is further configured to, by calling the SingleR tool, annotate the cells in the fifth dataset with cell type based on the transcriptome data of the reference cell type to obtain the annotation result, or, by calling the atlas annotation tool, map the scRNA-seq data corresponding to the cells in the fifth dataset to the pre-annotated single-cell transcriptome atlas based on the weighted nearest neighbor network algorithm to obtain the annotation result.
[0093] Understandably, the SingleR tool uses the “human primary cell atlas data” and “mouse RNAseq data” built into the Celldex package as transcriptome data for reference cell types; the SingleR package can annotate cell types based on the similarity between the cell data in the fifth dataset and the transcriptome data of the reference cell types.
[0094] When using map annotation tools, after the queried scRNA-seq data is mapped to a single-cell transcriptome map, the annotation content for cell types can be determined based on the mapping relationship.
[0095] The visualization module 180 displays the communication analysis results and the annotation results in a visual format based on a preset visualization strategy.
[0096] The preset visualization strategy can be flexibly set according to the actual situation. For example, generate icons of any two cells that are communicating in the communication analysis results, and connect the icons of any two cells that are communicating with a line segment.
[0097] Scatter points are generated for each cell in the second dataset. Based on the cell type in the second data in the annotation results, the rendering color corresponding to the scatter points of different cell types is determined. The scatter points of the corresponding cell types are rendered based on the rendering color, wherein the rendering color of the scatter points of different cell types is different.
[0098] Understandably, based on the results of intercellular communication analysis, the visualization module 180 can automatically generate icons for two communicating cells, where different types of cells can have different icons. Then, the two communicating cell icons are connected by a line segment, allowing users to easily and intuitively view the communication analysis results.
[0099] When performing dimensionality reduction clustering on cell data, a single cell can be treated as a scatter plot and rendered. Different cell types can be rendered in different colors to allow users to visually view the distribution of different cell types after clustering.
[0100] As an optional implementation, the cell data comprehensive analysis system 100 may further include: a selection module, an evaluation module, a cluster acquisition module 110, an input module, and an enrichment analysis module.
[0101] The selection module is used to select cell data of the cell type of interest from the fifth dataset as cells to be tested.
[0102] The evaluation module is used to determine the F score of the test cells after correction for a specified pathway using a preset algorithm.
[0103] The cluster acquisition module 110 is used to acquire cell subclass cluster information from the cells to be tested based on the corrected F score.
[0104] The input module is used to input the cell subclass data obtained from the cell to be tested into a pre-created single-cell pathway enrichment module. The cell subclass data includes the original gene expression matrix, the list of hypervariable genes, and the cluster information of the cell subclass.
[0105] The enrichment analysis module is used to call the clusterProfiler tool in the single-cell pathway enrichment module, take the list of hypervariable genes in the cell subclass data as input data, and perform GO enrichment analysis to obtain the enrichment analysis results. The enrichment analysis results include GO enrichment results, FDR p-value, and gene information contained in the pathway.
[0106] In this embodiment, the preset algorithm is as follows:
[0107] Among them, bg ratio Refers to the background value, N genes in a cell N refers to the number of genes detected in the cells being tested. total genes The total number of genes detected in the sample, fg ratio Foreground value; Ngenes from a term in a cell refers to the number of genes in the specified pathway detected in the cell being tested; Ntotal genes of a term refers to the total number of genes in the specified pathway; F score The F score of the specified pathway in the test cells, adjF score The F-score, Q, refers to the F-score of the test cells after correction for the specified pathway. term The p-value, Q, is the value corrected by FDR. term Calculated using the clusterProfiler tool.
[0108] Understandably, in formula (1), the number of genes (N) detected in the test cell is... genes in a cell Divide by the total number of genes detected in the sample (N) total genes ) as background value (bg) ratio ).
[0109] In formula (2), the number of genes from a term in a cell detected in the test cell (Ngenes from a term in a cell) divided by the total number of genes in the specified pathway (Ntotal genes of a term) is used as the foreground value (fg). ratio ).
[0110] In formula (3), the foreground value (fg) ratio Divide by the background value (bg) ratio The F score (F) is used as the designated pathway for the cells under test. score ).
[0111] In formula (4), F score (F score Divide by Q value (Q term The logarithm is taken as the F-score (adjF) of the test cells after correction for the specified pathway. score Understandably, the Q-value (Q...) term The p-value is the FDR-corrected value, calculated by clusterProfiler in a conventional manner, without any specific limitations.
[0112] In this embodiment, to perform "single-cell pathway enrichment" analysis, the user needs to prepare cell subclass data, including: a raw gene expression matrix (Count Matrix), a list of hypervariable genes (HVG markers), and cell subclass cluster information (Clusters & Annotated Clusters). Specifically, the user only needs to select the cell type of interest from the cell type annotation results in the second dataset as the cell to be tested and perform cell subclass analysis.
[0113] The original gene expression matrix and the list of highly variable genes are obtained in a conventional manner. The method for obtaining the cell subclass cluster information is described in the relevant implementation process of the above-mentioned preset algorithm, and will not be repeated here.
[0114] When performing pathway enrichment analysis, the obtained original gene expression matrix, list of hypervariable genes, and cell subclass clustering information can be used as input data and input into the single-cell pathway enrichment module for analysis and processing.
[0115] Before inputting cell subclass data into the single-cell pathway enrichment module, preprocessing is typically required. For example, starting with the selected cell type of interest, the cell data comprehensive analysis system 100 extracts the corresponding cells (i.e., the cells to be tested) from the initial expression matrix to construct a new expression matrix. After completing the data subset extraction and new matrix construction, this preprocessing process may include:
[0116] Cells with fewer than 3 genes expressed and fewer than 200 genes expressed were filtered out.
[0117] Cells with a detected mitochondrial gene ratio higher than 10% were filtered out;
[0118] Use SCTransform to standardize the filtered data;
[0119] The VST algorithm was used to extract the top 2000 genes by expression level in SD as hypervariable genes for PCA, and PCA dimensionality reduction was performed.
[0120] The top 10 PCs from the PCA dimensionality reduction results were selected for UMAP dimensionality reduction and Louvain clustering to obtain cluster information;
[0121] Cell type annotation was performed based on SingleR and Louvain clustering to obtain annotated clusters information;
[0122] The genes involved in the top N PCs are extracted from PCA to form a list of hypervariable genes, where N can be flexibly controlled by the user.
[0123] In this embodiment, by combining the enrichment analysis results of single-cell pathway enrichment analysis with the prediction results of cell development trajectory (calling the monocle2 results of the corresponding data subset, i.e. the dimensionality reduction part of DDRTree), it is possible to study the changes in biological process pathways and the expression of specific genes along the development trajectory under different states.
[0124] The Cell Data Comprehensive Analysis System 100 can support mapping the corrected F score of each cell to a dimensionality-reduced space, such as UMAP and DDRTree.
[0125] As an optional implementation, the cell data comprehensive analysis system 100 may further include a developmental trajectory prediction module. The developmental trajectory prediction module, upon receiving a third instruction for trajectory prediction, calls a pre-set trajectory prediction tool from the tool library and uses the trajectory prediction tool to determine the developmental trajectory of cells in the fourth dataset, wherein the trajectory prediction tool includes the Monocle2 tool or the SPRING tool.
[0126] Understandably, both Monocle2 and SPRING tools can be used for predicting cell developmental trajectories. A tool library can integrate both Monocle2 and SPRING tools to enrich the methods for predicting developmental trajectories.
[0127] As an optional implementation, the cell data comprehensive analysis system 100 may further include:
[0128] The statistics module is used to classify and statistically analyze the fourth dataset according to preset categories, and to display the classification results on a page. The preset categories include the number of cells and the number of feature genes.
[0129] Understandably, after filtering and other preprocessing, the statistics module can interactively display the preprocessed statistical results, such as cell count and feature gene count, to make it easier for users to view the results data more intuitively.
[0130] As an optional implementation, the cell data comprehensive analysis system 100 may further include a storage module and a query module. The storage module is used to store the first dataset, the second dataset, the third dataset, the fourth dataset, and the classification results based on a pre-created temporary project repository. The query module is used to retrieve the second dataset from the temporary project repository and display the second dataset in a violin diagram format when a query instruction for querying the second dataset is received.
[0131] Understandably, by providing users with a temporary project repository, datasets generated at each stage of the preprocessing process can be stored, allowing users to easily view or retrieve the datasets from the temporary project repository as needed.
[0132] When a user needs to view data in the temporary project repository, they can enter a query command through the web interface of the Cell Data Comprehensive Analysis System 100. The Cell Data Comprehensive Analysis System 100 then retrieves the dataset the user is looking for from the temporary project repository and displays it in a violin plot format. For example, the user can view violin plots of UMI counts and feature counts, as well as violin plots of mitochondrial gene percentages and β-action expression, in the plotting and visualization area of the web page.
[0133] As an optional implementation, the cell data comprehensive analysis system 100 may further include a data upload module and a data deletion module. The data upload module is used to upload personalized genomes to a temporary project repository based on a pre-created upload interface; the data deletion module is used to delete specified genomes from the temporary project repository based on a pre-created deletion interface.
[0134] Understandably, the data upload module and upload interface, as well as the data deletion module and deletion interface, can provide users with the ability to upload personalized genomes for genome deletion, thus facilitating downstream cell analysis.
[0135] Based on the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by hardware or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application can be embodied in the form of a software product. This software product can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various implementation scenarios of this application.
[0136] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A comprehensive cell data analysis system, characterized in that, The system includes: an acquisition module, a filtering module, a batch effect elimination module, a twin-cell elimination module, an inter-cell communication analysis module, a clustering module, a cell annotation module, and a visualization module; The acquisition module acquires the first dataset generated based on single-cell omics uploaded by the user. The data formats in the first dataset include tsv, txt, csv, RDS and HDF5 formats. When the filtering module receives the preprocessing instruction, it filters the first dataset based on the indicator data for filtering cells set by the user to obtain the second dataset. The indicator data includes a first threshold range for UMI count, a second threshold range for feature gene count, a third threshold range for mitochondrial gene percentage, and a fourth threshold range corresponding to β-action expression. The batch effect elimination module calls the batch effect elimination tool to perform batch effect elimination on the second dataset to obtain the third dataset; The twin removal module calls the DoubletFinder tool to remove twins from the third dataset to obtain the fourth dataset; When the intercellular communication analysis module receives the first instruction for intercellular communication analysis, it calls the pre-set CellPhoneDB tool, CellChat tool, and Cellcall tool from the tool library, and uses the CellPhoneDB tool, the CellChat tool, and the Cellcall tool to determine the strength of intercellular interactions and the difference pairs of ligand receptors in the fourth dataset, so as to obtain communication analysis results characterizing intercellular communication. When the clustering module receives the second instruction for cell type annotation, it calls a pre-set dimensionality reduction clustering tool from the tool library to perform dimensionality reduction clustering on the fourth dataset to obtain the clustered fifth dataset. The dimensionality reduction clustering tool includes PCA, SC3, tSNE, and UMAP tools. The cell annotation module annotates the cell types of the fifth dataset by calling the annotation tools in the tool library to obtain the annotation results. The annotation tools include any one of the SingleR tool and a graph annotation tool based on the weighted nearest neighbor network algorithm. The visualization module displays the communication analysis results and the annotation results in a visual format based on a preset visualization strategy.
2. The system according to claim 1, characterized in that, The system also includes: The selection module is used to select cell data of the cell type of interest from the fifth dataset as cells to be tested. The evaluation module is used to determine the F score of the test cells after correction for a specified pathway using a preset algorithm. The clustering acquisition module is used to acquire cell subclasses from the test cells based on the corrected F score. information; The input module is used to input the cell subclass data obtained from the cell to be tested into a pre-created single-cell pathway enrichment module. The cell subclass data includes the original gene expression matrix, the list of hypervariable genes, and the cluster information of the cell subclass. The enrichment analysis module is used to call the clusterProfiler tool in the single-cell pathway enrichment module, take the list of hypervariable genes in the cell subclass data as input data, and perform GO enrichment analysis to obtain the enrichment analysis results. The enrichment analysis results include GO enrichment results, FDR p-value, and gene information contained in the pathway.
3. The system according to claim 2, characterized in that, The preset algorithm is as follows: Among them, bg ratio Refers to the background value, N genes in a cell N refers to the number of genes detected in the cells being tested. total genes The total number of genes detected in the sample, fg ratio Foreground value; Ngenes from a term in a cell refers to the number of genes in the specified pathway detected in the cell being tested; Ntotal genes of a term refers to the total number of genes in the specified pathway; F score The F score of the specified pathway in the test cells, adjF score The F-score, Q, refers to the F-score of the test cells after correction for the specified pathway. term The p-value, Q, is the value corrected by FDR. term Calculated using the clusterProfiler tool.
4. The system according to claim 1, characterized in that, The system also includes: The developmental trajectory prediction module is used to, upon receiving a third instruction for trajectory prediction, call a pre-set trajectory prediction tool from the tool library and determine the developmental trajectory of cells in the fourth dataset using the trajectory prediction tool, wherein the trajectory prediction tool includes the Monocle2 tool or the SPRING tool.
5. The system according to claim 1, characterized in that, The clustering module is also used to call the PCA tool from the tool library and use the PCA tool to reduce the dimensionality of the fourth dataset to obtain the dimensionality-reduced fourth dataset, and to call the SC3 tool from the tool library and use the SC3 tool to cluster the dimensionality-reduced fourth dataset to obtain the fifth dataset; The visualization module is also used to call the tSNE tool or the UMAP tool from the tool library to visualize the fifth dataset; The cell annotation module is also used to annotate cells in the fifth dataset based on transcriptome data of a reference cell type by calling the SingleR tool, thereby obtaining the annotation results; or, by calling the graph... The spectrum annotation tool, based on the weighted nearest neighbor network algorithm, maps the scRNA-seq data corresponding to cells in the fifth dataset onto a pre-annotated single-cell transcriptome map to obtain the annotation results.
6. The system according to claim 1, characterized in that, The filtering module is also used for: Based on the first threshold range of UMI counts set by the user via the first bar slide button, filter out data in the first dataset whose UMI counts are not within the first threshold range; Based on the second threshold range of the feature gene count set by the user via the second bar slider button, data in the first dataset whose feature gene count is not within the second threshold range are filtered out; Based on the third threshold range of the percentage of mitochondrial genes set by the user via the third bar slider button, data in the first dataset whose percentage of mitochondrial genes is not within the third threshold range are filtered out; Based on the fourth threshold range corresponding to the β-action expression set by the user through the fourth bar slide button, data in the first dataset whose β-action expression is not within the fourth threshold range are filtered out; The first dataset, filtered by UMI count, feature gene count, mitochondrial gene percentage, and β-action expression, is used as the second dataset.
7. The system according to claim 1, characterized in that, The twin-cell elimination module is also used for: The DoubletFinder tool is used to randomly fuse artificial twin cells from single-cell data uploaded by the user in advance. The simulated twin cells and the cells in the third dataset are mixed to obtain mixed cell data; Using the PCA dimensionality reduction or PCA distance matrix in the DoubletFinder tool, find the proportion of artificial k nearest neighbor pANNs for each unit, where each unit is a bin divided from the mixed cell data, and each bin includes multiple feature genes; The mixed cell data is sorted based on a preset number of doublets, and a threshold for the pANN value is determined. Based on the threshold of the pANN value, twin-cell data are determined from the mixed cell data, and the twin-cell data are filtered out.
8. The system according to claim 1, characterized in that, The system also includes: The statistics module is used to classify and statistically analyze the fourth dataset according to preset categories, and to display the classification results on a page. The preset categories include the number of cells and the number of feature genes.
9. The system according to claim 8, characterized in that, The system also includes: A storage module is used to store the first dataset, the second dataset, the third dataset, the fourth dataset, and the classification results based on a pre-created temporary project repository; The query module is used to retrieve the second dataset from the temporary project repository and display the second dataset in a violin diagram when a query instruction for querying the second dataset is received.
10. The system according to any one of claims 1-9, characterized in that, The system also includes: The data upload module is used to upload personalized genomes to a temporary project repository based on a pre-created upload interface; The data deletion module is used to delete a specified genome from the temporary project repository based on a pre-created deletion interface.
Citation Information
Patent Citations
Single cell transcriptome data analysis method and device and electronic equipment
CN115458062A
Single cell transcriptome data analysis processing method and electronic equipment
CN116913388A
Single cell transcriptome data preprocessing method, electronic equipment and storage medium
CN116913389A
Cell data comprehensive analysis system
CN116959583A
Systems, software, and methods for multiomic single cell classification and prediction and longitudinal trajectory analysis
US20240249839A1