Cell Data Comprehensive Analysis System
By providing a comprehensive cellular data analysis system with multiple analysis modules, the existing analysis platform's single function and lack of flexibility are solved, improving the efficiency and accuracy of single-cell RNA sequencing, and improving the user experience.
Patent Information
- Application Number
- CN202310670920.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-07
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2043-06-07
AI Technical Summary
The existing single-cell RNA sequencing and analysis platform has a single function and lacks flexibility, which affects the sequencing efficiency and accuracy.
It provides a comprehensive cell data analysis system, including acquisition module, filter module, batch effect elimination module, twin culling module, intercellular communication analysis module, clustering module, cell annotation module and visualization module, supporting a variety of data formats and flexible analysis and processing.
It improves the analysis efficiency and accuracy of single-cell RNA sequencing data, enhances the flexibility of data analysis, and improves the user experience through visual presentation.
Smart Images

Figure CN116959583B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cell data processing, and more particularly, to a cell data comprehensive analysis system. Background Art
[0002] For multicellular organisms, there are usually differences between cells, and the differences between different groups of cells vary. Such differences are not only reflected in morphology but also in genetic information, such as genomic information, gene expression levels, etc. With the in-depth and refined application of single-cell RNA sequencing (scRNA-seq), it is often necessary to perform single-cell sequencing on complex organs, and simply sequencing a few cells no longer meets the scientific research needs. That is, large-scale single-cell RNA sequencing has become a powerful way to resolve the heterogeneity of individual cells. Currently, although there are already analysis platforms for large-scale single-cell RNA sequencing (such as analysis tools like GranatumX and Cellxgene), the existing analysis platforms have relatively single functions, still lack the preprocessing of single-cell transcriptome data, and the data analysis lacks flexibility, thus affecting the efficiency and accuracy of single-cell RNA sequencing. Summary of the Invention
[0003] In view of this, the purpose of the embodiments of the present application is to provide a cell data comprehensive analysis system, which can improve the problem of lack of flexibility in data analysis and is beneficial to improving the efficiency and accuracy of single-cell RNA sequencing.
[0004] To achieve the above technical objectives, the technical solutions adopted in the present application are as follows:
[0005] The embodiments of the present application provide a cell data comprehensive analysis system, which includes: an acquisition module, a filtering module, a batch effect elimination module, a doublet removal module, a cell-cell communication analysis module, a clustering module, a cell annotation module, and a visualization module;
[0006] The acquisition module acquires a first data set generated based on single-cell omics uploaded by the user, and the data formats in the first data set include tsv format, txt format, csv format, RDS format, and HDF5 format;
[0007] When receiving a preprocessing instruction, the filtering module filters the first data set based on the index data set by the user for filtering cells to obtain a second data set, and the index data includes a first threshold range of UMI count, a second threshold range of feature gene count, a third threshold range of mitochondrial gene percentage, and a fourth threshold range corresponding to β-action expression;
[0008] The batch effect elimination module calls a batch effect elimination tool to eliminate the batch effect from the second data set, obtaining a third data set;
[0009] The doublet removal module calls the DoubletFinder tool to remove doublets from the third data set, obtaining a fourth data set;
[0010] When receiving a first instruction for intercellular communication analysis, the intercellular communication analysis module calls the pre-set CellPhoneDB tool, CellChat tool, and Cellcall tool from the tool library, and determines the strength of intercellular interactions and ligand-receptor differential pairs in the fourth data set through the CellPhoneDB tool, the CellChat tool, and the Cellcall tool, so as to obtain a communication analysis result characterizing intercellular communication;
[0011] When receiving a second instruction for cell type annotation, the clustering module calls the pre-set dimensionality reduction and clustering tool from the tool library to perform dimensionality reduction and clustering on the fourth data set, obtaining a fifth data set after clustering, where the dimensionality reduction and clustering tool includes the PCA tool, SC3 tool, tSNE tool, and UMAP tool;
[0012] The cell annotation module performs cell type annotation on the fifth data set by calling the annotation tool in the tool library, obtaining an annotation result, where the annotation tool includes any one of the SingleR tool and the atlas annotation tool based on the weighted nearest neighbor network algorithm;
[0013] The visualization module visually displays the communication analysis result and the annotation result based on a preset visualization strategy.
[0014] In some alternative embodiments, the system further includes:
[0015] A selection module for selecting cell data of a cell type of interest from the fifth data set as the cells to be tested;
[0016] An evaluation module for determining the corrected F score of a specified pathway of the cells to be tested through a preset algorithm;
[0017] A subpopulation acquisition module for obtaining subpopulation information of cell subtypes from the cells to be tested based on the corrected F score;
[0018] An input module for inputting the cell subtype data obtained from the cells to be tested into a pre-created single-cell pathway enrichment module, where the cell subtype data includes a gene raw expression matrix, a list of highly variable genes, and the subpopulation information of the cell subtypes;
[0019] An enrichment analysis module, which is used to call the clusterProfiler tool in the single-cell pathway enrichment module, take the list of highly variable genes in the cell subtype data as input data, and perform GO enrichment analysis to obtain an enrichment analysis result, where the enrichment analysis result includes a GO enrichment result, an FDR p-value, and gene information included in the pathway.
[0020] In some alternative embodiments, the preset algorithm is as follows:
[0021]
[0022]
[0023]
[0024]
[0025] Among them, bg ratio refers to the background value, N genes in a cell refers to the number of genes detected in the cell to be tested, N total genes refers to the total number of genes detected in the sample, fg ratio refers to the foreground value, N genes from a term in a cell refers to the number of genes in the specified pathway detected in the cell to be tested, N total genes of a term refers to the total number of genes in the specified pathway, F score refers to the F score of the specified pathway in the cell to be tested, adjF score refers to the corrected F score of the specified pathway in the cell to be tested, Q term refers to the p-value corrected by FDR, Q term Calculated by the clusterProfiler tool.
[0026] In some alternative embodiments, the system further includes:
[0027] A development trajectory prediction module, which is used to, when receiving a third instruction for trajectory prediction, call a pre-set trajectory prediction tool from the tool library, and determine the development trajectory of the cells in the fourth dataset through the trajectory prediction tool, where the trajectory prediction tool includes the Monocle2 tool or the SPRING tool.
[0028] In some alternative embodiments, the clustering module is further configured to call the PCA tool from the tool library, and perform dimensionality reduction on the fourth data set through the PCA tool to obtain the dimensionally reduced fourth data set, and call the SC3 tool from the tool library, and perform clustering on the dimensionally reduced fourth data set through the SC3 tool to obtain the fifth data set;
[0029] The visualization module is further configured to call the tSNE tool or the UMAP tool from the tool library to perform visual display on the fifth data set;
[0030] The cell annotation module is further configured to perform cell type annotation on the cells in the fifth data set according to the transcriptome data of the reference cell type by calling the SingleR tool to obtain the annotation result, or, by calling the atlas annotation tool, based on the weighted nearest neighbor network algorithm, map the scRNA-seq data corresponding to the cells in the fifth data set to a pre-annotated single-cell transcriptome atlas to obtain the annotation result.
[0031] In some alternative embodiments, the filtering module is further configured to:
[0032] Filter out the data in the first data set whose UMI count is not within the first threshold range based on the first threshold range of the UMI count set by the user through the first bar sliding button;
[0033] Filter out the data in the first data set whose feature gene count is not within the second threshold range based on the second threshold range of the feature gene count set by the user through the second bar sliding button;
[0034] Filter out the data in the first data set whose mitochondrial gene percentage is not within the third threshold range based on the third threshold range of the mitochondrial gene percentage set by the user through the third bar sliding button;
[0035] Filter out the data in the first data set whose β-action expression is not within the fourth threshold range based on the fourth threshold range corresponding to the β-action expression set by the user through the fourth bar sliding button;
[0036] Use the first data set filtered by UMI count, feature gene count, mitochondrial gene percentage, and β-action expression as the second data set.
[0037] In some alternative embodiments, the doublet removal module is further configured to:
[0038] Through the DoubletFinder tool, artificial simulated doublets are randomly fused from the single-cell data pre-uploaded by the user;
[0039] Mix the simulated doublets with the cells in the third dataset to obtain mixed cell data;
[0040] Through the PCA dimensionality reduction or PCA distance matrix in the DoubletFinder tool, find the proportion of artificial k-nearest neighbors pANN for each unit, where each unit is a bin divided from the mixed cell data, and each bin includes multiple feature genes;
[0041] Sort the mixed cell data based on a preset number of doublets and determine the threshold of the pANN value;
[0042] Based on the threshold of the pANN value, determine the doublet data from the mixed cell data and filter out the doublet data.
[0043] In some alternative embodiments, the system further includes:
[0044] A statistics module for classifying and statistically analyzing the fourth dataset according to preset categories and presenting the statistical classification results on a page, where the preset categories include the number of cells and the number of feature genes.
[0045] In some alternative embodiments, the system further includes:
[0046] A storage module for storing the first dataset, the second dataset, the third dataset, the fourth dataset, and the classification results based on a pre-created temporary project repository;
[0047] A query module for, when receiving a query instruction for querying the second dataset, obtaining the second dataset from the temporary project repository and presenting the second dataset in the form of a violin plot.
[0048] In some alternative embodiments, the system further includes:
[0049] A data upload module for uploading personalized genomes to the temporary project repository based on a pre-created upload interface;
[0050] A data deletion module for deleting specified genomes from the temporary project repository based on a pre-created deletion interface.
[0051] The invention adopting the above technical solution has the following advantages:
[0052] In the technical solution provided by this application, various preprocessing operations are performed on the first data set generated based on single-cell omics through an acquisition module, a filtering module, a batch effect elimination module, and a doublet removal module, such as cell filtering, batch effect elimination, doublet removal, etc. Through these preprocessing operations, the validity and reliability of the data can be improved. In addition, by using a cell communication analysis module, a clustering module, a cell annotation module, and a visualization module, processing operations such as cell communication analysis, dimensionality reduction clustering, and cell annotation can be supported, enriching the analysis functions, improving the flexibility of data analysis, and thus facilitating the improvement of the efficiency of data analysis. In addition, by visually displaying the communication analysis results and annotation results, it is beneficial for users to intuitively view the analysis results and enhance the user experience. Brief Description of the Drawings
[0053] This application can be further illustrated by the non-limiting embodiments shown in the drawings. It should be understood that the following drawings only show some embodiments of this application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0054] Figure 1 It is a block diagram of the cell data comprehensive analysis system provided by the embodiment of this application.
[0055] Figure 2 It is a schematic diagram of the bar sliding button provided by the embodiment of this application.
[0056] Figure 3 It is a functional block diagram of the tool library in the cell data comprehensive analysis system provided by the embodiment of this application.
[0057] Icons: 100 - Cell data comprehensive analysis system; 110 - Acquisition module; 120 - Filtering module; 130 - Batch effect elimination module; 140 - Doublet removal module; 150 - Cell communication analysis module; 160 - Clustering module; 170 - Cell annotation module; 180 - Visualization module. Detailed Description of the Embodiments
[0058] The following will describe this application in detail in combination with the drawings and specific embodiments. It should be noted that in the drawings or the description, similar or identical parts use the same reference numerals. The implementation manners not shown or described in the drawings are the forms known to those of ordinary skill in the art. In the description of this application, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0059] Please refer to Figure 1, an embodiment of the present application provides a cell data comprehensive analysis system 100, which can also be referred to as the GRACE (GRaphical Analyzing Cell Explorer) system. The cell data comprehensive analysis system 100 can be deployed on devices such as personal computers and servers, and can be used for preprocessing corresponding data sets of single-cell RNA sequencing (scRNA-seq), filtering out interfering data to improve the effectiveness of the data set. Single-cell transcriptome data is a data set generated based on single-cell omics and can be referred to as scRNA-seq data.
[0060] The cell data comprehensive analysis system 100 can also perform downstream analysis and processing based on the preprocessed data set, such as cell type annotation, communication analysis between the same type of cells, etc. In this way, it is beneficial to improve the accuracy and reliability of the analysis when performing cell analysis subsequently.
[0061] In this embodiment, the cell data comprehensive analysis system 100 may include: an acquisition module 110, a filtering module 120, a batch effect elimination module 130, a doublet elimination module 140, an intercellular communication analysis module 150, a clustering module 160, a cell annotation module 170, and a visualization module 180.
[0062] The acquisition module 110 acquires a first data set generated based on single-cell omics uploaded by the user, and the data format in the first data set includes tsv format, txt format, csv format, RDS format, and HDF5 format.
[0063] In this embodiment, the user can upload the first data set containing various formats of data to the cell data comprehensive analysis system 100. Understandably, in terms of data upload, the user can customize the establishment of a project folder and enter relevant information of the data, and load various data and formats generated by single-cell omics. For example, the data format can include tab-separated values / text (tsv / txt) format directly downloaded from the GEO database, comma-separated values (csv) format, the normal RDS format of the Seurat package as input, and the Hd5 format generated from the 10x Cellranger package. That is, the cell data comprehensive analysis system 100 can be compatible with the upload and preprocessing of single-cell transcriptome data in multiple data formats.
[0064] When receiving a preprocessing instruction, the filtering module 120 filters the first data set based on the index data for filtering cells set by the user, and obtains a second data set. The index data includes a first threshold range of UMI count, a second threshold range of characteristic gene count, a third threshold range of mitochondrial gene percentage, and a fourth threshold range corresponding to β-action expression.
[0065] In this embodiment, the filtering module 120 is further configured to:
[0066] Based on the first threshold range of UMI count set by the user through the first bar sliding button, filter out the data in the first dataset whose UMI count is not within the first threshold range;
[0067] Based on the second threshold range of characteristic gene count set by the user through the second bar sliding button, filter out the data in the first dataset whose characteristic gene count is not within the second threshold range;
[0068] Based on the third threshold range of mitochondrial gene percentage set by the user through the third bar sliding button, filter out the data in the first dataset whose mitochondrial gene percentage is not within the third threshold range;
[0069] Based on the fourth threshold range corresponding to β - action expression set by the user through the fourth bar sliding button, filter out the data in the first dataset whose β - action expression is not within the fourth threshold range;
[0070] Take the first dataset filtered by UMI count, characteristic gene count, mitochondrial gene percentage, and β - action expression as the second dataset.
[0071] Please refer to Figure 2 , the first bar sliding button (UMI Count), the second bar sliding button (Gene Count), the third bar sliding button [MT - Genes(%)], and the fourth bar sliding button (ACTB Activity) can be as Figure 2 shown. Different lengths of a single bar sliding button correspond to different values. The user can drag one or both endpoints of the bar sliding button to flexibly select the corresponding threshold range. Among them, each threshold data (such as the first threshold range, the second threshold range, the third threshold range, and the fourth threshold range, etc.) can be flexibly set by the user by dragging the corresponding bar sliding button. That is, the user can independently set the corresponding threshold interval according to personalized needs. In this way, the flexibility of threshold setting and the flexibility of cell data filtering can be improved.
[0072] In this embodiment, the first threshold range can be set according to the total count of UMI in each cell, and the second threshold range can be set according to the number of characteristic gene counts. Through the bar sliding button, the user can respectively select appropriate ranges for the maximum values of "UMI Count" and "Gene Count" for each cell as the corresponding threshold ranges. Using the first threshold range and the second threshold range, low - quality cells can be filtered.
[0073] In this embodiment, through the corresponding bar-shaped sliding button, the user can select the appropriate percentage range of mitochondrial genes in the total detected genes and the relative expression level of β-actin, which is usually highly expressed in living cells, as the third threshold range and the fourth threshold range respectively. By using the third threshold range and the fourth threshold range, potential dead cells in the first dataset can be filtered out.
[0074] After the user sets the corresponding threshold range, the filtering module 120 can automatically filter out the data in the first dataset that is not within the threshold range, thereby obtaining the second dataset.
[0075] Understandably, experimental errors that may be caused by the single-cell transcriptome raw data (i.e., there may be data from multiple batches in the first dataset) obtained at different times, by different operators, using different reagents, and different instruments are reflected in the gene expression levels as batch effects. And there are usually certain errors in the single-cell transcriptome data from different batches. If these data are directly used for downstream analysis without any processing, it may have a certain impact on the accuracy of the analysis results. In this embodiment, the batch effect elimination module 130 calls a batch effect elimination tool to eliminate the batch effects of the first dataset, which can improve the effectiveness of the data, and the effective second dataset is beneficial to improving the accuracy and reliability of subsequent cell analysis.
[0076] For example, the batch effect elimination module 130 calls a batch effect elimination tool to eliminate the batch effects of the second dataset, obtaining the third dataset. Please refer to Figure 3 , the batch effect elimination tool is a software toolkit in the tool library and can include multiple tools for implementing the batch effect elimination function. For example, the batch effect elimination tool can include any one of the RPCA tool, FastMNN tool, Harmony tool, scVI tool, and svANVI tool in Seurat v4. When batch effect elimination is required, the user can flexibly select the corresponding batch effect elimination tool according to the actual situation; of course, the cell data comprehensive analysis system 100 can also automatically select the corresponding batch effect elimination tool. The batch effect elimination tool can be encapsulated in the nextflow analysis process to achieve simplified and efficient data processing.
[0077] Understandably, single-cell RNA sequencing expects only one true cell under each barcode label. However, in actual data, there may be a situation where two or more cells share one barcode. Doublets refer to a droplet or a microwell containing two or more cells. Cells in the same droplet will carry the same Cell Barcode in subsequent analysis and will thus be considered pseudo-cells of one cell. Therefore, in fact, doublets are not only a mixture of two cells but may also be a situation where multiple cells are confused as one cell. The main characteristic of such pseudo-cells is that the number of detected UMIs and genes is often more than twice that of normal cells. The existence of doublets will have a great impact on the results of cell analysis, such as affecting cell type identification. In this embodiment, the doublet removal module 140 calls the DoubletFinder tool to remove doublets from the third data set to obtain a fourth data set, so as to reduce the impact of doublets on the subsequent cell analysis results.
[0078] Specifically, the doublet removal module 140 can be used for:
[0079] Randomly fuse artificial simulated doublets from the single-cell data pre-uploaded by the user through the DoubletFinder tool;
[0080] Mix the simulated doublets and the cells in the third data set to obtain mixed cell data;
[0081] Find the proportion of artificial k nearest neighbors pANN for each unit through PCA dimensionality reduction or PCA distance matrix in the DoubletFinder tool, where each unit is a bin divided from the mixed cell data, and each bin includes multiple feature genes;
[0082] Sort the mixed cell data based on a preset number of doublets and determine the threshold of the pANN value;
[0083] Based on the threshold of the pANN value, determine doublet data from the mixed cell data and filter out the doublet data.
[0084] Understandably, DoubletFinder can generate artificially simulated doublets and incorporate the simulated doublets into the original single-cell expression data. In principle, the artificially simulated doublets will be closer to the true doublets in the third dataset. Based on this, by calculating the proportion of artificially simulated doublets (pANN) among the K nearest neighbor cells of each cell, the doublet probability in each sample of the third dataset can be ranked according to the pANN value. Additionally, according to the statistical principle of the Poisson distribution, the number of doublets in each sample can be calculated. Combining with the previous cell pANN value ranking, doublet filtering can be achieved.
[0085] Among them, the preset number of doublets and the threshold of the pANN value can both be flexibly set according to the actual situation. The characteristic genes included in each bin can be divided from the mixed cell data according to the actual situation.
[0086] Please refer to again Figure 3 , during the creation of the tool library, the user can prepare in advance the software packages corresponding to the CellPhoneDB tool, the CellChat tool, and the Cellcall tool. Then, integrate the software packages to achieve the integration of multiple tools. When integrating the tools, for the obtained software packages, use Shiny to build a GUI (Graphical User Interface) and the Plotly graphics library for the interactive interface of the communication analysis function, and integrate the software packages into the tool library.
[0087] In addition, after the user prepares in advance the PCA (Principal Component Analysis) tool, the SC3 (Single Cell Consensus Clustering) tool, the tSNE (t-Stochastic Neighbor Embedding) tool, and the UMAP (Uniform Manifold Approximation and Projection) tool, when integrating the tools, for the pre-prepared software packages, use Shiny to build a GUI and the Plotly graphics library for the visualization framework of the dimensionality reduction clustering function, and integrate the software packages into the tool library for convenient subsequent calls.
[0088] In this embodiment, a visualization framework of the tool is constructed based on GUI and Shiny, which is beneficial for non-programming experts to analyze single-cell transcriptome datasets. Among them, the functions of each tool (such as CellPhoneDB tool, CellChat tool, Cellcall tool, SC3 tool, tSNE tool, UMAP tool, etc.) are conventional techniques and will not be elaborated here.
[0089] When the intercellular communication analysis module 150 receives a first instruction for intercellular communication analysis, it calls the pre-set CellPhoneDB tool, CellChat tool, and Cellcall tool from the tool library, and determines the intensity of intercellular interaction and the ligand-receptor differential pairs in the fourth dataset through the CellPhoneDB tool, the CellChat tool, and the Cellcall tool, so as to obtain a communication analysis result characterizing intercellular communication.
[0090] When the data cell comprehensive analysis system is deployed on the server, the user can control cell analysis through the WEB page of a personal computer. For example, click the start button for intercellular communication analysis, thereby generating a first instruction for intercellular communication analysis. After receiving the first instruction, the server can call tools such as the CellPhoneDB tool, CellChat tool, and Cellcall tool from the tool library, and calculate the intensity of intercellular interaction and the ligand-receptor differential pairs in the fourth dataset, thereby realizing the analysis of intercellular communication to obtain a communication analysis result.
[0091] When the clustering module 160 receives a second instruction for cell type annotation, it calls the pre-set dimensionality reduction and clustering tool from the tool library to perform dimensionality reduction and clustering on the fourth dataset to obtain a fifth dataset after clustering, where the dimensionality reduction and clustering tool includes PCA tool, SC3 tool, tSNE tool, and UMAP tool.
[0092] In this embodiment, the clustering module 160 is further configured to call the PCA tool from the tool library, perform dimensionality reduction on the fourth dataset through the PCA tool to obtain a fourth dataset after dimensionality reduction, and call the SC3 tool from the tool library, and perform clustering on the fourth dataset after dimensionality reduction through the SC3 tool to obtain the fifth dataset. The visualization module 180 can be used to call the tSNE tool or the UMAP tool from the tool library to perform visual display on the fifth dataset.
[0093] Understandably, the user can customize the number of Highly Variable Genes (HVGs) for constructing PCA. For example, the current default number of HVG genes is 2,000. In addition, the user can use sklearn, the LinearSVC model, and the ExtraTreesClassifier to select a set of feature genes using prior classification information. Using the PCA tool, dimensionality reduction of HVGs can be achieved. The functions of the SC3 tool, the tSNE tool, and the UMAP tool are conventional techniques and will not be elaborated here.
[0094] The cell annotation module 170 annotates the cell types of the fifth dataset by calling the annotation tools in the tool library to obtain an annotation result, where the annotation tools include any one of the SingleR tool and the atlas annotation tool based on the weighted nearest neighbor network algorithm.
[0095] In this embodiment, the cell annotation module 170 is further configured to annotate the cell types of the cells in the fifth dataset according to the transcriptome data of the reference cell types by calling the SingleR tool to obtain the annotation result, or to map the scRNA-seq data corresponding to the cells in the fifth dataset to a pre-annotated single-cell transcriptome atlas based on the weighted nearest neighbor network algorithm by calling the atlas annotation tool to obtain the annotation result.
[0096] Understandably, the SingleR tool uses the "Human Primary Cell Atlas Data" and "Mouse RNAseq Data" built into the Celldex software package as the transcriptome data of the reference cell types; the SingleR software package can perform cell type annotation based on the similarity between the characteristics of the cell data in the fifth dataset and the transcriptome data of the reference cell types.
[0097] When using the atlas annotation tool for annotation, after mapping the queried scRNA-seq data to the single-cell transcriptome atlas, the annotation content of the cell type can be determined based on the mapping relationship.
[0098] The visualization module 180 visually displays the communication analysis result and the annotation result based on a preset visualization strategy.
[0099] Among them, the preset visualization strategy can be flexibly set according to the actual situation. For example: generating icons for any two cells with communication in the communication analysis result and connecting the icons of any two cells with communication by line segments;
[0100] Generate scatter points corresponding to each cell in the second dataset, determine the rendering colors corresponding to the scatter points of different cell types according to the cell types in the second data in the annotation result, and render the scatter points of the corresponding cell types based on the rendering colors, where the rendering colors of the scatter points of different cell types are different.
[0101] Understandably, for the results of cell - cell communication analysis, the visualization module 180 can automatically generate two cell icons with communication. Among them, different types of cell icons can be different. Then, connect the two cell icons with communication by line segments. In this way, it is convenient for users to intuitively view the communication analysis results.
[0102] When performing dimensionality reduction clustering on cell data, a single cell can be used as a scatter point and rendered. The scatter points of different cell types can be rendered in different colors to facilitate users to intuitively view the distribution of different types of cells after clustering.
[0103] As an alternative implementation, the cell data comprehensive analysis system 100 may further include: a selection module, an evaluation module, a sub - population acquisition module 110, an input module, and an enrichment analysis module.
[0104] The selection module is used to select the cell data of the cell type that the user expects to focus on from the fifth dataset as the cells to be tested;
[0105] The evaluation module is used to determine the corrected F - score of the specified pathway of the cells to be tested through a preset algorithm;
[0106] The sub - population acquisition module 110 is used to obtain the sub - population information of cell subtypes from the cells to be tested based on the corrected F - score;
[0107] The input module is used to input the cell subtype data obtained from the cells to be tested into a pre - created single - cell pathway enrichment module. The cell subtype data includes the gene raw expression matrix, the list of highly variable genes, and the sub - population information of the cell subtypes;
[0108] The enrichment analysis module is used to call the clusterProfiler tool in the single - cell pathway enrichment module, use the list of highly variable genes in the cell subtype data as input data, and perform GO enrichment analysis to obtain the enrichment analysis results. The enrichment analysis results include GO enrichment results, FDR p - value, and the gene information included in the pathway.
[0109] In this embodiment, the preset algorithm is the following formula:
[0110]
[0111]
[0112]
[0113]
[0114] Among them, bg ratio refers to the background value, N genes in a cell refers to the number of genes detected in the cells to be tested, N total genes refers to the total number of genes detected in the sample, fg ratio refers to the foreground value, N genes from a term in a cell refers to the number of genes of the specified pathway detected in the cells to be tested, N total genes of a term refers to the total number of all genes of the specified pathway, F score refers to the F score of the specified pathway of the cells to be tested, adjF score refers to the corrected F score of the specified pathway of the cells to be tested, Q term refers to the p value corrected by FDR, Q term Obtained by calculation with the clusterProfiler tool.
[0115] Understandably, in formula (1), the number of genes detected in the cells to be tested (N genes in a cell ) divided by the total number of genes detected in the sample (N total genes ) serves as the background value (bg ratio ).
[0116] In formula (2), the number of genes of the specified pathway detected in the cells to be tested (N genes from a term in a cell ) divided by the total number of all genes of the specified pathway (N total genes of a term ) serves as the foreground value (fg ratio ).
[0117] In formula (3), the foreground value (fg ratio ) divided by the background value (bg ratio ) serves as the F score of the specified pathway of the cells to be tested (F score ).
[0118] In formula (4), the F score (F score ) divided by the Q value (Q term ) takes the logarithm as the corrected F score of the specified pathway of the cells to be tested (adjF score ). Understandably, the Q value (Q term ) is the p value corrected by FDR, obtained by calculation with clusterProfiler, and the obtaining method is the conventional method, which is not specifically limited here.
[0119] In this embodiment, for the execution of "single-cell pathway enrichment" analysis, the user needs to prepare data of cell subtypes, including: the original gene expression matrix (Count Matrix), the list of highly variable genes (HVG markers), and the cell subtype clustering information (Clusters&Annotated Clusters). Among them, the user only needs to select the cell types of interest from the cell type annotation results in the second dataset as the cells to be tested and perform cell subtype analysis.
[0120] Among them, the acquisition methods of the original gene expression matrix and the list of highly variable genes are conventional methods. For the acquisition method of the cell subtype clustering information, refer to the relevant implementation process of the above preset algorithm, which will not be elaborated here.
[0121] When performing pathway enrichment analysis, the obtained original gene expression matrix, list of highly variable genes, and cell subtype clustering information can be used as input data and input into the single-cell pathway enrichment module for analysis and processing.
[0122] Before inputting the cell subtype data into the single-cell pathway enrichment module, preprocessing is usually required. For example, starting from the selected cell types of interest, the cell data comprehensive analysis system 100 extracts the corresponding cells (i.e., the cells to be tested) from the initial expression matrix to construct a new expression matrix. After completing the extraction of the data subset and the construction of the new matrix, this preprocessing process may include:
[0123] Filter out genes expressed in fewer than 3 cells and cells expressing fewer than 200 genes;
[0124] Filter out cells with a mitochondrial gene detection ratio higher than 10%;
[0125] Use SCTransform to standardize the filtered data;
[0126] Use the VST algorithm to extract the top 2000 genes with the highest expression SD as the highly variable genes for PCA and perform PCA dimensionality reduction;
[0127] Select the top 10 PCs of the PCA dimensionality reduction results to perform UMAP dimensionality reduction and Louvain clustering to obtain clusters information;
[0128] Based on SingleR and Louvain clustering, perform cell type annotation to obtain annotated clusters information;
[0129] Extract the genes involved in the top N PCs from the PCA as the list of highly variable genes, where N can be flexibly controlled by the user.
[0130] In this embodiment, by combining the enrichment analysis results of single-cell pathway enrichment analysis with the prediction results of cell developmental trajectories (invoking the monocle2 results of the corresponding data subset, i.e., the DDRTree dimensionality reduction part), the changes in the biological process pathways and the expression of specific genes in different states along the developmental trajectory can be studied.
[0131] The cell data comprehensive analysis system 100 can support mapping the corrected F-score of each cell into the dimensionality reduction space, such as UMAP and DDRTree.
[0132] As an alternative embodiment, the cell data comprehensive analysis system 100 may further include a developmental trajectory prediction module. The developmental trajectory prediction module is configured to, when receiving a third instruction for trajectory prediction, call a pre-set trajectory prediction tool from the tool library and determine the developmental trajectory of the cells in the fourth data set through the trajectory prediction tool, where the trajectory prediction tool includes the Monocle2 tool or the SPRING tool.
[0133] Understandably, both the Monocle2 tool and the SPRING tool can be used for cell developmental trajectory prediction. The Monocle2 tool and the SPRING tool can be integrated into the tool library to enrich the prediction methods of developmental trajectories.
[0134] As an alternative embodiment, the cell data comprehensive analysis system 100 may further include:
[0135] A statistical module, configured to classify and statistically analyze the fourth data set according to a preset category and display the statistical classification results on a page, where the preset category includes the number of cells and the number of characteristic genes.
[0136] Understandably, after preprocessing such as filtering is completed, the statistical module can interactively display the preprocessed statistical results, such as the number of cells and the number of characteristic genes, to facilitate users to view the result data more intuitively.
[0137] As an alternative embodiment, the cell data comprehensive analysis system 100 may further include: a storage module and a query module. Among them, the storage module is configured to store the first data set, the second data set, the third data set, the fourth data set, and the classification results based on a pre-created temporary project repository; the query module is configured to, when receiving a query instruction for querying the second data set, obtain the second data set from the temporary project repository and display the second data set in the form of a violin plot.
[0138] Understandably, by providing a temporary project repository for the user, in this way, the data sets generated in each link of the preprocessing process can be stored, so as to facilitate the user to view or call the data sets from the temporary project repository at any time according to the needs.
[0139] When the user needs to view the data in the temporary project repository, the user can input a query instruction through the WEB interface of the cell data comprehensive analysis system 100. Then, the cell data comprehensive analysis system 100 obtains the data set that the user expects to find from the temporary project repository and displays the data set in the form of a violin plot. For example, the user can view the violin plots of UMI count and feature count in the second data set, as well as the violin plots of mitochondrial gene percentage and β-action expression in the drawing and visualization area of the WEB page.
[0140] As an optional implementation manner, the cell data comprehensive analysis system 100 may further include: a data upload module and a data deletion module. Among them, the data upload module is used to upload the personalized genome to the temporary project repository based on a pre-created upload interface; the data deletion module is used to delete the specified genome from the temporary project repository based on a pre-created deletion interface.
[0141] Understandably, the data upload module, the upload interface, as well as the data deletion module and the deletion interface can provide the user with the ability to upload the personalized genome for genome deletion. In this way, it can facilitate the downstream analysis of cells.
[0142] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented through hardware, or can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of this application.
[0143] The above are only the embodiments of this application and are not used to limit the protection scope of this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.
Claims
1. A cell data comprehensive analysis system, characterized in that: The system includes: an acquisition module, a filtering module, a batch effect elimination module, a twin elimination module, an intercellular communication analysis module, a clustering module, a cell annotation module and a visualization module; The acquisition module acquires a first data set generated based on single-cell omics uploaded by a user, wherein the data format of the first data set includes tsv format, txt format, csv format, RDS format and HDF5 format; The filtering module, upon receiving the preprocessing instruction, filters the first data set based on the indicator data for filtering cells set by the user to obtain a second data set, wherein the indicator data includes a first threshold range of UMI counts, a second threshold range of feature gene counts, a third threshold range of mitochondrial gene percentages, and a fourth threshold range corresponding to β-action expression; The batch effect elimination module calls a batch effect elimination tool to eliminate the batch effect of the second data set to obtain a third data set; The doublet elimination module performs doublet elimination on the third data set by calling the DoubletFinder tool to obtain a fourth data set; When receiving the first instruction for intercellular communication analysis, the intercellular communication analysis module calls the preset CellPhoneDB tool, CellChat tool and Cellcall tool from the tool library, and determines the intensity of the intercellular interaction and the difference pairs of ligand receptors in the fourth data set through the CellPhoneDB tool, the CellChat tool and the Cellcall tool to obtain a communication analysis result characterizing the intercellular communication; When receiving the second instruction for cell type annotation, the clustering module calls a preset dimensionality reduction clustering tool from the tool library to perform dimensionality reduction clustering on the fourth data set to obtain a clustered fifth data set, wherein the dimensionality reduction clustering tool includes a PCA tool, an SC3 tool, a tSNE tool, and a UMAP tool; The cell annotation module performs cell type annotation on the fifth data set by calling the annotation tool in the tool library to obtain an annotation result, wherein the annotation tool includes any one of the SingleR tool and the atlas annotation tool based on the weighted nearest neighbor network algorithm; The visualization module visualizes the communication analysis results and the annotation results based on a preset visualization strategy; The system further comprises: A selection module, used for selecting cell data of a desired cell type from the fifth data set as cells to be tested; An evaluation module, used to determine the F score of the designated pathway of the cell to be tested after correction by a preset algorithm; A clustering acquisition module, used for acquiring clustering information of cell subclasses from the cells to be tested based on the corrected F score; An input module, used to input the cell subclass data obtained from the cells to be tested into a pre-created single cell pathway enrichment module, wherein the cell subclass data includes an original gene expression matrix, a list of highly variable genes, and clustering information of the cell subclasses; An enrichment analysis module is used to call the clusterProfiler tool in the single-cell pathway enrichment module, take the highly variable gene list in the cell subclass data as input data, and perform GO enrichment analysis to obtain enrichment analysis results, wherein the enrichment analysis results include GO enrichment results, FDR p-value, and gene information included in the pathway; The preset algorithm is as follows: Among them, bg ratio Refers to the background value, N genesinacell Refers to the number of genes detected in the cells to be tested, N totalgenes Refers to the total number of genes detected in the sample, fg ratio Refers to the foreground value, N genesfromaterminacell Refers to the number of genes of a specified pathway detected in the cells to be tested, N totalgenesofaterm refers to the total number of genes in the specified pathway, F score refers to the F score of the specified pathway of the cell to be tested, adjF score refers to the F score after correction of the specified pathway of the cell to be tested, Q term refers to the p-value after FDR correction, Q term Calculated by clusterProfiler tool; The filtering module is also used for: Based on the first threshold range of UMI counts set by a user through a first bar slide button, filter out data in the first data set whose UMI counts are not within the first threshold range; Based on the second threshold range of the characteristic gene count set by the user through the second bar sliding button, filter out data in the first data set whose characteristic gene count is not within the second threshold range; Based on the third threshold range of the mitochondrial gene percentage set by the user through the third bar slide button, filter out the data in the first data set whose mitochondrial gene percentage is not within the third threshold range; Based on the fourth threshold range corresponding to the β-action expression set by the user through the fourth bar sliding button, filtering out the data in the first data set whose β-action expression is not within the fourth threshold range, wherein, in the first bar sliding button, the second bar sliding button, the third bar sliding button and the fourth bar sliding button, different lengths of a single bar sliding button correspond to different values; The first data set filtered by UMI count, characteristic gene count, mitochondrial gene percentage and β-action expression is used as the second data set.
2. The system according to claim 1, characterized in that The system further comprises: A developmental trajectory prediction module is used to, upon receiving a third instruction for trajectory prediction, call a preset trajectory prediction tool from the tool library and determine the developmental trajectory of the cells in the fourth data set by using the trajectory prediction tool, wherein the trajectory prediction tool includes a Monocle2 tool or a SPRING tool.
3. The system according to claim 1, characterized in that The clustering module is further used to call the PCA tool from the tool library, and perform dimension reduction on the fourth data set by using the PCA tool to obtain a fourth data set after dimension reduction, and to call the SC3 tool from the tool library, and perform clustering on the fourth data set after dimension reduction by using the SC3 tool to obtain the fifth data set; The visualization module is further used to call the tSNE tool or the UMAP tool from the tool library to visualize the fifth data set; The cell annotation module is also used to annotate the cells in the fifth data set with cell types according to the transcriptome data of the reference cell type by calling the SingleR tool to obtain the annotation result, or to map the scRNA-seq data corresponding to the cells in the fifth data set to a pre-annotated single-cell transcriptome map based on a weighted nearest neighbor network algorithm by calling the map annotation tool to obtain the annotation result.
4. The system according to claim 1, characterized in that The twin elimination module is also used for: Through the DoubletFinder tool, artificial simulated twins are randomly fused from the single-cell data uploaded in advance by the user; Mixing the simulated twin cells with cells in the third data set to obtain mixed cell data; Find the proportion of the artificial k nearest neighbor pANN of each unit through PCA dimensionality reduction or PCA distance matrix in the DoubletFinder tool, wherein each unit is a bin divided from the mixed cell data, and each bin includes a plurality of characteristic genes; Sorting the mixed cell data based on a preset number of doublets, and determining a threshold value of a pANN value; Based on the threshold of the pANN value, twin data are determined from the mixed cell data, and the twin data are filtered out.
5. The system according to claim 1, characterized in that The system further comprises: The statistical module is used to classify and count the fourth data set according to preset categories, and display the classification results obtained by statistics on a page, wherein the preset categories include the number of cells and the number of characteristic genes.
6. The system according to claim 5, characterized in that The system further comprises: A storage module, configured to store the first data set, the second data set, the third data set, the fourth data set and the classification result based on a pre-created temporary project repository; The query module is used for, when receiving a query instruction for querying the second data set, obtaining the second data set from the temporary project repository and displaying the second data set in a violin plot.
7. The system according to any one of claims 1 to 6, characterized in that: The system further comprises: A data upload module, used to upload personalized genomes to a temporary project repository based on a pre-created upload interface; The data deletion module is used to delete the specified genome from the temporary project repository based on a pre-created deletion interface.
Citation Information
Patent Citations
Tumor neoantigen screening method fused with single cell TCR sequencing data
CN113160887A
Esophageal squamous cell carcinoma data processing method and system
CN116129998A