Single-cell transcriptome data analysis and processing method, and electronic device
By integrating various tools and tool libraries, multi-format analysis of single-cell transcriptome data has been achieved, solving the problem of limited functionality in existing platforms, improving the flexibility and efficiency of data analysis, and enhancing the user experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2026-04-30
AI Technical Summary
Existing single-cell RNA sequencing analysis platforms have limited functionality and lack flexibility in data analysis, which affects data analysis efficiency.
This paper provides a method for analyzing and processing single-cell transcriptome data. It uses CellPhoneDB, CellChat, and Cellcall tools to perform inter-cell communication analysis, combines PCA, SC3, tSNE, and UMAP tools for dimensionality reduction and clustering, and uses SingleR tool for cell type annotation. It supports the analysis of datasets in multiple data formats and uses GUI and Shiny to build a visualization interface to display the analysis results.
The enhanced analytical capabilities have improved the flexibility and efficiency of data analysis, allowing users to view the results intuitively and enhancing the user experience.
Smart Images

Figure CN2024126955_30042026_PF_FP_ABST
Abstract
Description
Single-cell transcriptome data analysis and processing methods and electronic equipment Technical Field
[0001] This invention relates to the field of cell data processing technology, and more specifically, to a method and electronic device for analyzing and processing single-cell transcriptome data. Background Technology
[0002] With the deepening and refinement of single-cell RNA sequencing (scRNA-seq) applications, there is often a need to perform single-cell sequencing on complex organs, as sequencing just a few cells no longer meets research requirements. In other words, large-scale single-cell RNA sequencing has become a powerful way to break down the heterogeneity of individual cells. Currently, although large-scale single-cell RNA sequencing analysis platforms exist (such as GranatumX and Cellxgene), these existing platforms have relatively limited functionality and lack flexibility in data analysis, thus affecting the efficiency of data analysis.
[0003] Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a single-cell transcriptome data analysis and processing method and electronic device, which can improve the problem of relatively simple cell data analysis functions and lack of flexibility in data analysis, and help improve the efficiency of data analysis.
[0005] To achieve the above technical objectives, the technical solution adopted in this application is as follows:
[0006] In a first aspect, embodiments of this application provide a method for analyzing and processing single-cell transcriptome data, the method comprising:
[0007] Obtain the first dataset after preprocessing, wherein the first dataset is a dataset generated based on single-cell omics, and the data formats in the first dataset include tsv, txt, csv, RDS and HDF5 formats;
[0008] Upon receiving the first instruction for intercellular communication analysis, the pre-set CellPhoneDB, CellChat, and Cellcall tools are invoked from the tool library. The strength of intercellular interactions and the difference pairs of ligand-receptor in the first dataset are determined by the CellPhoneDB, CellChat, and Cellcall tools to obtain communication analysis results characterizing intercellular communication.
[0009] Upon receiving a second instruction for cell type annotation, a pre-set dimensionality reduction clustering tool is invoked from the tool library to perform dimensionality reduction clustering on the first dataset, resulting in a clustered second dataset. The dimensionality reduction clustering tool includes PCA, SC3, tSNE, and UMAP tools.
[0010] The second dataset is annotated with cell types by calling the annotation tools in the tool library to obtain the annotation results. The annotation tools include any one of the SingleR tool and a graph annotation tool based on the weighted nearest neighbor network algorithm.
[0011] Based on a preset visualization strategy, the communication analysis results and the annotation results are visualized.
[0012] In conjunction with the first aspect, in some alternative implementations, the method further includes, before obtaining the preprocessed first dataset:
[0013] Obtain the software packages corresponding to the CellPhoneDB tool, the CellChat tool, and the Cellcall tool;
[0014] Based on the graphical user interface (GUI) and Shiny, the corresponding software packages for the CellPhoneDB tool, the CellChat tool, and the Cellcall tool are added to the pre-created tool library.
[0015] In conjunction with the first aspect, in some alternative implementations, the method further includes, before obtaining the preprocessed first dataset:
[0016] Obtain the software packages corresponding to the PCA tool, the SC3 tool, the tSNE tool, and the UMAP tool;
[0017] Based on the graphical user interface (GUI) and Shiny, the software packages corresponding to the PCA tool, the SC3 tool, the tSNE tool, and the UMAP tool are added to the pre-created tool library.
[0018] In conjunction with the first aspect, in some alternative implementations, the method further includes:
[0019] Upon receiving a third instruction for trajectory prediction, a pre-set trajectory prediction tool is invoked from the tool library, and the developmental trajectory of cells in the first dataset is determined by the trajectory prediction tool, wherein the trajectory prediction tool includes the Monocle2 tool or the SPRING tool.
[0020] In conjunction with the first aspect, in some optional implementations, a pre-set dimensionality reduction clustering tool is invoked from the tool library to perform dimensionality reduction clustering on the first dataset, resulting in a clustered second dataset, including:
[0021] The PCA tool is called from the tool library, and the PCA tool is used to reduce the dimensionality of the first dataset to obtain the dimensionality-reduced first dataset.
[0022] The SC3 tool is called from the tool library, and the SC3 tool is used to cluster the dimensionality-reduced first dataset to obtain the second dataset;
[0023] The second dataset is visualized by calling the tSNE tool or the UMAP tool from the tool library.
[0024] In conjunction with the first aspect, in some optional implementations, the second dataset is annotated with cell types by invoking the annotation tools in the tool library to obtain annotation results, including:
[0025] By calling the SingleR tool, cell type annotation is performed on the cells in the second dataset based on the transcriptome data of the reference cell type, and the annotation results are obtained.
[0026] Alternatively, by invoking the map annotation tool, based on the weighted nearest neighbor network algorithm, the scRNA-seq data corresponding to cells in the second dataset can be mapped onto a pre-annotated single-cell transcriptome map to obtain the annotation results.
[0027] In conjunction with the first aspect, in some optional implementations, the communication analysis results and the annotation results are visualized based on a preset visualization strategy, including:
[0028] Generate icons for any two cells that are communicating in the communication analysis results, and connect the icons of any two cells that are communicating with each other by a line segment;
[0029] Scatter points are generated for each cell in the second dataset. Based on the cell type in the second data in the annotation results, the rendering color corresponding to the scatter points of different cell types is determined. The scatter points of the corresponding cell types are rendered based on the rendering color, wherein the rendering color of the scatter points of different cell types is different.
[0030] In conjunction with the first aspect, in some alternative implementations, the method further includes:
[0031] The communication analysis results and the annotation results are stored in a pre-created temporary project repository.
[0032] Secondly, this application also provides an electronic device, which includes a processor and a memory coupled to each other, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the electronic device performs the above-described method.
[0033] The invention employing the above technical solution has the following advantages:
[0034] The technical solution provided in this application supports the analysis and processing of datasets in various formats, including TSV, TXT, CSV, RDS, and HDF5. Furthermore, the tool library includes pre-set cell analysis tools that support cell communication analysis, dimensionality reduction clustering, cell annotation, and other processing operations, enriching analytical functions, improving data analysis flexibility, and ultimately enhancing data analysis efficiency. In addition, visualizing the communication analysis and annotation results allows users to intuitively view the analysis results, improving the user experience. Attached Figure Description
[0035] This application can be further illustrated by the non-limiting embodiments given in the accompanying drawings. It should be understood that the following drawings only illustrate some embodiments of this application and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained from these drawings without any inventive effort.
[0036] Figure 1 is a flowchart illustrating the single-cell transcriptome data analysis and processing method provided in the embodiments of this application. Detailed Implementation
[0037] The present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that similar or identical parts are referred to by the same reference numerals in the drawings or description. Implementations not shown or described in the drawings are forms known to those skilled in the art. In the description of this application, terms such as "first" and "second" are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0038] This application provides an electronic device that may include a processing module and a storage module. The storage module stores a computer program, which, when executed by the processing module, enables the electronic device to perform the corresponding steps in the following single-cell transcriptome data analysis and processing method.
[0039] The electronic device can be, but is not limited to, a personal computer or a server. The electronic device may be equipped with an analysis platform for single-cell transcriptome data analysis and processing, or the analysis platform may be deployed on a cloud server, and the electronic device can access the analysis platform via the web to perform single-cell transcriptome data analysis and processing.
[0040] Understandably, this analysis platform can be used to analyze and process datasets from single-cell RNA sequencing (scRNA-seq), such as cell filtering, batch effect elimination, and double-cell knockout. Single-cell transcriptome data, which is a dataset generated based on single-cell omics, can be called scRNA-seq data.
[0041] Referring to Figure 1, this application also provides a method for analyzing and processing single-cell transcriptome data, which can be applied to the aforementioned electronic device. The single-cell transcriptome data analysis and processing method may include the following steps:
[0042] Step 110: Obtain the first dataset after preprocessing. The first dataset is a dataset generated based on single-cell omics. The data formats in the first dataset include tsv, txt, csv, RDS and HDF5 formats.
[0043] Step 120: Upon receiving the first instruction for intercellular communication analysis, the pre-set CellPhoneDB tool, CellChat tool, and Cellcall tool are called from the tool library, and the strength of intercellular interactions and the difference pairs of ligand receptors in the first dataset are determined by the CellPhoneDB tool, the CellChat tool, and the Cellcall tool to obtain communication analysis results characterizing intercellular communication.
[0044] Step 130: Upon receiving the second instruction for cell type annotation, a pre-set dimensionality reduction clustering tool is invoked from the tool library to perform dimensionality reduction clustering on the first dataset to obtain the clustered second dataset. The dimensionality reduction clustering tool includes PCA, SC3, tSNE, and UMAP tools.
[0045] Step 140: By calling the annotation tools in the tool library, the second dataset is annotated with cell types to obtain the annotation results. The annotation tools include any one of the SingleR tool and a graph annotation tool based on the weighted nearest neighbor network algorithm.
[0046] Step 150: Based on a preset visualization strategy, the communication analysis results and the annotation results are visualized.
[0047] The following is a detailed explanation of each step in the single-cell transcriptome data analysis and processing method:
[0048] Prior to step 110, the method may further include:
[0049] Obtain the software packages corresponding to the CellPhoneDB tool, the CellChat tool, and the Cellcall tool;
[0050] Based on GUI (Graphical User Interface) and Shiny, the corresponding software packages for the CellPhoneDB tool, the CellChat tool, and the Cellcall tool are added to the pre-created tool library.
[0051] Understandably, users can prepare the corresponding software packages for the CellPhoneDB, CellChat, and Cellcall tools in advance, and then obtain the corresponding pre-prepared software packages via electronic devices during tool integration. During tool integration, for the obtained software packages, Shiny is used to build the GUI and Plotly's graphical library to create an interactive interface for communication analysis functions, and the software packages are then integrated into the tool library.
[0052] Prior to step 110, the method may further include:
[0053] Obtain the software packages corresponding to the PCA (Principal Component Analysis) tool, the SC3 (Single Cell Consensus Clustering) tool, the tSNE (t-Stochastic Neighbor Embedding) tool, and the UMAP (Uniform Manifold Approximation and Projection) tool;
[0054] Based on the graphical user interface (GUI) and Shiny, the software packages corresponding to the PCA tool, the SC3 tool, the tSNE tool, and the UMAP tool are added to the pre-created tool library.
[0055] Understandably, after users have prepared PCA, SC3, tSNE, and UMAP tools in advance, the electronic device can obtain the corresponding software packages when integrating the tools. For the obtained software packages, a visualization framework for dimensionality reduction and clustering functions is built using Shiny's GUI and Plotly's graphics library, and the software packages are integrated into the tool library for easy subsequent use.
[0056] In this embodiment, a visualization framework for the tools is constructed based on the GUI and Shiny, which facilitates the analysis of single-cell transcriptome datasets by non-programming experts. The functions of each tool (such as CellPhoneDB, CellChat, Cellcall, SC3, tSNE, and UMAP) are conventional techniques and will not be elaborated upon here.
[0057] In step 110, the preprocessed first dataset can refer to the dataset obtained after preprocessing operations such as cell filtering, batch effect elimination, and double cell removal from the original dataset generated based on single-cell omics. This preprocessed first dataset helps improve the accuracy and reliability of subsequent cell analysis results.
[0058] The first dataset can contain multiple data formats, meaning it supports uploading and processing data in various formats, which improves the user experience. These data formats can include tab-delimited value / text (tsv / txt) and comma-delimited value (csv) formats downloaded directly from the GEO database, normal RDS formats from the Seurat package as input, and HD5 formats generated from the 10x Cellranger package.
[0059] In step 120, the user can control the cell analysis through the electronic device's web page. For example, clicking the start button for intercellular communication analysis generates a first instruction for the analysis. Upon receiving the first instruction, the electronic device's processing module can call tools such as CellPhoneDB, CellChat, and Cellcall from the tool library and calculate the strength of intercellular interactions and the differences in ligand-receptor pairs in the first dataset, thereby performing intercellular communication analysis and obtaining the results.
[0060] In step 130, a pre-set dimensionality reduction clustering tool is invoked from the tool library to perform dimensionality reduction clustering on the first dataset, resulting in a clustered second dataset, including:
[0061] The PCA tool is called from the tool library, and the PCA tool is used to reduce the dimensionality of the first dataset to obtain the dimensionality-reduced first dataset.
[0062] The SC3 tool is called from the tool library, and the SC3 tool is used to cluster the dimensionality-reduced first dataset to obtain the second dataset;
[0063] The second dataset is visualized by calling the tSNE tool or the UMAP tool from the tool library.
[0064] Understandably, users can customize the number of Highly Variable Genes (HVGs) used to construct PCA. For example, the current default number of HVG genes is 2000. Furthermore, users can use sklearn, LinearSVC models, and ExtraTreesClassifier to select the feature gene set using prior classification information. Dimensionality reduction of HVGs can be achieved using PCA tools. The functions of the SC3, tSNE, and UMAP tools are standard techniques and will not be elaborated upon here.
[0065] In step 140, the second dataset is annotated with cell types by calling the annotation tools in the tool library to obtain annotation results, including:
[0066] By calling the SingleR tool, cell type annotation is performed on the cells in the second dataset based on the transcriptome data of the reference cell type, and the annotation results are obtained.
[0067] Alternatively, by invoking the map annotation tool, based on the Weighted-Nearest Neighbor (WNN) algorithm, the scRNA-seq data corresponding to cells in the second dataset can be mapped onto a pre-annotated single-cell transcriptome map to obtain the annotation results.
[0068] Understandably, the SingleR tool uses the “human primary cell atlas data” and “mouse RNAseq data” built into the Celldex package as transcriptome data for reference cell types; the SingleR package can annotate cell types based on the similarity between the cell data in the second dataset and the transcriptome data of the reference cell types.
[0069] When using map annotation tools, after the queried scRNA-seq data is mapped to a single-cell transcriptome map, the annotation content for cell types can be determined based on the mapping relationship.
[0070] In step 150, based on a preset visualization strategy, the communication analysis results and the annotation results are visualized, including:
[0071] Generate icons for any two cells that are communicating in the communication analysis results, and connect the icons of any two cells that are communicating with each other by a line segment;
[0072] Scatter points are generated for each cell in the second dataset. Based on the cell type in the second data in the annotation results, the rendering color corresponding to the scatter points of different cell types is determined. The scatter points of the corresponding cell types are rendered based on the rendering color, wherein the rendering color of the scatter points of different cell types is different.
[0073] Understandably, based on the results of intercellular communication analysis, the electronic device can automatically generate icons for two cells that are communicating, where different types of cells can have different icons. Then, the two communicating cell icons are connected by a line segment, allowing users to easily and intuitively view the communication analysis results.
[0074] When performing dimensionality reduction clustering on cell data, a single cell can be treated as a scatter plot and rendered. Different cell types can be rendered in different colors to allow users to visually view the distribution of different cell types after clustering.
[0075] As an optional implementation, the method may further include:
[0076] Upon receiving a third instruction for trajectory prediction, a pre-set trajectory prediction tool is invoked from the tool library, and the developmental trajectory of cells in the first dataset is determined by the trajectory prediction tool, wherein the trajectory prediction tool includes the Monocle2 tool or the SPRING tool.
[0077] Understandably, both Monocle2 and SPRING tools can be used for predicting cell developmental trajectories. A tool library can integrate both Monocle2 and SPRING tools to enrich the methods for predicting developmental trajectories.
[0078] As an optional implementation, the method may further include:
[0079] Select cell data of the cell type of interest from the second dataset as the cells to be tested;
[0080] The corrected F-score of the specified pathway of the test cell is determined by a preset algorithm.
[0081] Based on the corrected F score, cell subclass clustering information is obtained from the cells to be tested;
[0082] The cell subclass data obtained from the cells to be tested is input into a pre-created single-cell pathway enrichment module. The cell subclass data includes the original gene expression matrix, the list of hypervariable genes, and the cluster information of the cell subclass.
[0083] The clusterProfiler tool in the single-cell pathway enrichment module is invoked, and the list of hypervariable genes in the cell subclass data is used as input data. GO (Gene Ontology) enrichment analysis is performed to obtain the enrichment analysis results, which include GO enrichment results, FDR p-value, and gene information contained in the pathway.
[0084] The preset algorithm is as follows:
[0085] Among them, bg ratio Refers to the background value, N genes in a cell N refers to the number of genes detected in the cells being tested. total genes The total number of genes detected in the sample, fg ratio Foreground value; Ngenes from a term in a cell refers to the number of genes in the specified pathway detected in the cell being tested; Ntotal genes of a term refers to the total number of genes in the specified pathway; F score The F score of the specified pathway in the test cells, adjF score The corrected F-score, Q, refers to the score of the specified pathway in the test cells. term The p-value, Q, is the value corrected by FDR. term Calculated using the clusterProfiler tool.
[0086] Understandably, in formula (1), the number of genes (N) detected in the test cell is... genes in a cell Divide by the total number of genes detected in the sample (N) total genes ) as background value (bg) ratio ).
[0087] In formula (2), the number of genes from a term in a cell detected in the test cell (Ngenes from a term in a cell) divided by the total number of genes in the term (Ntotal genes of a term) is used as the foreground value (fg). ratio ).
[0088] In formula (3), the foreground value (fg) ratio Divide by the background value (bg)ratio The F score (F) is used as the designated pathway for the cells under test. score ).
[0089] In formula (4), F score (F score Divide by Q value (Q term The logarithm is taken as the corrected F-score (adjF) for the specified pathway in the test cells. score Understandably, the Q-value (Q) term The p-value is the FDR-corrected value, calculated by clusterProfiler in a conventional manner, without any specific limitations.
[0090] In this embodiment, to perform "single-cell pathway enrichment" analysis, the user needs to prepare cell subclass data, including: a raw gene expression matrix (Count Matrix), a list of hypervariable genes (HVG markers), and cell subclass cluster information (Clusters & Annotated Clusters). Specifically, the user only needs to select the cell type of interest from the cell type annotation results in the second dataset as the cell to be tested and perform cell subclass analysis.
[0091] The original gene expression matrix and the list of highly variable genes are obtained in a conventional manner. The method for obtaining the cell subclass cluster information is described in the relevant implementation process of the above-mentioned preset algorithm, and will not be repeated here.
[0092] When performing pathway enrichment analysis, the obtained original gene expression matrix, list of hypervariable genes, and cell subclass clustering information can be used as input data and input into the single-cell pathway enrichment module for analysis and processing.
[0093] Before inputting cell subclass data into the single-cell pathway enrichment module, preprocessing is typically required. For example, starting with the selected cell type of interest, the analysis platform extracts the corresponding cells (i.e., the cells to be tested) from the initial expression matrix to construct a new expression matrix. After completing data subset extraction and new matrix construction, this preprocessing process may include:
[0094] Cells with fewer than 3 genes expressed and fewer than 200 genes expressed were filtered out.
[0095] Cells with a detected mitochondrial gene ratio higher than 10% were filtered out;
[0096] Use SCTransform to standardize the filtered data;
[0097] The VST algorithm was used to extract the top 2000 genes by expression level in SD as hypervariable genes for PCA, and PCA dimensionality reduction was performed.
[0098] The top 10 PCs from the PCA dimensionality reduction results were selected for UMAP dimensionality reduction and Louvain clustering to obtain cluster information;
[0099] Cell type annotation was performed based on SingleR and Louvain clustering to obtain annotated clusters information;
[0100] The genes involved in the top N PCs are extracted from PCA to form a list of hypervariable genes, where N can be flexibly controlled by the user.
[0101] In this embodiment, by combining the enrichment analysis results of single-cell pathway enrichment analysis with the prediction results of cell development trajectory (calling the monocle2 results of the corresponding data subset, i.e. the dimensionality reduction part of DDRTree), it is possible to study the changes in biological process pathways and the expression of specific genes along the development trajectory under different states.
[0102] The analysis platform can support mapping the corrected F score of each cell to a dimensionality-reduced space, such as UMAP and DDRTree.
[0103] As an optional implementation, the method may further include:
[0104] The communication analysis results and the annotation results are stored in a pre-created temporary project repository.
[0105] Understandably, a temporary project repository can store and back up datasets generated during cell data analysis, making it convenient for users to view them later. When a user needs to view data in the temporary project repository, they can enter a query command through the web interface of their electronic device. The device can then retrieve the dataset the user is looking for from the temporary project repository via the analysis platform.
[0106] In this embodiment, the processing module can be an integrated circuit chip with signal processing capabilities. The processing module can be a general-purpose processor. For example, the processor can be a Central Processing Unit (CPU), a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application.
[0107] The storage module can be, but is not limited to, random access memory, read-only memory, programmable read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, etc. In this embodiment, the storage module can be used to store the first dataset, the second dataset, communication analysis results, and annotation results, etc. Of course, the storage module can also be used to store programs, which the processing module executes after receiving execution instructions.
[0108] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the electronic device described above can be referred to the corresponding steps in the aforementioned method, and will not be elaborated further here.
[0109] This application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program that, when run on a computer, causes the computer to perform the single-cell transcriptome data analysis and processing method as described in the above embodiments.
[0110] Based on the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by hardware or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application can be embodied in the form of a software product. This software product can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as a personal computer, electronic device, or network device, etc.) to execute the methods described in the various implementation scenarios of this application.
[0111] In the embodiments provided in this application, it should be understood that the disclosed apparatus, systems, and methods can also be implemented in other ways. The apparatus, systems, and methods embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which includes one or more executable instructions for implementing a specified logical function. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0112] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for analyzing and processing single-cell transcriptome data, characterized in that, The method includes: Obtain the first dataset after preprocessing, wherein the first dataset is a dataset generated based on single-cell omics, and the data formats in the first dataset include tsv, txt, csv, RDS and HDF5 formats; Upon receiving the first instruction for intercellular communication analysis, the pre-set CellPhoneDB, CellChat, and Cellcall tools are invoked from the tool library. The strength of intercellular interactions and the difference pairs of ligand-receptor in the first dataset are determined by the CellPhoneDB, CellChat, and Cellcall tools to obtain communication analysis results characterizing intercellular communication. Upon receiving a second instruction for cell type annotation, a pre-set dimensionality reduction clustering tool is invoked from the tool library to perform dimensionality reduction clustering on the first dataset, resulting in a clustered second dataset. The dimensionality reduction clustering tool includes PCA, SC3, tSNE, and UMAP tools. The second dataset is annotated with cell types by calling the annotation tools in the tool library to obtain the annotation results. The annotation tools include any one of the SingleR tool and a graph annotation tool based on the weighted nearest neighbor network algorithm. Based on a preset visualization strategy, the communication analysis results and the annotation results are visualized.
2. The method according to claim 1, characterized in that, Before obtaining the preprocessed first dataset, the method further includes: Obtain the software packages corresponding to the CellPhoneDB tool, the CellChat tool, and the Cellcall tool; Based on the graphical user interface (GUI) and Shiny, the corresponding software packages for the CellPhoneDB tool, the CellChat tool, and the Cellcall tool are added to the pre-created tool library.
3. The method according to claim 1, characterized in that, Before obtaining the preprocessed first dataset, the method further includes: Obtain the software packages corresponding to the PCA tool, the SC3 tool, the tSNE tool, and the UMAP tool; Based on the graphical user interface (GUI) and Shiny, the software packages corresponding to the PCA tool, the SC3 tool, the tSNE tool, and the UMAP tool are added to the pre-created tool library.
4. The method according to claim 1, characterized in that, The method further includes: Upon receiving a third instruction for trajectory prediction, a pre-set trajectory prediction tool is invoked from the tool library, and the developmental trajectory of cells in the first dataset is determined by the trajectory prediction tool, wherein the trajectory prediction tool includes the Monocle2 tool or the SPRING tool.
5. The method according to claim 1, characterized in that, The method further includes: The communication analysis results and the annotation results are stored in a pre-created temporary project repository.
6. The method according to claim 1, characterized in that, The tool library is used to call a pre-set dimensionality reduction clustering tool to perform dimensionality reduction clustering on the first dataset, resulting in a clustered second dataset, including: The PCA tool is called from the tool library, and the PCA tool is used to reduce the dimensionality of the first dataset to obtain the dimensionality-reduced first dataset. The SC3 tool is called from the tool library, and the SC3 tool is used to cluster the dimensionality-reduced first dataset to obtain the second dataset; The second dataset is visualized by calling the tSNE tool or the UMAP tool from the tool library.
7. The method according to claim 1, characterized in that, By invoking the annotation tools in the aforementioned tool library, cell type annotations are performed on the second dataset to obtain the annotation results, including: By calling the SingleR tool, cell type annotation is performed on the cells in the second dataset based on the transcriptome data of the reference cell type, and the annotation results are obtained. Alternatively, by invoking the map annotation tool, based on the weighted nearest neighbor network algorithm, the scRNA-seq data corresponding to cells in the second dataset can be mapped onto a pre-annotated single-cell transcriptome map to obtain the annotation results.
8. The method according to claim 1, characterized in that, Based on a preset visualization strategy, the communication analysis results and the annotation results are visualized and displayed, including: Generate icons for any two cells that are communicating in the communication analysis results, and connect the icons of any two cells that are communicating with each other by a line segment; Scatter points are generated for each cell in the second dataset. Based on the cell type in the second data in the annotation results, the rendering color corresponding to the scatter points of different cell types is determined. The scatter points of the corresponding cell types are rendered based on the rendering color, wherein the rendering color of the scatter points of different cell types is different.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory coupled together, the memory storing a computer program that, when executed by the processor, causes the electronic device to perform the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Cell subset annotation method based on single cell transcriptome sequencing
CN112700820A
Analysis method, device and equipment based on single cell transcriptome sequencing data
CN116189764A
Single cell transcriptome data analysis processing method and electronic equipment
CN116913388A
Single cell transcriptome data preprocessing method, electronic equipment and storage medium
CN116913389A
Cell data comprehensive analysis system
CN116959583A