A method and system for identifying breast cancer cell subgroups, and identifying characteristics of metastatic or drug-resistant subtype tumor cells
By using correlation calculations and dimensionality reduction clustering with breast cancer tissue cell-specific marker datasets, combined with EMT and taxane resistance gene analysis, the accuracy and standardization issues of cell population identification in breast cancer single-cell transcriptome sequencing data in existing technologies have been resolved, enabling precise identification of metastatic or drug-resistant tumor cells and personalized treatment guidance.
Patent Information
- Application Number
- CN202410966873.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-07-18
AI Technical Summary
Existing methods for identifying cell populations using single-cell transcriptome sequencing data in breast cancer suffer from problems such as inaccurate classification, high classification costs, and inconsistent classification standards, making it difficult to accurately identify metastatic or drug-resistant characteristic tumor cells.
By calculating the correlation between clustered cell populations and breast cancer tissue cell-specific marker datasets, and combining dimensionality reduction and clustering techniques, the cell types corresponding to each clustered cell population are identified. Differential expression analysis is then performed using epithelial-mesenchymal transition and taxane resistance gene characteristics, and a visualization system is constructed for identification.
It improves the accuracy and efficiency of cell population identification in breast cancer single-cell transcriptome sequencing data, provides a unified classification standard, and can more accurately identify EMT and taxane-resistant tumor cells, guiding personalized treatment.
Smart Images

Figure CN118942555B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of bio-information analysis, and more particularly to a breast cancer cell cluster identification method and system, and a metastasis or drug resistance subtype tumor cell feature identification method. BACKGROUND
[0002] With the development of biological technology, today's single-cell transcriptome sequencing technology can perform high-throughput sequencing of the transcriptome at the single-cell level, and can obtain the transcriptome expression data of several thousand to tens of thousands of cells at a time. Due to the heterogeneity and complexity of cells in breast cancer tissue, single-cell transcriptome sequencing is widely used in breast cancer. By analyzing the single-cell transcriptome sequencing data of breast cancer tissue, the molecular characteristics and biological functions of different cells in the tumor tissue of breast cancer patients can be revealed, which is beneficial to exploring the potential pathogenesis and individualized treatment targets of patients.
[0003] According to the data of the International Agency for Research on Cancer and the World Health Organization in 2023, there were 2,752,000 new cases of breast cancer and 684,996 deaths from breast cancer worldwide in 2023. Tumor metastasis and treatment resistance are the main reasons for the progression and death of breast cancer patients. Epithelial-mesenchymal transition (EMT) is a biological process that makes epithelial cells lose polarity and intercellular connections, enabling breast cancer cells to migrate from the primary lesion to distant organs, and is an important process of tumor metastasis. Taxanes inhibit the division and proliferation of cancer cells by blocking microtubule polymerization, and are widely standardized into standard treatment regimens for breast cancer, especially for locally advanced or metastatic breast cancer. However, patients often develop resistance to taxanes during treatment, affecting treatment efficacy and prognosis.
[0004] By identifying the molecular characteristics of metastasis or drug resistance tumor subpopulations in breast cancer patients through single-cell transcriptome sequencing, the differentially expressed genes of metastasis or drug resistance tumor cells and their enriched biological processes can be obtained, which can guide treatment strategies and more accurately select appropriate treatment regimens, reducing the risk of metastasis and drug resistance. The basic process of single-cell transcriptome sequencing data analysis includes quality control of sequencing data, standardization of data, identification of high-variable genes, data dimensionality reduction, cell clustering, cell cluster identification, determination of characteristic cell subpopulation differential genes, and differential gene enrichment pathway analysis. Among them, cell cluster identification is a key step, and incorrect cell cluster identification can lead to incorrect judgment of the function, state and physiological significance of the cell, misleading the subsequent biological research direction and conclusion.
[0005] The most commonly used method for identifying cell clusters from single-cell transcriptome sequencing data is the standardized classification method based on reference database and machine learning. However, such methods have some problems. First, the coverage and quality of the reference database cannot cover all cell types or newly discovered cell subtypes, and the annotation accuracy for rare cell types is low. Second, the selection of model structure and parameters has a significant impact on the results, and a large amount of labeled data is needed to train the model, which has a high threshold for use. Another commonly used method for identifying cell clusters from single-cell transcriptome sequencing data is to manually associate each cell cluster with a known cell type based on cell-specific marker genes. However, this method also has some problems. First, cell-specific marker genes often rely on other studies and have certain limitations. Second, noise signals and rare cell types are often present in single-cell sequencing data and are easily ignored or confused with other types in clustering analysis. Third, manual annotation of cell types is easily influenced by researchers' subjective understanding and experience, and it is difficult to determine the expression level of cell-specific marker genes. The process is tedious, time-consuming, and inefficient.
[0006] In summary, the existing cell cluster identification methods have the problems of inaccurate classification, high classification cost, and non-uniform classification standards. There is a need to identify the molecular characteristics of tumor subpopulations with metastasis or drug resistance characteristics in breast cancer patients. SUMMARY
[0007] To overcome the shortcomings of existing cell cluster identification techniques and meet the demand for identifying the molecular characteristics of breast cancer tumor cells with metastasis or drug resistance characteristics, the present application provides a method and system for identifying breast cancer single-cell transcriptome sequencing data cell clusters and epithelial-mesenchymal transition (EMT) or paclitaxel drug resistance characteristic subtype tumor cells. The present application correlates each cluster cell group with each cell type in the breast cancer tissue cell-specific marker dataset to determine the corresponding cell type of each cluster cell group. The present application solves the problems of inaccurate classification, high classification cost, and non-uniform classification standards in existing cell cluster identification techniques, and provides a visualization system for cell cluster annotation based on breast cancer single-cell transcriptome sequencing data and identification of EMT or paclitaxel drug resistance characteristic tumor subpopulation molecular characteristics.
[0008] According to a first aspect of the present application, a method for identifying breast cancer cell clusters is provided, comprising the following steps:
[0009] (1) Dimensionality reduction and clustering of single-cell sequencing expression matrix data to obtain cluster cell groups;
[0010] (2) calculating the correlation of each clustered cell population with each cell type in the breast cancer tissue cell-specific marker dataset to determine the cell type corresponding to each clustered cell population.
[0011] Preferably, the breast cancer tissue cell-specific marker dataset is constructed by the following steps:
[0012] S1: obtaining marker genes of basic cells in breast cancer tissue, the basic cells in breast cancer tissue are 15, which are epithelial cells, endothelial cells, fibroblasts, tumor-associated fibroblasts, T cells, Treg cells, B cells, NK cells, NKT cells, macrophages, monocytes, dendritic cells, mast cells, tumor cells and tumor stem cells;
[0013] S2: sorting the marker genes of each basic cell obtained in step S1, the sorting is based on the sorting in the CellMarker2.0 database, obtaining the sorting from high to low importance, and obtaining the top 10 marker genes of each basic cell;
[0014] S3: obtaining the weight of each marker gene of each basic cell by counting the number of times each marker gene of each basic cell appears in the CellMarker2.0 database / the total number of times the top 10 marker genes of each basic cell appear in the CellMarker2.0 database;
[0015] S4: the basic cells of breast cancer tissue obtained in step S1 and the marker genes of each basic cell, and the weight of each marker gene of each basic cell obtained in step S3, constructing the breast cancer tissue cell-specific marker dataset.
[0016] Preferably, the step (2) specifically comprises the following steps:
[0017] S1: finding the top 10 marker genes of each of the 15 basic cells in the breast cancer tissue cell-specific marker dataset in the clustered cell population, and obtaining the average expression level of each marker gene in each clustered cell population;
[0018] S2: normalizing the average expression level of the top 10 marker genes of each of the 15 basic cells in each clustered cell population, respectively, to obtain the expression score of the top 10 marker genes of each basic cell type in each clustered cell population, respectively;
[0019] S3: for each cell type, multiplying the expression score of the top 10 marker genes in each clustered cell population by the weight of the top 10 marker genes, and then summing up to obtain the classification score of each cell population;
[0020] S4: Label the cell type corresponding to the highest classification score of each cluster cell population as the cell type of the respective cluster cell population.
[0021] Preferably, the dimensionality reduction is preceded by quality control and normalization processing.
[0022] According to another aspect of the present application, a method for identifying breast cancer cell metastasis or drug resistance subtype tumor cell characteristics is provided, comprising the following steps:
[0023] (1) Calculate the relative expression amount of epithelial-mesenchymal transition marker genes or taxane drug marker genes in the breast cancer tumor cell population;
[0024] (2) Extract and set cells with a relative expression amount of epithelial-mesenchymal transition marker genes greater than 0 as epithelial-mesenchymal transition marker gene characteristic tumor cells, and extract and set cells with a relative expression amount of epithelial-mesenchymal transition marker genes less than 0 as non-epithelial-mesenchymal transition marker gene characteristic tumor cells; extract and set cells with a taxane drug resistance expression amount greater than 0 as taxane drug resistance characteristic tumor cells, and extract and set cells with a taxane drug resistance expression amount less than 0 as non-taxane drug resistance characteristic tumor cells;
[0025] (3) Separate the epithelial-mesenchymal transition and non-epithelial-mesenchymal transition characteristic tumor cells, and obtain the epithelial-mesenchymal transition score violin plot of both; separate the taxane drug resistance and non-taxane drug resistance characteristic tumor cells, and obtain the taxane drug resistance score violin plot of both;
[0026] (4) Identify the differentially expressed genes of the epithelial-mesenchymal transition and non-epithelial-mesenchymal transition characteristic tumor cells, and identify the differentially expressed genes of the taxane drug resistance characteristic and non-taxane drug resistance characteristic tumor cells;
[0027] (5) Screen a number of genes with significant differences in the differentially expressed genes obtained in step (4) for KEGG or GO enrichment analysis, and obtain the enrichment pathways and visual bubble chart.
[0028] According to another aspect of the present application, a breast cancer cell cluster identification system is provided, comprising:
[0029] A cluster cell population module: the cluster cell population module is used for dimensionality reduction and clustering of single cell sequencing expression matrix data;
[0030] A cell cluster identification module: the cell cluster identification module is used for correlation calculation of each cell type in the cluster cell population and the breast cancer tissue cell-specific marker data set, and to determine the corresponding cell type of each cluster cell population.
[0031] According to another aspect of the present application, a breast cancer cell metastasis or drug resistance subtype tumor cell feature recognition system is provided, comprising:
[0032] The computing module is configured to calculate the relative expression amount of the epithelial-mesenchymal transition marker gene or the taxane drug marker gene in the breast cancer tumor cell population;
[0033] The extraction and setting module is configured to extract and set the cells with the relative expression amount of the epithelial-mesenchymal transition marker gene greater than 0 as the epithelial-mesenchymal transition marker gene characteristic tumor cells, and extract and set the cells with the relative expression amount of the epithelial-mesenchymal transition marker gene less than 0 as the non-epithelial-mesenchymal transition marker gene characteristic tumor cells; extract and set the cells with the taxane drug resistance expression amount greater than 0 as the taxane drug resistance characteristic tumor cells, and extract and set the cells with the taxane drug resistance expression amount less than 0 as the non-taxane drug resistance characteristic tumor cells;
[0034] The score violin plot acquisition module is configured to separate the epithelial-mesenchymal transition and non-epithelial-mesenchymal transition characteristic tumor cells, and acquire the epithelial-mesenchymal transition score violin plot of both; separate the taxane drug resistance and non-taxane drug resistance characteristic tumor cells, and acquire the taxane drug resistance score violin plot of both;
[0035] The recognition module is configured to recognize the differentially expressed genes of the epithelial-mesenchymal transition and non-epithelial-mesenchymal transition characteristic tumor cells, and recognize the differentially expressed genes of the taxane drug resistance characteristic and non-taxane drug resistance characteristic tumor cells;
[0036] The enrichment module is configured to perform KEGG or GO enrichment analysis on a number of genes with significant differences in the obtained differentially expressed genes, and acquire the enrichment pathways and visual bubble chart.
[0037] Overall, compared with the prior art, the above technical solutions conceived by the present application mainly have the following technical advantages:
[0038] The present application provides a visualization system for cell cluster annotation and EMT or taxane drug resistance characteristic tumor subgroup molecular feature recognition based on breast cancer single cell transcriptome sequencing data, solves the problems of inaccurate classification, high classification cost and non-uniform classification standard in existing cell cluster identification techniques, and meets the demand for breast cancer metastasis or drug resistance characteristic tumor cell molecular feature recognition. BRIEF DESCRIPTION OF DRAWINGS
[0039] Figure 1 A breast cancer single cell transcriptome sequencing data cell cluster identification and EMT or taxane drug resistance characteristic tumor subgroup molecular feature recognition method flowchart is provided for the embodiments of the present application.
[0040] Figure 2 The internal structure diagram of the system provided by the embodiment of the application is shown.
[0041] Figure 3 The visual webpage display interface of the application. DETAILED DESCRIPTION
[0042] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.
[0043] The purpose of the present application is to provide a breast cancer single cell transcriptome sequencing data cell cluster identification and EMT or taxane drug resistance characteristic subtype tumor cell recognition method and system, which solves the problems of inaccurate classification, high classification cost and non-uniform classification standard in existing cell cluster identification technology, and provides a visualization system for cell cluster annotation based on breast cancer single cell transcriptome sequencing data and EMT or taxane drug resistance characteristic tumor subtype molecular feature recognition.
[0044] The breast cancer single cell transcriptome sequencing data cell cluster identification method of the present application comprises the following steps:
[0045] Step 1, obtaining breast cancer single cell sequencing data, preprocessing to obtain a gene expression matrix;
[0046] Step 2, quality control, standardization processing, dimensionality reduction and clustering of the single cell sequencing data to obtain clustered cell cluster data;
[0047] Step 3, correlation calculation of the clustered cell cluster data and each cell type in the constructed breast cancer tissue cell-specific marker data set to obtain the cell type closest to the clustered cell cluster;
[0048] Step 4, cell cluster annotation based on the closest cell type;
[0049] Further, in step 1, the preprocessing includes quality control of the original sample data, alignment with the reference genome, quantitative analysis and single cell gene expression matrix construction;
[0050] Further, in step 2, the quality control includes setting a gene threshold for a single cell to remove non-single cells and setting a mitochondrial percentage threshold for a single cell to remove cells with poor activity;
[0051] Further, the normalization in step 2 includes eliminating noise caused by uneven sequencing depth, cell size difference, batch effect, etc., ensuring that the expression values of different genes have similar numerical ranges, and normalizing the gene expression values of each single cell using the total gene expression of all cells;
[0052] Further, the dimension reduction and clustering method in step 2 includes reducing the dimension of the data and finding genes with significant expression changes in different cell states through principal component analysis, constructing cell neighborhood relationships by calculating the Euclidean distance between each cell and all other cells, and dividing cells into different groups according to the cell neighborhood relationships and based on graph-based clustering algorithm, K-means clustering and spectral clustering method;
[0053] Further, in step 3, the breast cancer tissue cell-specific marker dataset includes:
[0054] The basic fifteen cell types of breast cancer tissue: epithelial cells, endothelial cells, fibroblasts, tumor-associated fibroblasts, T cells, Treg cells, B cells, NK cells, NKT cells, macrophages, monocytes, dendritic cells, mast cells, tumor cells and tumor stem cells;
[0055] The marker genes of each cell type-specific marker gene that appear in the cancer tissue sample (except for epithelial cells), and are ranked from high to low according to the number of times they appear in the literature as cell type marker genes;
[0056] The top 10 marker genes of each cell type (less than 10 replaced by NA);
[0057] Each marker gene as a classification weight of cell type (assigning weight to the number of times the marker gene appears in the literature as a cell type marker gene);
[0058] Further, in step 3, the process of calculating the correlation between the clustered cell group data and each cell type in the dataset includes:
[0059] Find the first to tenth marker genes of the basic fifteen cell types of breast cancer tissue in the data set in the clustered cell data frame respectively;
[0060] Get the average expression level of the first to tenth marker genes of the fifteen cell types in each cell group;
[0061] 0-1 normalization is performed on the average expression level of the first to tenth marker genes of the fifteen cell types in each cell group respectively;
[0062] For each cell type, multiply the expression score of its 1st to 10th marker genes in each cell group with its classification weight respectively and sum them up;
[0063] Return the fifteen cell class classification scores of each cell group;
[0064] Further, in step 4, the cell class group annotation based on the closest cell type is to mark the cell type corresponding to the highest cell class group classification score of each cell group as the cell type of the respective cell group.
[0065] The application discloses a method for identifying molecular characteristics of a tumor subpopulation with epithelial-mesenchymal transition (EMT) or paclitaxel drug resistance of breast cancer tumor cells, which comprises the following steps:
[0066] Step 1: obtaining a tumor cell gene expression data frame in breast cancer single cell sequencing data which has been subjected to cell class group identification;
[0067] Step 2: grouping and clustering the tumor cells;
[0068] Step 3: scoring the tumor cells by using EMT marker genes and paclitaxel drug resistance marker genes respectively;
[0069] Step 4: separating EMT characteristic tumor cells and non-EMT characteristic tumor cells, and separating paclitaxel drug resistance tumor cells and non-paclitaxel drug resistance tumor cells;
[0070] Step 5: identifying differential expression genes (DEGs) of EMT and non-EMT characteristic tumor cells and DEGs of paclitaxel drug resistance and non-paclitaxel drug resistance tumor cells respectively;
[0071] Step 6: performing enrichment analysis on the DEGs by using biological function annotation databases KEGG and GO respectively;
[0072] Further, in step 1, the tumor cell gene expression data frame is extracted from the tumor cell or tumor stem cell group in the breast cancer single cell sequencing data frame which has been classified into tumor cells;
[0073] Further, in step 2, the method for grouping and clustering the tumor cells comprises constructing cell neighbor relationships by calculating the Euclidean distance between each tumor cell and all other tumor cells, and grouping the tumor cells into different groups according to the cell neighbor relationships and based on graph-based clustering algorithms, K-means clustering and spectral clustering methods;
[0074] Further, in step 3, the tumor cell EMT or paclitaxel drug resistance marker genes are collected and obtained through literature and cancer treatment response gene marker database (CTR-DB);
[0075] Further, the process of scoring the tumor cell population in step 3 comprises:
[0076] Defining the EMT marker gene module of tumor cells and the taxane drug resistance gene module of tumor cells;
[0077] Randomly extracting the expression values of 200 genes outside the module in the tumor cells as background values, and calculating the average expression values of all genes in the tumor cells in the module;
[0078] Further, in step 4, the EMT gene module with a score higher than 0 is identified as an EMT characteristic tumor cell, the EMT gene module with a score lower than 0 is identified as a non-EMT characteristic tumor cell, the taxane drug resistance gene module with a score higher than 0 is identified as a taxane drug resistance tumor cell, and the taxane drug resistance gene module with a score lower than 0 is identified as a non-taxane drug resistance tumor cell.
[0079] Preferably, the 200 DEGs obtained in step 5 are screened for KEGG or GO analysis.
[0080] The breast cancer single-cell transcriptome sequencing data cell cluster identification and EMT or taxane drug resistance subtype tumor cell characteristic recognition system comprises:
[0081] An algorithm server and an uploading, parameter selection and display webpage;
[0082] The algorithm server is used to store algorithms and calculate breast cancer single-cell transcriptome sequencing data cell clusters and EMT or taxane drug resistance characteristic subtype tumor cell recognition, to obtain breast cancer single-cell transcriptome sequencing cell cluster identification results and EMT or taxane drug resistance characteristic subtype tumor cell recognition results.
[0083] The uploading, parameter selection and display webpage are integrated in the same website and are respectively used for uploading single-cell transcriptome sequencing data, selecting a result observation interface and displaying calculation results.
[0084] Further, the algorithm server comprises a data transmission and R script calling module, an R script calculation module and a result storage and calling module.
[0085] The data transmission and R script calling module is used to transmit breast cancer single-cell sequencing data uploaded by a user to a backend and call a cell cluster identification and EMT or taxane drug resistance subtype tumor cell characteristic recognition calculation method in an R script;
[0086] The R script calculation module is used to perform cell cluster identification and EMT or taxane drug resistance subtype tumor cell characteristic recognition calculation.
[0087] The result storage and calling module is used for storing the result of the R script calculation and facilitating the user to call the result display file.
[0088] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0089] As shown in the accompanying drawings Figure 1 The present application provides a breast cancer single cell transcriptome sequencing data cell cluster identification and EMT or taxane drug resistance characteristic subtype tumor cell identification method, comprising:
[0090] Step1: uploading breast cancer single cell sequencing expression matrix data to scBCcelltype software
[0091] In this embodiment, first, the supplementary files of breast cancer patient breast tissue single cell sequencing data (GEO access number, GSM5354528) are downloaded from the Gene Expression Omnibus (GEO) database, and the count_matrix_barcodes.tsv, count_matrix_genes.tsv and count_matrix_sparse.mtx files are obtained after decompression, preferably, the above three files are compressed and renamed as barcodes.tsv.gz, features.tsv.gz and matrix.mtx.gz respectively using 7-Zip software, and the three files are uploaded to a visualization webpage.
[0092] Step2: quality control, standardization processing and dimensionality reduction clustering of the uploaded breast cancer single cell sequencing expression matrix
[0093] In this embodiment, the R package mainly used is Seurat. First, an analysis data frame is created using CreateSeuratObject, named "Object". Non-single cells and cells with poor status are removed by setting thresholds for the number of genes in each cell and the percentage of mitochondrial genes, preferably with the following parameters: nFeature_RNA>200&nFeature_RNA<6000&percent.mt<15&nCount_RNA>400. The single cell gene expression level is standardized using the NormalizedData and ScaleData algorithms. The PCA algorithm is used to reduce the dimension of the data, the FindNeighbors algorithm is used to calculate the Euclidean distance between each cell and all other cells to construct the cell neighborhood relationship, and the FindClusters algorithm is used to divide the cells into 8 cell groups, cluster0-7, and the UMAP (Uniform Manifold Approximation and Projection) or TSNE (t-Distributed Stochastic Neighbor Embedding) algorithm is used to obtain the distribution of cell groups in the UMAP or TSNE dimension reduction space.
[0094] Step 3: Determine the cell type of the clustered cell group
[0095] The construction process of the breast cancer tissue cell-specific marker data set of the application is as follows:
[0096] Search for marker genes of epithelial cells, endothelial cells, fibroblasts, tumor-associated fibroblasts, T cells, Treg cells, B cells, NK cells, NKT cells, macrophages, monocytes, dendritic cells, mast cells, tumor cells, and tumor stem cells in breast tissue in the CellMarker2.0 database;
[0097] Screen the marker genes of each cell type that appear in the cancer tissue sample (except for epithelial cells);
[0098] Sort the marker genes of each cell type according to the order in the CellMarker2.0 database, and obtain the ranking from high to low importance;
[0099] Select the top 10 marker genes of each cell type, and if there are less than 10 marker genes for a cell type, replace them with NA.
[0100] According to the number of times the marker gene appears as a cell type marker gene in the CellMarker2.0 database, assign it a weight.
[0101] The method of correlating the clustered cell population data with each cell type in the above constructed data set is constructed as follows:
[0102] First, a cell cluster scoring function is constructed for the clustered cell population:
[0103] The input parameters of the function are set, including: the analysis data frame that has been clustered, the number of cell clusters, the cell type, the top1 marker gene (gene1), the top2 marker gene (gene2), …, the top10 marker gene (gene10), the weight of the top1 marker gene, the weight of the top2 marker gene, …, the weight of the top10 marker gene, specifically: data, cluster_num, celltype, gene1, gene2, gene3, gene4, gene5, gene6, gene7, gene8, gene9, gene10, weight1, weight2, weight3, weight4, weight5, weight6, weight7, weight8, weight9, weight10;
[0104] Each cell cluster gene1 score calculation includes: finding gene1 in single cell sequencing data, if there is gene1 expression, then using a loop statement to calculate the average expression level of gene1 of the cell type in each cell cluster, and then through 0-1 normalization to get the expression score of gene1 in each cell cluster, if there is no gene1 expression, the expression score of gene1 is 0, specifically: desired_feature <- gene1
[0105] if (desired_feature %in% rownames(data@assays$RNA@features)){
[0106] gene1 <- VlnPlot(data, features = gene1, group.by = "seurat_clusters", log = T, ncol = 1)[[1]][["data"]]
[0107] gene1_score <- rep(0, cluster_num)
[0108] cell_count <- rep(0, cluster_num)
[0109] for (i in 1:nrow(gene1)){
[0110] gene1_score[as.numeric(gene1[i,][2])] = gene1_score[as.numeric(gene1[i,][2])] + as.numeric(gene1[i,][1])
[0111] cell_count[as.numeric(gene1[i,][2])] = cell_count[as.numeric(gene1[i,][2])] + 1
[0112] }
[0113] gene1_score = gene1_score / cell_count
[0114] gene1_score_normalized = (gene1_score - min(gene1_score)) / (max(gene1_score) - min(gene1_score))
[0115] }else{
[0116] gene1_score_normalized = 0
[0117] };
[0118] Each cell group gene2-gene10 score calculation process is the same as gene1;
[0119] The expression score of each cell group gene1-gene10 in each cell group is multiplied by its classification weight and summed to obtain the score of each cell type in each cell group, which is:
[0120] score = gene1_score_normalized*weight1 + gene2_score_normalized*weight2 + gene3_score_normalized*weight3 + gene4_score_normalized*weight4 + gene5_score_normalized*weight5 + gene6_score_normalized*weight6 +
[0121] gene7_score_normalized*weight7 + gene8_score_normalized*weight8 +
[0122] gene9_score_normalized * weight9 + gene10_score_normalized * weight10;
[0123] Returning the score of the cell type for each cell population;
[0124] Next, the cell class scoring is performed using the above functions, in particular:
[0125] celltype_score <- data.frame(marker_gene(Object, num_clusters, "Epithelial_cells", "CDH1", "CD117", "CD166", "KRT14", "KRT18", "DSP", "EPCAM", "GCDFP15", "MCAM", "TP63", 7 / 16, 1 / 16, 1 / 16, 1 / 16, 1 / 16, 1 / 16, 1 / 16, 1 / 16, 1 / 16, 1 / 16),
[0126] marker_gene(Object, num_clusters, "Endothelial_cells", "PECAM1", "CD31", "VWF", "IL3RA", "CD34", "CDH5", "CLEC14A", "EMCN", "MCAM", "PECAM 1", 5 / 21, 5 / 21, 1 / 7, 2 / 21, 1 / 21, 1 / 21, 1 / 21, 1 / 21, 1 / 21, 1 / 21),
[0127] marker_gene(Object, num_clusters, "Fibroblasts", "DCN", "PDGFRB", "ACTA2", "COL1A1", "COL1A2", "ITGB1", "PDGFRA", "S100A4", "TAGLN", "THY1", 1 / 6, 1 / 6, 1 / 12, 1 / 12, 1 / 12, 1 / 12, 1 / 12, 1 / 12, 1 / 12, 1 / 12),
[0128] marker_gene(Object, num_clusters, "CAF", "ACTA2", "COL1A1", "PDGFR A", "PDGFRB", "PDPN", "FAP", "THY1", "NA", "NA", "NA", 1 / 6, 1 / 6, 1 / 6, 1 / 6, 1 / 6, 1 / 12, 1 / 12, 0, 0, 0),
[0129] marker_gene (Object, num_clusters, "T_cells", "CD3D", "CD3G", "CD4", "CD2", "CD3", "CD3E", "CD8", "CD137", "CD28", "CD8A", 5 / 22, 3 / 22, 3 / 22, 1 / 11, 1 / 11, 1 / 11, 1 / 11, 1 / 22, 1 / 22, 1 / 22),
[0130] marker_gene (Object, num_clusters, "Treg_cells", "FOXP3", "CCR8", "IL2RA", "BATF", "CCR7", "CD3D", "CD3G", "KLRB1", "LAYN", "MAGEH1", 6 / 17, 2 / 17, 2 / 17, 1 / 17, 1 / 17, 1 / 17, 1 / 17, 1 / 17, 1 / 17, 1 / 17),
[0131] marker_gene (Object, num_clusters, "B_cell", "CD20", "CD79A", "MS4A1", "CD19", "BANK1", "CD14", "CD27", "CD38", "CD40", "CD74", 7 / 24, 1 / 6, 1 / 6, 1 / 8, 1 / 24, 1 / 24, 1 / 24, 1 / 24, 1 / 24, 1 / 24),
[0132] marker_gene (Object, num_clusters, "NK_cell", "KLRB1", "KLRC1", "KLRD 1", "NKG7", "CCR7", "CD3D", "CD3G", "CD8A", "GNLY", "KIR2DL1", 1 / 7, 1 / 7, 1 / 7, 1 / 7, 1 / 14, 1 / 14, 1 / 14, 1 / 14, 1 / 14, 1 / 14),
[0133] marker_gene (Object, num_clusters, "NKT_cells", "CD3D", "GNLY", "KLRD 1", "NCR1", "NA", "NA", "NA", "NA", "NA", "NA", 1 / 4, 1 / 4, 1 / 4, 1 / 4, 0, 0, 0, 0, 0, 0),
[0134] marker_gene (Object, num_clusters, "Macrophage", "CD68", "CD14", "CD163", "CD209", "CD33", "CD4", "CD86", "CDH5", "CSF1", "FCGR3A", 4 / 13, 1 / 13, 1 / 13, 1 / 13, 1 / 13, 1 / 13, 1 / 13, 1 / 13, 1 / 13, 1 / 13),
[0135] marker_gene (Object, num_clusters, "Monocyte", "CD14", "CD33", "CD4", "FCGR3A", "ITGAM", "ITGAX", "LAPTM5", "LYZ", "PTPRC", "VIM", 2 / 11, 1 / 11, 1 / 11, 1 / 11, 1 / 11, 1 / 11, 1 / 11, 1 / 11, 1 / 11, 1 / 11),
[0136] marker_gene (Object, num_clusters, "DC", "CD1C", "HLA-DPA1", "HLA-DPB1", "HLA-DQB1", "HLA-DRA", "ITGAX", "NRP1", "NA", "NA", "NA", 1 / 7, 1 / 7, 1 / 7, 1 / 7, 1 / 7, 1 / 7, 1 / 7, 0, 0, 0),
[0137] marker_gene (Object, num_clusters, "Mast_cells", "CAM1", "CPA3", "CTSG", "KIT", "TPSD1", "TPSAB1", "NA", "NA", "NA", "NA", 1 / 6, 1 / 6, 1 / 6, 1 / 6, 1 / 6, 0, 0, 0, 0, 0),
[0138] marker_gene (Object, num_clusters, "cancer_stem_cells", "CD44", "CD24", "ALDH1", "CD133", "ALDH", "SOX2", "EPCAM", "POU5F1", "BMI1", "ITGA6", 52 / 151, -29 / 151, 22 / 151, 13 / 151, 8 / 151, 8 / 151, 6 / 151, 6 / 151, 4 / 151, 3 / 151),
[0139] marker_gene (Object, num_clusters, "cancer_cells", "ERBB2", "EPCAM", "GATA3", "KRT18", "ADH1B", "AKR1A1", "ALPI", "ARHGAP10", "B3GNT2", "BCAM", 5 / 17, 2 / 17, 2 / 17, 2 / 17, 1 / 17, 1 / 17, 1 / 17, 1 / 17, 1 / 17, 1 / 17, 1 / 17)
[0140] Next, the cell type corresponding to the highest cell class group classification score of each cell population is marked as the cell type of the respective cell population, as follows:
[0141] Set the column name of the table obtained by scoring the cell class groups in the previous step as cell type name;
[0142] Set the cell type corresponding to the highest cell class group score of each cell population as the cell type of the cell population, specifically: celltype = data.frame(cluster = c(0:(num_clusters-1)), celltype = apply(celltype_score, 1, function(t) colnames(celltype_score)[which.max(t)])), obtain the cell population-cell type correspondence table;
[0143] In this embodiment, the cell types corresponding to the highest cell class group classification scores of cell populations 0-7 are endothelial cells, tumor-associated fibroblasts, fibroblasts, monocytes, B cells, tumor cells, fibroblasts, and NK cells, respectively.
[0144] Step 4: Obtain the distribution of cell class groups in the UMAP or TSNE dimension reduction space
[0145] In this embodiment, the cell class group annotation information is added to the single cell sequencing data frame, and the distribution of 7 cell class groups including tumor cells, endothelial cells, fibroblasts, etc. in the UMAP or TSNE dimension reduction space is obtained and output.
[0146] Step 5: Determine whether there are tumor cells in the cell class group annotation results
[0147] In this embodiment, there are tumor cells, so the tumor cells are extracted and Steps 6-10 are sequentially executed.
[0148] Step 6: Further cluster and cluster the tumor cells extracted above.
[0149] In this embodiment, the tumor cells are further sub-clustered and clustered to obtain two tumor cell populations and their distribution in the UMAP or TSNE dimension reduction space.
[0150] Step7: Scoring tumor cells using tumor cell EMT marker gene set or taxane drug resistance marker gene set
[0151] Preferably, the tumor cell EMT marker gene set and the taxane drug resistance marker gene set are obtained from literature and collected and screened from the Cancer Treatment Response Gene Marker Database (CTR-DB), respectively.
[0152] In this embodiment, the tumor cells are scored for EMT and taxane drug resistance using the AddModuleScore in the Seurat package and the EMT and taxane resistance gene sets collected above, and EMT characteristic tumor cells, non-EMT characteristic tumor cells, taxane drug resistance tumor cells, and non-taxane drug resistance tumor cells are obtained, respectively.
[0153] Step8: Isolation of EMT characteristic tumor cells and non-EMT characteristic tumor cells, isolation of taxane drug resistance tumor cells and non-taxane drug resistance tumor cells
[0154] Preferably, cells with EMT scores greater than 0 are extracted and set as EMT characteristic tumor cells, cells with EMT scores less than 0 are extracted and set as non-EMT characteristic tumor cells, cells with taxane resistance scores greater than 0 are extracted and set as taxane resistance characteristic tumor cells, and cells with taxane resistance scores less than 0 are extracted and set as non-taxane resistance characteristic tumor cells, and EMT score small violin plot and taxane drug resistance score small violin plot of tumor cells are obtained, respectively.
[0155] Step9: Obtain differential expression genes of EMT and non-EMT characteristic tumor cells and differential genes of taxane resistance and non-taxane resistance tumor cells
[0156] In this embodiment, the FindMarkers function in the Seurat package is used, preferably the statistical test method for comparison is "wilcox", genes expressed in at least 10% of cells are retained, genes with adjusted p-value less than 0.05 are used as significantly differentially expressed genes, and a differential expression gene table is output.
[0157] Step10: Pathway enrichment analysis of differentially expressed genes
[0158] In this embodiment, preferably, the 200 genes with the most significant differences are screened and KEGG and GO enrichment analyses are performed using clusterProfiler's enrichKEGG and enrichGO functions, respectively. The p-value threshold for screening significantly enriched pathways is set to 0.06, the q-value threshold for multiple test correction is 0.05, and the method for multiple test correction is Benjamini-Hochberg. The enriched pathway table and the visualized bubble chart are then obtained and output.
[0159] This invention also provides a system for identifying cell populations and EMT or taxane resistance subtypes of breast cancer cells based on single-cell transcriptome sequencing data, comprising:
[0160] Algorithm server and webpage for uploading, parameter selection, and display;
[0161] The algorithm server is used to store algorithms and calculate the identification of cell groups and EMT or taxane-resistant tumor cell subtypes in breast cancer single-cell transcriptome sequencing data, thereby obtaining the identification results of breast cancer single-cell transcriptome sequencing cell groups and EMT or taxane-resistant tumor cell subtypes.
[0162] The upload, parameter selection, and display webpages are integrated into the same website, which are used to upload single-cell transcriptome sequencing data, select the result observation interface, and display the calculation results, respectively.
[0163] Appendix Figure 2 The diagram shows the modules and specific processes for data uploading and system analysis in this invention, including:
[0164] S1: Users upload breast cancer single-cell transcriptome expression matrix data through the upload option on the visual webpage. These data files are cell gene expression files after the original data has been filtered and gene aligned by the 10X Genomics official analysis software Cell Ranger, including: barcodes.tsv.gz, features.tsv.gz, matrix.mtx.gz, and then execute S2;
[0165] S2: The algorithm server transfers the file to the backend, calls the R script, and then executes S3;
[0166] S3: The computation module within the R script identifies cell populations from the uploaded expression matrix, and then executes S4;
[0167] S4: The calculation module in the R script determines whether there are tumor cells in the cell population. If so, execute S5; otherwise, execute S6.
[0168] S5: The calculation module in the R script calculates the EMT or taxane drug resistance subtype tumor cell characteristics of the tumor cells, and then performs S6;
[0169] S6: The algorithm server stores the results locally and returns to the front end, and then performs S7;
[0170] S7: The user selects the result observation interface parameters, which can include cell group identification pictures and tables, tumor cell EMT or taxane drug resistance score pictures, EMT characteristic tumor cells or taxane drug resistance characteristic tumor cells differential expression gene tables, EMT characteristic tumor cells or taxane drug resistance characteristic tumor cells pathway enrichment pictures and tables, and then performs S8;
[0171] S8: Display the result file and provide user download.
[0172] Figure 1 shows the display webpage, including: breast cancer single cell transcriptome sequencing data selection and upload area and result display area Figure 3 Figure 3
[0173] Figure 2 shows the breast cancer single cell transcriptome sequencing data successful upload interface Figure 3
[0174] Figure 3 Figure 3 shows the calculation and result loading process interface
[0175] Figure 4 shows the uploaded breast cancer single cell sequencing data cell type classification result picture Figure 3
[0176] Figure 5 shows the uploaded breast cancer single cell EMT characteristic tumor cell enrichment pathway picture Figure 3
[0177] Figure 6 shows the uploaded breast cancer single cell taxane drug resistance characteristic tumor cell enrichment pathway picture Figure 3
[0178] Figure 7 shows the uploaded breast cancer single cell analysis result table and download interface Figure 3
[0179] The application provides a breast cancer single-cell transcriptome sequencing data cell cluster identification and EMT or taxane drug resistance characteristic subtype tumor cell identification method and system, and the main innovations are as follows: first, the application proposes a new breast cancer single-cell sequencing data cell cluster identification method, integrates the marker genes of basic cell types of breast cancer tissues, and first gives a calculation weight based on the number of cell type specific genes in the literature and constructs a cell cluster identification function. Second, the application can automatically obtain EMT and taxane drug resistance subtype tumor cells and their gene expression characteristics, which helps to understand the pathogenesis of breast cancer patients and optimize the diagnosis and treatment process of patients. Third, the application provides an algorithm server and a visualization webpage, which can automatically analyze and present the results of breast cancer single-cell transcriptome sequencing.
[0180] Those skilled in the art will readily understand that the above description is only the preferred embodiment of the present application, and is not intended to limit the present application, and any modifications, equivalent replacements and improvements made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for identifying the characteristics of metastatic or drug-resistant subtypes of breast cancer cells, characterized in that, Includes the following steps: (1) Calculate the relative expression levels of epithelial-mesenchymal transition marker genes or taxane drug marker genes in breast cancer tumor cell populations; (2) Cells with a relative expression level of epithelial-mesenchymal transition marker gene greater than 0 were extracted and set as tumor cells with epithelial-mesenchymal transition marker gene characteristics, and cells with a relative expression level of epithelial-mesenchymal transition marker gene less than 0 were extracted and set as tumor cells without epithelial-mesenchymal transition marker gene characteristics. Cells with taxane resistance expression levels greater than 0 were extracted and designated as taxane-resistant tumor cells, while cells with taxane resistance expression levels less than 0 were extracted and designated as non-taxane-resistant tumor cells. (3) Separate tumor cells with epithelial-mesenchymal transition (EMT) and non-EMT characteristics, and obtain violin plots of EMT scores for both. Separate tumor cells with taxane resistance and non-taxane resistance characteristics, and obtain violin plots of taxane resistance scores for both. (4) Identify differentially expressed genes in tumor cells with epithelial-mesenchymal transition and non-epithelial-mesenchymal transition characteristics, and identify differentially expressed genes in tumor cells with taxane resistance and paclitaxel non-resistance characteristics; (5) Screening steps (4) select several differentially expressed genes that show significant differences and perform KEGG or GO enrichment analysis to obtain enriched pathways and visualized bubble charts.
2. A system for identifying the characteristics of metastatic or drug-resistant subtype tumor cells in breast cancer, characterized in that, include: Calculation module: The calculation module is used to calculate the relative expression levels of epithelial-mesenchymal transition marker genes or taxane drug marker genes in breast cancer tumor cell populations; Extraction and setting module: The extraction and setting module is used to extract cells with a relative expression level of epithelial-mesenchymal transition marker gene greater than 0 and set them as tumor cells characterized by epithelial-mesenchymal transition marker gene, and to extract cells with a relative expression level of epithelial-mesenchymal transition marker gene less than 0 and set them as tumor cells not characterized by epithelial-mesenchymal transition marker gene. Cells with taxane resistance expression levels greater than 0 were extracted and designated as taxane-resistant tumor cells, while cells with taxane resistance expression levels less than 0 were extracted and designated as non-taxane-resistant tumor cells. The scoring violin plot module is used to separate tumor cells with epithelial-mesenchymal transition (EMT) and non-EMT characteristics, and to obtain the scoring violin plots of both. It is also used to separate taxane-resistant and non-taxane-resistant tumor cells, and to obtain the taxane-resistant violin plots of both. Identification module: The identification module is used to identify differentially expressed genes in tumor cells with epithelial-mesenchymal transition (EMT) and non-EMT characteristics, and to identify differentially expressed genes in tumor cells with taxane resistance and paclitaxel non-resistance characteristics; Enrichment Module: The enrichment module is used to perform KEGG or GO enrichment analysis on several genes with significant differences in differential expression, and to obtain enriched pathways and visualized bubble charts.