Screening method of triple negative breast cancer immunotherapy drug-resistant population
By acquiring and analyzing gene expression data from triple-negative breast cancer patients using single-cell technology, and screening for drug resistance-associated subgroups and characteristic genes, this technology addresses the lack of precision in screening for drug-resistant populations in triple-negative breast cancer immunotherapy in existing technologies. It enables precise localization of drug-resistant populations and supports the selection of treatment regimens.
Patent Information
- Application Number
- CN202511826985.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-02-27
AI Technical Summary
Current technologies cannot effectively distinguish biological signals from different cell types, resulting in insufficient precision in screening patients resistant to immunotherapy in triple-negative breast cancer. It is also difficult to analyze cell type-specific biological signals, affecting the accuracy of immunotherapy response studies.
Gene expression data of triple-negative breast cancer patients were obtained using single-cell technology. Immune and matrix-related indicators were calculated to generate an immune matrix score comparison dataset. Single-cell samples were collected for clustering and secondary clustering to screen drug resistance-related subgroups. Specific expression genes were extracted to establish drug resistance characteristic gene combinations. The expression status of characteristic genes was detected by combining patient transcriptome data, and thresholds were set to determine drug-resistant populations.
This technology improves the accuracy of screening for drug-resistant patients in triple-negative breast cancer, enabling the differentiation of different cell types and the analysis of unique biological signals, thus providing a reliable basis for subsequent treatment plans.
Smart Images

Figure CN121583322A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical detection technology, and in particular to a method for screening individuals resistant to immunotherapy in triple-negative breast cancer. Background Technology
[0002] Breast cancer, a highly heterogeneous tumor, ranks first in incidence and second in mortality among women worldwide. Different subtypes exhibit significant differences in treatment options and prognosis. Triple-negative breast cancer is a special subtype characterized by resistance to immunotherapy for all three receptors (estrogen receptor, progesterone receptor, and human epidermal growth factor receptor). Screening for patients with triple-negative breast cancer who are resistant to immunotherapy yields negative results. This type of patient exhibits highly aggressive and rapidly progressing tumors, and due to the lack of effective targets, traditional treatments are often less effective, resulting in a poor prognosis. Given the high heterogeneity of triple-negative breast cancer patients, current research in the field of immunotherapy often employs methods such as screening for patients resistant to immunotherapy in triple-negative breast cancer (Bulk screening, RNA screening, seq screening) to compare and analyze patients with different treatment responses in order to identify relevant biomarkers.
[0003] However, screening for patients resistant to immunotherapy in triple-negative breast cancer using methods such as Bulk, RNA, and seq has significant limitations. It cannot distinguish between different cell types and struggles to analyze the unique biological signals of these cell types, which to some extent restricts the precise study of immunotherapy response. Based on this, this patent proposes a method for predicting the immunotherapy response of triple-negative breast cancer patients by using single-cell technology to identify biomarkers and construct predictive models. Through in-depth analysis of single-cell data from immunotherapy-resistant and beneficiary populations, model scores are obtained in tumor cells from both populations to determine whether the patient is suitable for immunotherapy. Summary of the Invention
[0004] To address the technical problems existing in the prior art, this invention provides a method for screening patients resistant to immunotherapy in triple-negative breast cancer. The technical solution is as follows: A screening method for patients resistant to immunotherapy in triple-negative breast cancer includes the following steps: S1: Obtain gene expression data of triple-negative breast cancer patients, calculate immune and matrix-related indicators, obtain corresponding quantitative values, and generate immune matrix score comparison dataset; S2: Collect mRNA sequence data of single-cell samples from triple-negative breast cancer patients, remove low-quality data and normalize, filter variable genes by dimensionality reduction and then cluster, annotate fibroblasts and macrophages, perform secondary clustering to quantify the proportion of subpopulations, screen drug resistance-associated subpopulations, and obtain a list of drug resistance-associated cell subpopulations. S3: Extract gene expression data of MCAM+ fibroblast subsets, compare gene expression differences with other subsets, screen for specifically expressed genes, and establish a combination of drug resistance characteristic genes; S4: Obtain bulk transcriptome data of patients to be screened, detect the expression status of drug resistance characteristic genes, count the number of highly expressed genes, and obtain the matching degree of drug resistance characteristics of patients; S5: Set a drug resistance feature matching threshold, compare the patient's drug resistance feature matching degree with the threshold, determine whether the patient is drug-resistant or non-drug-resistant, and generate screening results for immunotherapy drug-resistant patients.
[0005] As a further aspect of the present invention, the immune matrix scoring comparison dataset includes patient individual identifiers, immune-related quantitative values, and matrix-related quantitative values; the drug resistance-associated cell subpopulation list includes cell subpopulation names, cell subpopulation percentages, and the degree of association with drug resistance phenotypes; the drug resistance characteristic gene combination includes gene names, gene-specific expression markers, and gene function annotations; the patient drug resistance characteristic matching degree includes the number of highly expressed characteristic genes, the total number of characteristic genes, and the matching ratio; and the immunotherapy drug resistance population screening results include patient individual identifiers, drug resistance determination conclusions, and matching degree values.
[0006] As a further aspect of the present invention, the steps for obtaining the immune matrix scoring comparison dataset are as follows: S101: Obtain gene expression data of triple-negative breast cancer patients, select immune-related indicators and matrix-related indicators, perform quantitative calculations on the two types of indicators respectively, obtain the quantitative values of the corresponding indicators for each patient, and generate a basic immune matrix data table. S102: Call the basic immune matrix data table, integrate the patient's individual identifier and corresponding quantitative value, classify and organize them according to the indicator type, form a standardized data structure, and generate an immune matrix score comparison dataset.
[0007] As a further aspect of the present invention, the step of obtaining the list of drug resistance-associated cell subpopulations is as follows: S201: Collect single-cell samples from triple-negative breast cancer patients, extract single-cell mRNA sequence data, remove low-quality data, normalize the remaining data, reduce data dimensionality through principal component analysis, screen variable genes, and obtain a standardized single-cell gene dataset. S202: Based on a standardized single-cell gene dataset, cluster cells according to their expression characteristics, call up the clustering results of cell marker genes, annotate fibroblasts and macrophages, and obtain the target cell classification dataset; S203: For the target cell classification dataset, perform secondary clustering of the two cell types, quantify the proportion of each subgroup, screen corresponding subgroups based on the immunotherapy resistance phenotype, and generate a list of drug resistance-related cell subgroups.
[0008] As a further aspect of the present invention, the step of obtaining the drug resistance characteristic gene combination is as follows: S301: Call the drug resistance-associated cell subpopulation list, extract gene expression data of MCAM+ fibroblast subpopulation, collect gene expression data of other cell subpopulations, and obtain a multi-subpopulation gene comparison dataset. S302: Based on a multi-subpopulation gene comparison dataset, compare the gene expression differences between the MCAM+ fibroblast subpopulation and other subpopulations, screen for specifically expressed genes, and establish a combination of drug resistance characteristic genes.
[0009] As a further aspect of the present invention, the step of obtaining the patient drug resistance characteristic matching degree is as follows: S401: Obtain bulk transcriptome gene expression data of patients to be screened, call the combination of drug resistance characteristic genes, detect the expression status of each characteristic gene one by one, record the actively expressed genes, and obtain the characteristic gene expression status table. S402: Calculate the number of actively expressed genes in the statistical characteristic gene expression status table, and then calculate the ratio of this number to the total number of characteristic genes to obtain the patient's drug resistance characteristic matching degree.
[0010] As a further aspect of the present invention, the step of obtaining the screening results of immunotherapy-resistant populations is as follows: S501: Based on clinical immunotherapy response data and the expression patterns of drug resistance characteristic genes, a criterion for distinguishing between drug resistance and non-drug resistance was set to obtain the drug resistance characteristic matching threshold. S502: Call the patient's drug resistance feature matching degree and drug resistance feature matching threshold, compare the values, classify those above the threshold as drug-resistant, and classify those below the threshold as non-drug-resistant, and generate the screening results of immunotherapy drug-resistant population.
[0011] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: This invention obtains gene expression data from triple-negative breast cancer patients and then uses a specific algorithm to calculate immune and matrix-related indicators to generate corresponding datasets. Based on single-cell samples, mRNA sequence data is collected, processed, clustered, and annotated for specific cell types. A second clustering is then performed to screen for drug-resistant subpopulations. Gene data is extracted from specific subpopulations to screen for highly expressed genes and establish feature combinations. Transcriptome data from patients to be screened is obtained to detect the expression status of characteristic genes, and the number of highly expressed genes is statistically analyzed to obtain the matching degree. The matching degree is compared with a threshold to generate screening results. This invention can distinguish different cell types, analyze unique biological signals, accurately locate drug-resistant cell subpopulations and genes, improve the accuracy of screening for immunotherapy-resistant populations, and provide a more reliable basis for subsequent treatment selection. Attached Figure Description
[0012] Figure 1 This is an overall flowchart of the screening method of the present invention; Figure 2 This is a flowchart illustrating the acquisition process of the immune matrix scoring comparison dataset of the present invention. Figure 3 This is a flowchart illustrating the process of obtaining the drug resistance-associated cell subpopulation list of the present invention. Figure 4 This is a flowchart illustrating the process of obtaining the drug resistance gene combination of the present invention. Figure 5 This is a flowchart illustrating the process of obtaining patient drug resistance characteristic matching in this invention. Figure 6 This is a flowchart illustrating the process of obtaining the screening results for immunotherapy-resistant populations according to the present invention. Detailed Implementation
[0013] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0014] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0015] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0016] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0017] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0018] Please see Figure 1 This invention provides a technical solution: a method for screening patients resistant to immunotherapy in triple-negative breast cancer, comprising the following steps: S1: Obtain gene expression data of triple-negative breast cancer patients, call the xCell algorithm and ESTIMATE algorithm to calculate the immune-related indicators and matrix-related indicators in the gene expression data respectively, obtain the immune-related quantitative values and matrix-related quantitative values for each patient, and generate an immune matrix score comparison dataset. S2: Based on single-cell samples from triple-negative breast cancer patients, single-cell mRNA sequence data were collected, low-quality sequence data were removed, the remaining data were normalized, principal component analysis was used to reduce data dimensionality, variable genes were screened, and clustering was performed according to cell expression characteristics. Cell marker genes were called to compare the clustering results, fibroblasts and macrophages were annotated, and the two types of cells were clustered again to quantify the proportion of each subpopulation. Subpopulations whose proportions were associated with immunotherapy resistance phenotypes were screened to obtain a list of drug resistance-associated cell subpopulations. S3: For the MCAM+ fibroblast subpopulation in the drug resistance-associated cell subpopulation list, extract gene expression data of this subpopulation, compare the gene expression differences of this subpopulation with other cell subpopulations, screen for specific high-expression genes, and establish a combination of drug resistance characteristic genes. S4: Obtain bulk transcriptome gene expression data of patients to be screened, call the combination of drug resistance characteristic genes, detect the expression status of each characteristic gene, count the number of highly expressed genes among the characteristic genes, and obtain the matching degree of patient drug resistance characteristics; S5: Set a drug resistance feature matching threshold, compare the patient's drug resistance feature matching degree with the threshold, and determine the patient as a drug-resistant population if the value is higher than the threshold, and as a non-drug-resistant population if the value is lower than the threshold, thus generating the screening results for immunotherapy drug-resistant populations.
[0019] The immune matrix score comparison dataset includes patient individual identifiers, immune-related quantitative values, and matrix-related quantitative values. The drug resistance-associated cell subpopulation list includes cell subpopulation names, cell subpopulation percentages, and the degree of association with drug resistance phenotypes. The drug resistance characteristic gene combination includes gene names, gene-specific expression markers, and gene function annotations. The patient drug resistance characteristic matching degree includes the number of highly expressed characteristic genes, the total number of characteristic genes, and the matching ratio. The immunotherapy drug resistance population screening results include patient individual identifiers, drug resistance determination conclusions, and matching degree values.
[0020] Please see Figure 2 The steps for obtaining the immune matrix scoring comparison dataset are as follows: S101: Obtain gene expression data of triple-negative breast cancer patients, select immune-related indicators and matrix-related indicators, perform quantitative calculations on the two types of indicators respectively, obtain the quantitative values of the corresponding indicators for each patient, and generate a basic immune matrix data table. Raw gene expression data from triple-negative breast cancer patients were selected. The data included transcript counts of immune-related genes (such as CD3, CD4, CD8, PD-1, PD-L1) and matrix-related genes (such as COL1, COL3, POSTN, FN1) in the tumor tissue of each patient. First, it was defined that immune-related indicators are the set of expression levels of the aforementioned immune genes, and matrix-related indicators are the set of expression levels of the aforementioned matrix genes. For the quantification of immune-related indicators, the raw transcript counts of each immune gene for each patient were extracted. A threshold of 1.5 for the expression level of a single immune gene was set. When the raw transcript count of the CD3 gene in a patient was 320, COL3, POSTN, FN1, etc., the threshold was set. When D4 is 280, CD8 is 250, PD-1 is 120, and PD-L1 is 90, genes with transcript counts below 1.5 are removed (there are no such genes in this example). Then, the ratio of each immune gene transcript count to the patient's average transcript count is calculated. The patient's average transcript count is 200, so the ratios are: CD3 320 / 200 = 1.6, CD4 280 / 200 = 1.4, CD8 250 / 200 = 1.25, PD-1 120 / 200 = 0.6, and PD-L1 90 / 200 = 0.45. These ratios are then standardized using the following formula: ,in Let be the standardized value of the i-th immune gene. Let be the ratio of the i-th immune gene. This is the average of the ratios of all immune genes (in this example, it is (1.6+1.4+1.25+0.6+0.45) / 5=1.06). The standard deviation of all immune gene ratios (calculated as 0.43 in this example) is used. Substituting the data, the standardized values are: CD3 (1.6-1.06) / 0.43≈1.26, CD4 (1.4-1.06) / 0.43≈0.79, CD8 (1.25-1.06) / 0.43≈0.44, PD-1 (0.6-1.06) / 0.43≈-1.07, and PD-L1 (0.45-1.06) / 0.43≈-1.42. Finally, the standardized values of each immune gene are summed to obtain the quantitative value of the patient's immune-related indicators, i.e., 1.26+0.79+0.44-1.07-1.42≈-0.00. The same procedure is used for the quantitative calculation of matrix-related indicators, yielding the patient's original transcript counts as 450 for COL1, 380 for COL3, and 320 for POSTN. With FN1 at 290 and the average transcript count for all genes remaining at 200, the calculated ratios of each matrix gene are 450 / 200 = 2.25, 380 / 200 = 1.9, 320 / 200 = 1.6, and 290 / 200 = 1.45. The average ratio is calculated to be (2.25 + 1.9 + 1.6 + 1.45) / 4 = 1.8, with a standard deviation of 0.34. After standardization, the values are (2.25 - 1.8). ) / 0.34≈1.32, (1.9-1.8) / 0.34≈0.29, (1.6-1.8) / 0.34≈-0.59, (1.45-1.8) / 0.34≈-1.03, summing them up gives the quantitative value of matrix-related indicators as 1.32+0.29-0.59-1.03≈-0.01. Follow the above process to complete the indicator quantification for all patients and generate the basic data table of immune matrix.
[0021] S102: Call the basic immune matrix data table, integrate the patient's individual identifier and corresponding quantitative value, classify and organize them according to the indicator type, form a standardized data structure, and generate an immune matrix score comparison dataset.
[0022] The generated basic immune matrix data table is accessed. This table contains unique patient identifiers (e.g., P001, P002, P003, etc.), quantitative values of immune-related indicators, and quantitative values of matrix-related indicators. First, the individual identifier for each patient is extracted along with its corresponding two quantitative values. For example, patient P001 corresponds to an immune quantitative value of -0.00 and a matrix quantitative value of -0.01; patient P002 has an immune quantitative value of 1.23 and a matrix quantitative value of 0.89; patient P003 has an immune quantitative value of -0.76 and a matrix quantitative value of 1.54, etc. Then, the data is categorized and organized by indicator type, grouping all patients' individual identifiers and immune-related indicator quantitative values into... One category consists of individual identifiers and quantitative values of matrix-related indicators, while the other category consists of individual identifiers and matrix-related indicators. The standardized data structure is that each row corresponds to one patient and each column corresponds to one data point. The first column is fixed as the patient's individual identifier, the second column corresponds to the quantitative value of immune-related indicators, and the third column corresponds to the quantitative value of matrix-related indicators. The data is checked for completeness, and records without individual identifiers or missing quantitative values are removed. It is confirmed that all patient data have a uniform format and reasonable numerical range (both immune and matrix quantitative values are within the range of -5 to 5). Finally, a standardized table containing patient individual identifiers, quantitative values of immune-related indicators, and quantitative values of matrix-related indicators is generated, thus creating an immune matrix score comparison dataset.
[0023] Please see Figure 3 The steps for obtaining the list of drug resistance-associated cell subpopulations are as follows: S201: Collect single-cell samples from triple-negative breast cancer patients, extract single-cell mRNA sequence data, remove low-quality data, normalize the remaining data, reduce data dimensionality through principal component analysis, screen variable genes, and obtain a standardized single-cell gene dataset. Tumor tissue samples were collected from surgically removed patients with triple-negative breast cancer. Single-cell suspensions were obtained through mechanical grinding and enzymatic digestion. mRNA sequencing was performed using the 10xGenomics platform to obtain raw sequencing data (stored in FASTQ format). The data included gene expression read counts for each cell. Low-quality data rejection criteria were established: cells with fewer than 200 or more than 6000 genes were considered abnormal cells, and cells with more than 10% mitochondrial genes were considered damaged cells. Valid cells were selected based on these criteria. The remaining data were normalized using the "parts per million transformation" method, which divided the gene expression level of each cell by the total expression level of that cell, multiplied by 1e6, and then logarithmically transformed (log2(x+1)) to make the data conform to a normal distribution. Principal component analysis was used to reduce data dimensionality. The covariance matrix of the gene expression matrix was calculated, and eigenvalues and eigenvectors were solved. The top 20 principal components (with a cumulative variance contribution rate exceeding 80%) were selected as principal features. Variable genes were screened by calculating the coefficient of variation (the ratio of standard deviation to mean) for each gene across all cells. A threshold of 1.5 was set for the coefficient of variation, and genes with coefficients of variation greater than this threshold were selected as variable genes. Taking the original dataset containing 1000 cells and 20000 genes as an example, after removing cells with abnormal gene counts and excessive mitochondrial gene proportions, 800 cells remained. After normalization, the gene expression levels of each cell were converted into relative values. Principal component analysis yielded 20 principal components, from which 1500 variable genes were selected, resulting in a standardized single-cell gene dataset.
[0024] S202: Based on a standardized single-cell gene dataset, cluster cells according to their expression characteristics, call up the clustering results of cell marker genes, annotate fibroblasts and macrophages, and obtain the target cell classification dataset; Based on a standardized single-cell gene dataset, the K-means clustering algorithm was used to cluster cells according to their expression characteristics. The number of clusters was set to 10. The Euclidean distance between each cell and the center of each cluster was calculated, and cells were assigned to the nearest cluster. A known list of cellular marker genes was used: fibroblast marker genes included COL1A1, COL3A1, and POSTN; macrophage marker genes included CD68, CD163, and CSF1R. The average expression level of marker genes in each cluster was calculated. If the average expression levels of COL1A1, COL3A1, and POSTN in a cluster were all higher than those in other clusters (with a fold change threshold of 2), then that cluster was labeled as fibroblasts; if the average expression levels of CD68, CD163, and CSF1R were all higher than those in other clusters (with a fold change threshold of 2), then that cluster was labeled as macrophages. Taking a dataset containing 10 clusters as an example, the average expression levels of COL1A1, COL3A1, and POSTN in cluster 2 are 5.2, 4.8, and 4.5, respectively, which are more than twice that of other clusters, and are labeled as fibroblasts; the average expression levels of CD68, CD163, and CSF1R in cluster 5 are 6.1, 5.7, and 5.3, respectively, which are more than twice that of other clusters, and are labeled as macrophages, thus obtaining the target cell classification dataset.
[0025] S203: For the target cell classification dataset, perform secondary clustering of the two cell types, quantify the proportion of each subgroup, screen corresponding subgroups based on the immunotherapy resistance phenotype, and generate a list of drug resistance-related cell subgroups.
[0026] For the target cell classification dataset, secondary clustering was performed on fibroblasts and macrophages separately using hierarchical clustering. The Pearson correlation coefficient between cells was calculated as a distance metric, and cluster merging was performed using the Ward method. The number of clusters was set to 5 (based on the maximum silhouette coefficient). The proportion of each subpopulation was quantified by counting the number of cells in each subpopulation and dividing it by the total number of the corresponding major cell type (fibroblasts or macrophages) to obtain the subpopulation proportion. To associate immunotherapy resistance phenotypes, immunotherapy response data from patients (divided into resistant and sensitive groups) was collected. The difference in the proportion of each subpopulation in the resistant and sensitive groups was calculated, and statistical analysis was performed using the chi-square test. A p-value threshold of 0.05 was set, and subpopulations with p-values less than this threshold were selected as resistance-associated subpopulations. Taking fibroblasts as an example, secondary clustering yielded five subpopulations: A, B, C, D, and E, with cell numbers of 80, 120, 100, 60, and 40, respectively, accounting for 20%, 30%, 25%, 15%, and 10% of the total fibroblast population (400). After associating with drug resistance phenotypes, subpopulation B accounted for 45% of the drug-resistant group and 15% of the sensitive group, with a P-value of 0.02 (less than 0.05), and was thus selected as a drug resistance-associated subpopulation. Similarly, after secondary clustering of macrophages, subpopulation D accounted for 35% of the drug-resistant group and 10% of the sensitive group, with a P-value of 0.03 (less than 0.05), and was also selected as a drug resistance-associated subpopulation, generating a list of drug resistance-associated cell subpopulations.
[0027] Please see Figure 4 The steps for obtaining drug resistance gene combinations are as follows: S301: Call the drug resistance-associated cell subpopulation list, extract gene expression data of MCAM+ fibroblast subpopulation, collect gene expression data of other cell subpopulations, and obtain a multi-subpopulation gene comparison dataset. The generated list of drug resistance-associated cell subpopulations is retrieved. This list clearly identifies the cell IDs and classifications of the MCAM+ fibroblast subpopulation and other cell subpopulations (including POSTN+ fibroblasts, COL15A1+ fibroblasts, WSB1+ fibroblasts, macrophages, etc.). First, all cell IDs (e.g., C001-C120) of the MCAM+ fibroblast subpopulation are extracted. Based on these IDs, all gene expression data for the corresponding cells are retrieved from the standardized single-cell gene dataset. Gene expression data for each cell are presented as quantified standardized values (ranging from -3 to 3). For example, the standardized expression values for MCAM in cell C001 are 2.3, RGS5 is 1.9, and NDUFA4L2 is 1.7; for cell C002, MCAM is 2.1, RGS5 is 1.8, and NDUFA4L2 is 1.6, etc. Then, according to other cell subpopulations in the catalog, the POSTN+ fibroblast subpopulation (cell numbers C121-C200) and COL15A1+ fibroblast subpopulation are extracted sequentially. Cell numbers were assigned to cell subpopulations (C201-C280), WSB1+ fibroblast subpopulation (C281-C350), and macrophage subpopulation (C351-C450). Complete gene expression data for all cells in each subpopulation were retrieved. Specifically, in cell C121 of the POSTN+ fibroblast subpopulation, the normalized expression values of POSTN gene were 2.0 and COL1 was 1.8; in cell C201 of the COL15A1+ fibroblast subpopulation, the normalized expression values of COL15A1 and FN1 were 1.7. In the B1+ fibroblast subpopulation, cell C281 showed WSB1 at 1.8 and CD44 at 1.6, while in the macrophage subpopulation, cell C351 showed CD68 at 2.2 and CD163 at 2.0. Gene expression data from the MCAM+ fibroblast subpopulation were categorized and organized with gene expression data from other cell subpopulations. Each subpopulation data set included the gene names and corresponding standardized expression values for all cells within that subpopulation, forming a structured dataset containing multi-subpopulation gene expression information, resulting in a multi-subpopulation gene comparison dataset.
[0028] S302: Based on a multi-subpopulation gene comparison dataset, compare the gene expression differences between the MCAM+ fibroblast subpopulation and other subpopulations, screen for specifically expressed genes, and establish a combination of drug resistance characteristic genes.
[0029] Based on a multi-subpopulation gene comparison dataset, the mean normalized expression value of each gene in the MCAM+ fibroblast subpopulation was first calculated. 1000 highly expressed genes (mean normalized expression value ≥ 0.5) in this subpopulation were selected as candidate genes. Then, the mean normalized expression value of each candidate gene in other cell subpopulations was calculated. The gene expression difference fold (mean expression value of MCAM+ subpopulation ÷ mean expression value of other subpopulations) was obtained by calculating the difference in mean expression value of the same gene between the MCAM+ fibroblast subpopulation and other subpopulations. The difference screening criteria were set as follows: fold difference ≥ 2.0 and mean normalized expression value of the gene in the MCAM+ subpopulation ≥ 1.0. Simultaneously, the statistical difference in gene expression between the MCAM+ subpopulation and other subpopulations was calculated, with a statistical difference threshold of 0.05 (i.e., the difference in gene expression between the two subpopulations is statistically significant). Taking the candidate gene MCAM as an example, its average normalized expression value was 2.2 in the MCAM+ fibroblast subset, 0.8 in the POSTN+ fibroblast subset, 0.7 in the COL15A1+ subset, 0.6 in the WSB1+ subset, and 0.5 in the macrophage subset, with fold differences of 2.75, 3.14, 3.67, and 4.4, respectively, all ≥2.0, and the statistical difference value was 0.02 < 0.05; the gene RGS5 in The mean expression value of the MCAM+ subpopulation was 1.9, while the mean expression values of the other subpopulations were 0.7, 0.6, 0.5, and 0.4, respectively, with fold differences ranging from 2.71 to 4.75 (statistical value < 0.01). The mean expression value of the gene NDUFA4L2 was 1.8 in the MCAM+ subpopulation, while the mean expression values of the other subpopulations were 0.6, 0.5, 0.4, and 0.3, with fold differences ranging from 3.0 to 6.0 (statistical value < 0.03). The average expression level of gene TINAGL1 in the MCAM+ subpopulation was 1.7, while the average expression levels in other subpopulations were 0.5, 0.4, 0.3, and 0.2, with a fold change of 3.4-8.5 and a statistical significance level of 0.04 < 0.05. The average expression level of gene ADIRF in the MCAM+ subpopulation was 1.6, while the average expression levels in other subpopulations were 0.4, 0.3, 0.2, and 0.1, with a fold change of 4.0-16.0 and a statistical significance level of 0.03 < 0.05. The average expression level of gene ACTA2 in the MCAM+ subpopulation was 1.5, while the average expression levels in other subpopulations were 0.3, 0.2, 0.1, and 0.05, with a fold change of 5.0-30.0 and a statistical significance level of 0.02 < 0.05. Six genes (MCAM, RGS5, NDUFA4L2, TINAGL1, ADIRF, and ACTA2) that met all the criteria were screened out and integrated to establish a drug resistance characteristic gene combination.
[0030] Please see Figure 5 The steps for obtaining patient drug resistance characteristic matching degree are as follows: S401: Obtain bulk transcriptome gene expression data of patients to be screened, call the combination of drug resistance characteristic genes, detect the expression status of each characteristic gene one by one, record the actively expressed genes, and obtain the characteristic gene expression status table. Bulk transcriptome gene expression data of tumor tissues from patients with triple-negative breast cancer were obtained. The data was indexed by gene name and included the raw transcript count for each gene (ranging from 0 to 5000). Simultaneously, a pre-established combination of drug resistance characteristic genes (including MCAM, RGS5, NDUFA4L2, TINAGL1, ADIRF, and ACTA2, a total of 6 genes) was retrieved. The bulk transcriptome gene expression data was preprocessed, and the average raw transcript count for all genes was calculated. The average raw transcript count for all genes in this patient was 320. Then, the raw transcript count for each characteristic gene was... Transcript counts were converted into relative expression levels (original transcript count of the characteristic gene ÷ average original transcript count of all genes). A threshold for determining active expression of the characteristic gene was set, which was determined based on bulk transcriptome data from 100 known immunotherapy-responsive patients. The relative expression level of each characteristic gene in each patient was calculated, and the minimum relative expression level of each characteristic gene in drug-resistant patients was taken as the active threshold of the corresponding gene. The thresholds were 1.2 for MCAM, 1.1 for RGS5, 1.0 for NDUFA4L2, 0.9 for TINAGL1, 0.8 for ADIRF, and 0.7 for ACTA2. Taking a patient to be screened as an example, the bulk transcriptome data showed the following raw transcript counts: MCAM 416, RGS5 384, NDUFA4L2 352, TINAGL1 288, ADIRF 224, and ACTA2 256. The relative expression levels of each gene were calculated as follows: MCAM 416 ÷ 320 = 1.3, RGS5 384 ÷ 320 = 1.2, NDUFA4L2 352 ÷ 320 = 1.1, TINAGL1 288 ÷ 320 = 0.9, ADIRF 224 ÷ 320 = 0.7, and ACTA2 256. ÷320=0.8. Compare the relative expression levels of each gene with the corresponding thresholds. MCAM1.3≥1.2, RGS51.2≥1.1, NDUFA4L21.1≥1.0, TINAGL10.9≥0.9, ADIRF0.7<0.8, ACTA20.8≥0.7. MCAM, RGS5, NDUFA4L2, TINAGL1, and ACTA2 are identified as actively expressed genes, while ADIRF is identified as an inactive gene. Organize the expression status of the genes by gene name, original transcript count, relative expression level, and expression activity status to obtain a characteristic gene expression status table.
[0031] S402: Calculate the number of actively expressed genes in the statistical characteristic gene expression status table, and then calculate the ratio of this number to the total number of characteristic genes to obtain the patient's drug resistance characteristic matching degree.
[0032] Retrieve the characteristic gene expression status table. First, determine that the total number of drug resistance characteristic gene combinations is 6. Then, count the number of gene entries marked "active" in the "Active Expression Status" column of the table row by row. Taking the characteristic gene expression status table of a patient to be screened as an example, MCAM, RGS5, NDUFA4L2, TINAGL1, and ACTA2 are active genes, totaling 5, while ADIRF is an inactive gene, totaling 1. Therefore, the number of actively expressed genes is determined to be 5. Calculate the patient's drug resistance characteristic matching degree using the formula "Drug resistance characteristic matching degree = Number of actively expressed genes ÷ Total number of characteristic genes". Substitute the value into the formula, i.e., 5 ÷ 6 ≈ 0.833, and retain three decimal places to obtain the drug resistance characteristic matching degree for this patient as 0.833. Repeat the above process to count the number of actively expressed genes and calculate the ratio for all patients to be screened, retaining three decimal places each time, to obtain the drug resistance characteristic matching degree for each patient.
[0033] Please see Figure 6 The steps for obtaining the screening results of immunotherapy-resistant individuals are as follows: S501: Based on clinical immunotherapy response data and the expression patterns of drug resistance characteristic genes, a criterion for distinguishing between drug resistance and non-drug resistance was set to obtain the drug resistance characteristic matching threshold. Clinical immunotherapy response data were collected from 150 patients with triple-negative breast cancer. Among them, 80 were non-drug-resistant patients with complete or partial remission after treatment, and 70 were drug-resistant patients with stable or progressive tumors. The drug resistance characteristic matching data (ranged from 0.1 to 0.95) of these 150 patients were also retrieved. The mean and standard deviation of the drug resistance characteristic matching data were calculated separately for the non-drug-resistant and drug-resistant groups. For the 80 patients in the non-drug-resistant group, the matching data were 0.21, 0.25, 0.30…0.58, with a mean of (0.21+0.25+…+0.58)÷80≈0.38 and a standard deviation of 0.08. For the 70 patients in the drug-resistant group, the matching data were 0.62, 0.65, 0.70…0.95, with a mean of (0.62+0.65+…+0.95)÷70≈0.76 and a standard deviation of 0.11. Further analysis of the overlapping areas of the two data sets revealed that the maximum matching score in the non-drug-resistant group was 0.58, while the minimum matching score in the drug-resistant group was 0.62. Since there was no overlap between the two groups, the median value of the two boundary values was taken as the initial candidate threshold value: (0.58 + 0.62) ÷ 2 = 0.60. To verify the discriminative effect of this candidate threshold, 80 patients (100%) in the non-drug-resistant group had a matching score below 0.60, while 68 patients (68 ÷ 70 ≈ 97.14%) in the drug-resistant group had a matching score above 0.60. Further adjustments to the threshold were made for validation. If the threshold was set to 0.59, 79 patients (98.75%) in the non-drug-resistant group had a score below this value, while 69 patients (98.57%) in the drug-resistant group had a score above this value. If the threshold was set to 0.61, 80 patients (100%) in the non-drug-resistant group had a score below this value, while 67 patients (95.71%) in the drug-resistant group had a score above this value. Taking into account the accuracy of the two groups, when 0.60 is used as the threshold, the accuracy of the non-drug-resistant group is 100% and the accuracy of the drug-resistant group is 97.14%, which is the best overall discrimination effect. Therefore, this value is determined as the drug resistance feature matching threshold, and the drug resistance feature matching threshold is 0.60.
[0034] S502: Call the patient's drug resistance feature matching degree and drug resistance feature matching threshold, compare the values, classify those above the threshold as drug-resistant, and classify those below the threshold as non-drug-resistant, and generate the screening results of immunotherapy drug-resistant population.
[0035] The drug resistance trait matching data of all patients to be screened were retrieved, along with the established drug resistance trait matching threshold of 0.60. The values were compared one by one according to the individual patient identifier. Taking 10 patients to be screened as an example, the matching degree of patient P001 was 0.83, P002 was 0.75, P003 was 0.68, P004 was 0.62, P005 was 0.59, P006 was 0.55, P007 was 0.72, P008 was 0.48, P009 was 0.65, and P010 was 0.53. Each patient's match score was compared to a threshold of 0.60. Patients with match scores of P001 (0.83 > 0.60), P002 (0.75 > 0.60), P003 (0.68 > 0.60), P004 (0.62 > 0.60), P007 (0.72 > 0.60), and P009 (0.65 > 0.60) were classified as drug-resistant. Patients with match scores of P005 (0.59 < 0.60), P006 (0.55 < 0.60), P008 (0.48 < 0.60), and P010 (0.53 < 0.60) were classified as non-drug-resistant. This comparison process was repeated for all patients to be screened. Each patient's individual identifier, drug resistance characteristic match score, comparison results, and classification were recorded. The data was then categorized and organized into a structured table containing basic patient information, match score data, and classification, generating the screening results for immunotherapy-resistant individuals.
[0036] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for screening patients resistant to immunotherapy in triple-negative breast cancer, characterized in that, Includes the following steps: S1: Obtain gene expression data of triple-negative breast cancer patients, calculate immune and matrix-related indicators, obtain corresponding quantitative values, and generate immune matrix score comparison dataset; S2: Collect mRNA sequence data of single-cell samples from triple-negative breast cancer patients, remove low-quality data and normalize, filter variable genes by dimensionality reduction and then cluster, annotate fibroblasts and macrophages, perform secondary clustering to quantify the proportion of subpopulations, screen drug resistance-associated subpopulations, and obtain a list of drug resistance-associated cell subpopulations. S3: Extract gene expression data of MCAM+ fibroblast subsets, compare gene expression differences with other subsets, screen for specifically expressed genes, and establish a combination of drug resistance characteristic genes; S4: Obtain bulk transcriptome data of patients to be screened, detect the expression status of drug resistance characteristic genes, count the number of highly expressed genes, and obtain the matching degree of drug resistance characteristics of patients; S5: Set a drug resistance feature matching threshold, compare the patient's drug resistance feature matching degree with the threshold, determine whether the patient is drug-resistant or non-drug-resistant, and generate screening results for immunotherapy drug-resistant patients.
2. The screening method for patients resistant to immunotherapy in triple-negative breast cancer according to claim 1, characterized in that: The immune matrix score comparison dataset includes patient individual identifiers, immune-related quantitative values, and matrix-related quantitative values. The drug resistance-associated cell subpopulation list includes cell subpopulation names, cell subpopulation percentages, and the degree of association with drug resistance phenotypes. The drug resistance characteristic gene combination includes gene names, gene-specific expression markers, and gene function annotations. The patient drug resistance characteristic matching degree includes the number of highly expressed characteristic genes, the total number of characteristic genes, and the matching ratio. The immunotherapy drug resistance population screening results include patient individual identifiers, drug resistance determination conclusions, and matching degree values.
3. The screening method for patients resistant to immunotherapy in triple-negative breast cancer according to claim 1, characterized in that: The steps for obtaining the immune matrix scoring comparison dataset are as follows: S101: Obtain gene expression data of triple-negative breast cancer patients, select immune-related indicators and matrix-related indicators, perform quantitative calculations on the two types of indicators respectively, obtain the quantitative values of the corresponding indicators for each patient, and generate a basic immune matrix data table. S102: Call the basic immune matrix data table, integrate the patient's individual identifier and corresponding quantitative value, classify and organize them according to the indicator type, form a standardized data structure, and generate an immune matrix score comparison dataset.
4. The screening method for patients resistant to immunotherapy in triple-negative breast cancer according to claim 1, characterized in that: The steps for obtaining the list of drug resistance-associated cell subsets are as follows: S201: Collect single-cell samples from triple-negative breast cancer patients, extract single-cell mRNA sequence data, remove low-quality data, normalize the remaining data, reduce data dimensionality through principal component analysis, screen variable genes, and obtain a standardized single-cell gene dataset. S202: Based on a standardized single-cell gene dataset, cluster cells according to their expression characteristics, call up the clustering results of cell marker genes, annotate fibroblasts and macrophages, and obtain the target cell classification dataset; S203: For the target cell classification dataset, perform secondary clustering of the two cell types, quantify the proportion of each subgroup, screen corresponding subgroups based on the immunotherapy resistance phenotype, and generate a list of drug resistance-related cell subgroups.
5. The screening method for patients resistant to immunotherapy in triple-negative breast cancer according to claim 1, characterized in that: The steps for obtaining the drug resistance gene combination are as follows: S301: Call the drug resistance-associated cell subpopulation list, extract gene expression data of MCAM+ fibroblast subpopulation, collect gene expression data of other cell subpopulations, and obtain a multi-subpopulation gene comparison dataset. S302: Based on a multi-subpopulation gene comparison dataset, compare the gene expression differences between the MCAM+ fibroblast subpopulation and other subpopulations, screen for specifically expressed genes, and establish a combination of drug resistance characteristic genes.
6. The screening method for patients resistant to immunotherapy in triple-negative breast cancer according to claim 1, characterized in that: The steps for obtaining the patient's drug resistance characteristic matching degree are as follows: S401: Obtain bulk transcriptome gene expression data of patients to be screened, call the combination of drug resistance characteristic genes, detect the expression status of each characteristic gene one by one, record the actively expressed genes, and obtain the characteristic gene expression status table. S402: Calculate the number of actively expressed genes in the statistical characteristic gene expression status table, and then calculate the ratio of this number to the total number of characteristic genes to obtain the patient's drug resistance characteristic matching degree.
7. The screening method for patients resistant to immunotherapy in triple-negative breast cancer according to claim 1, characterized in that: The steps for obtaining the screening results of immunotherapy-resistant individuals are as follows: S501: Based on clinical immunotherapy response data and the expression patterns of drug resistance characteristic genes, a criterion for distinguishing between drug resistance and non-drug resistance was established, and a drug resistance characteristic matching threshold was obtained.
8. The method for screening patients resistant to immunotherapy in triple-negative breast cancer according to claim 7, characterized in that: S502: Call the patient's drug resistance feature matching degree and drug resistance feature matching threshold, compare the values, classify those above the threshold as drug-resistant, and classify those below the threshold as non-drug-resistant, and generate the screening results of immunotherapy drug-resistant population.