Gene marker combination, kit and system for breast cancer molecular typing

The GPS17 molecular subtyping system, constructed using a combination of 17 gene markers, addresses the shortcomings of existing molecular subtyping systems for breast cancer, enabling more precise molecular subtyping of breast cancer, improving diagnostic accuracy and treatment outcomes, and demonstrating significant industrialization potential.

CN121629054APending Publication Date: 2026-03-10SHUQI MEDICAL TECHNOLOGY (SUZHOU) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610024737.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-09
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing molecular subtyping systems for breast cancer, especially the PAM50 molecular subtyping system, are insufficient in analyzing the molecular heterogeneity of breast cancer. In particular, the heterogeneity within the Luminal A subtype is difficult to distinguish effectively, resulting in large differences in treatment effects and survival outcomes, and failing to meet the clinical demand for highly accurate molecular subtyping.

Method used

A GPS17 molecular subtyping system was constructed using a combination of 17 gene markers (CLDN19, DSG1, FABP7, FGFBP1, KLK7, KLK8, MAPK4, MAPT-AS1, MAPT-IT1, MIA, PTPRZ1, ROPN1, ROPN1B, SCN2B, SERPINA11, SOX11, SRARP, etc.). By training a classification model, breast cancer samples were divided into 8 subtypes, providing more refined molecular subtyping analysis.

Benefits of technology

It significantly improves the ability to resolve molecular heterogeneity, optimizes the precision diagnosis and treatment of breast cancer, and the classification model constructed by combining 17 gene markers outperforms the PAM50 system in key performance indicators. It clearly reveals the molecular heterogeneity within the Luminal A subtype, reduces detection and analysis costs, and has good prospects for industrialization and clinical translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a gene marker combination, a kit and a system for breast cancer molecular typing, and belongs to the field of precise medical big data analysis. The gene marker combination is composed of 17 gene markers, namely, CLDN19, DSG1, FABP7, FGFBP1, KLK7, KLK8, MAPK4, MAPT-AS1, MAPT-IT1, MIA, PTPRZ1, ROPN1, ROPN1B, SCN2B, SERPINA11, SOX11, and SRARP. The invention further discloses a preparation method of the gene marker combination. Based on the combination, a brand new breast cancer GPS17 molecular typing system is constructed, the breast cancer can be finely divided into eight molecular subtypes with remarkable expression difference, and at least four different molecular subgroups in the Luminal A subtype are systematically analyzed for the first time. Experiments show that the molecular typing system is obviously superior to a traditional PAM50 molecular typing system in balance accuracy, macro F1 score and macro AUC score. The invention further provides a corresponding detection kit and an automatic typing system integrated with a machine learning model, and an auxiliary diagnosis tool with excellent performance is provided for accurate diagnosis and treatment decision of breast cancer.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of precision medicine big data analysis, in particular to a gene marker combination for breast cancer molecular subtyping, a kit for detecting the expression level of the combination, and a system for analyzing the expression data of the combination. BACKGROUND

[0002] Breast cancer is a malignant tumor with high molecular heterogeneity. Molecular subtyping is an important prerequisite for accurate diagnosis and effective treatment. Currently, the widely used molecular subtyping system for breast cancer in clinical and scientific research is PAM50, which divides breast cancer into Normal-like, Basal-like, HER2-enriched, Luminal A, and Luminal B subtypes. Different treatment regimens are used clinically for different PAM50 subtypes. However, with further research into the molecular heterogeneity of breast cancer, the limitations of the PAM50 molecular subtyping system have become increasingly apparent: patients with the same PAM50 subtype, especially those with the Luminal A subtype, have significantly different treatment outcomes and survival outcomes. This suggests that the existing molecular subtyping system is insufficient in analyzing the molecular heterogeneity of breast cancer, particularly within the Luminal A subtype, and fails to fully meet the urgent need for high-precision molecular subtyping systems in clinical practice.

[0003] Therefore, there is an urgent need in the art for a more effective molecular subtyping system for breast cancer that can more finely analyze the heterogeneity of breast cancer and improve the level of accurate diagnosis and treatment of breast cancer. SUMMARY

[0004] The present application aims to overcome the limitations of existing molecular subtyping systems in analyzing the molecular heterogeneity of breast cancer, particularly within the Luminal A subtype, and provide a new molecular subtyping system with stronger molecular heterogeneity analysis capabilities.

[0005] In a first aspect, the present application provides a gene marker combination for breast cancer molecular subtyping, characterized in that the combination consists of the following 17 gene markers: CLDN19, DSG1, FABP7, FGFBP1, KLK7, KLK8, MAPK4, MAPT-AS1, MAPT-IT1, MIA, PTPRZ1, ROPN1, ROPN1B, SCN2B, SERPINA11, SOX11, and SRARP. The GPS17 (17-Gene Profile Subtyping) molecular subtyping system refers to a molecular subtyping system that divides breast cancer samples into 8 subtypes with different molecular characteristics and clinical significance based on the expression data of the 17 gene markers through a trained classification model.

[0006] In a second aspect, the present application provides a detection kit for breast cancer molecular typing, comprising reagents for detecting the expression levels of the 17 gene markers in the first aspect.

[0007] In a third aspect, the present application provides a breast cancer molecular typing system, comprising: (1) a data input module for receiving expression data of the 17 gene markers in a sample to be analyzed; and (2) a data analysis module embedded with a classification model using the expression data of the 17 gene markers as necessary input features, for outputting a breast cancer molecular typing result based on the GPS17 molecular typing system defined in the present application.

[0008] In a fourth aspect, the present application provides use of the gene marker combination, kit and system in the preparation of a product for breast cancer molecular typing.

[0009] In a fifth aspect, the present application provides an electronic device and a computer readable storage medium having stored thereon a computer program, which, when executed by a processor, can realize the functions of the molecular typing system. Advantages

[0010] Compared with the prior art, the present application has the following advantages: (1) significantly enhanced molecular heterogeneity analysis capability: the present application first defines a new breast cancer molecular typing system GPS17 based on a set of 17 carefully selected gene markers, which finely divides breast cancer into 8 molecular subtypes with significantly different expression patterns. Through experimental verification, the classification model constructed based on this combination is significantly superior to the PAM50 molecular typing system in terms of balanced accuracy, macro F1 score and macro AUC score, proving that the new system has stronger molecular heterogeneity analysis capability; (2) first systematic analysis of molecular heterogeneity within Luminal A subtype: the present application clearly reveals and confirms that there are at least 4 molecular subtypes within the Luminal A subtype that have significantly different expression patterns of tumor driver genes, providing a new perspective for solving the heterogeneity problem of this subtype; (3) efficient and easy-to-apply gene marker set: the present application selects 17 gene markers from a large number of gene markers, which greatly reduces the detection and analysis cost while maintaining excellent subtype discrimination ability, and has good industrialization and clinical transformation prospects; (4) provides an end-to-end solution: the present application covers the whole process from core biomarker discovery and verification to detection product and intelligent analysis system development, providing an excellent and accessible auxiliary diagnostic solution for precise molecular diagnosis and individualized treatment decision-making of breast cancer. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1Figure 1 shows a schematic diagram of the branch process of K-Means clustering analysis based on the expression data of 49 gene markers, and the number of samples in each cluster, the dominant PAM50 subtype, and the sample purity are marked.

[0012] Figure 2 To drive Figure 1 The z-score heat map of differentially expressed gene markers in the cluster branch process intuitively shows the unique gene expression patterns between different clusters.

[0013] Figure 3 Figure 2 shows a comparison of the typing effect of GPS17 and PAM50 molecular typing systems based on the expression data of 17 gene markers in the present application after t-SNE dimensionality reduction visualization, showing that the new molecular typing system has a clearer class separation degree. DETAILED DESCRIPTION

[0014] The present application will be further described in detail below in conjunction with the drawings and examples. The following examples are only used to explain the present application and do not constitute a limitation on the scope of protection of the present application. Example 1

[0015] Discovery of core gene markers and establishment of new molecular typing system: (1) Data preparation: Obtain multi-omics data of 875 breast cancer samples from the Cancer Genome Atlas (TCGA) database, including mRNA expression data, DNA methylation data and microRNA expression data, then perform variance filtering on the original features, and retain the top 1000 features with the highest variability for mRNA features; (2) Feature selection: Use the multi-omics data integration and classification method based on hierarchical attention (see application number: 202511477493.2) to select the 49 mRNA features with the highest contribution to the differentiation of PAM50 subtypes from the above 1000 mRNA features; (3) Unsupervised clustering and new subtype definition: Use the above 49 mRNA features to perform hierarchical K-Means clustering analysis on all samples, and finally obtain 8 final clusters corresponding to 8 different GPS17 subtypes. By analyzing the differentially expressed genes driving the cluster branch process, 17 mRNA features with the highest contribution to the differentiation of the 8 final clusters are further selected from the 49 mRNA features, i.e., the gene marker combination of the present application. DAVID functional clustering analysis shows that these genes are mainly involved in the regulation and remodeling of tumor microenvironment, providing important clues for analyzing breast cancer molecular heterogeneity. (4) Analysis of molecular characteristics of new subtypes: As shown in Figure 3, the 8 GPS17 subtypes have different molecular characteristics, such as different PAM50 subtypes, different DNA methylation patterns, different microRNA expression patterns, and different gene expression patterns. Figure 1As shown, the newly defined 8 GPS17 subtypes are associated with PAM50 subtypes but are more discriminative. Four of them are mainly dominated by non-Luminal A subtype, including R1C1_R2C2_R3C2 (Normal-like, 0.514), R1C1_R2C1 (Basal-like, 0.934), R1C2_R2C1 (HER2-enriched, 0.548) and R1C2_R2C3_R3C2 (Luminal B, 0.631), while the other four are all dominated by Luminal A subtype, among which the more aggressive R1C1_R2C2_R3C1 (Luminal A, 0.420) is derived from R1C1 (Basal-like, 0.385), while the less aggressive R1C2_R2C2 (Luminal A, 0.957), R1C2_R2C3_R3C1 (Luminal A, 0.763) and R1C2_R2C4 (Luminal A, 0.724) are derived from R1C2 (Luminal A, 0.680). As shown, these four Luminal A dominated GPS17 subtypes are significantly different in the expression pattern of the 17 gene markers, which for the first time confirmed the existence of stably distinguishable molecular heterogeneity subgroups within Luminal A subtype from the computational biology level. Figure 2 As shown, these four Luminal A dominated GPS17 subtypes are significantly different in the expression pattern of the 17 gene markers, which for the first time confirmed the existence of stably distinguishable molecular heterogeneity subgroups within Luminal A subtype from the computational biology level. Example 2

[0016] Performance verification of the new subtyping method: (1) Model construction: Using only the expression data of the 17 gene markers determined in Example 1 as input features, two random forest classifiers were constructed: the first used the PAM50 subtype label (5 classes) as the prediction target, and the second used the GPS17 subtype label (8 classes) defined in this invention as the prediction target; (2) Stability evaluation: To comprehensively evaluate the stability of the model, the dataset was split into a training set (60%), a validation set (20%), and a test set (20%) using 10 different random seeds, while maintaining a consistent proportion of each subtype in different subsets. For each data split, the model was initialized for training using 10 different random seeds. Finally, statistical analysis was performed based on the results of 100 independent experiments. (3) Performance comparison results: On the test set, the GPS17 classification model significantly outperformed the PAM50 classification model in several core performance metrics, including balanced accuracy (0.845 ± 0.046 vs 0.670 ± 0.031, p < 0.00001), macro F1 score (0.837 ± 0.048 vs 0.648 ± 0.035, p < 0.00001), and macro AUC score (0.986 ± 0.005 vs 0.887 ± 0.017, p < 0.00001), with performance improvements of 26.1%, 29.2%, and 11.2%, respectively; (4) Visual comparison: such as Figure 3 As shown, the expression data of 17 gene markers from all samples were subjected to t-SNE dimensionality reduction and visualized. The decision boundary of the GPS17 molecular subtyping system was clear, and the different categories were well separated; while the decision boundary of the PAM50 molecular subtyping system was blurred, and there were a large number of overlapping regions between different categories. This intuitively demonstrates that the gene combination and subtyping system provided by this invention has a stronger ability to resolve molecular heterogeneity in breast cancer. Example 3

[0017] Construction of the detection kit: based on the specific sequences of the 17 gene markers of the present application, a fluorescence quantitative PCR kit suitable for clinical sample detection is constructed. (1) Primer and probe design: specific primers and TaqMan hydrolysis probes are designed and synthesized for each target gene specific transcript region using professional software. All sequences are verified by bioinformatics comparison to ensure their specificity and avoid cross-reaction; (2) The kit mainly includes the following components: RNA reverse transcription premix, fluorescence quantitative PCR reaction premix, specific primers and TaqMan probe premix of 17 target genes and 3 internal reference genes, positive quality control and negative quality control; (3) The detection process is as follows: total RNA is extracted from breast cancer fresh frozen tissue or paraffin-embedded tissue, and the qualified total RNA is reverse transcribed into cDNA; then the cDNA is subjected to fluorescence quantitative PCR amplification, and the corrected expression data of the target genes are obtained, and finally the obtained data are used as the input of the molecular typing system in Example 4. Example 4

[0018] Implementation of the molecular typing software system: a software system for breast cancer molecular typing is developed to realize automatic analysis from data to results. The core functional modules include: (1) Data input module: support manual input or batch import of the expression data of 17 gene markers detected by the kit of Example 3, and perform data standardization processing; (2) Data analysis module: embedded a pre-trained high-performance random forest classification model as described in Example 2, this module receives the standardized expression data vector, automatically calculates the probability of the sample belonging to the 8 GPS17 molecular subtypes defined in the present application, and provides auxiliary diagnostic content such as clinical association information related to the subtypes; this system can be packaged as a standalone desktop software installed on the local server of the hospital pathology department or third-party medical testing laboratory, or can be deployed on a private cloud platform that meets the medical data security specifications, providing online analysis services through a browser or API interface. Clinical application value

[0019] The hospital pathology department or third-party medical testing laboratory can use the kit of the present application to detect breast cancer patient tissue samples, obtain the expression data of 17 gene markers, and use the matching molecular typing software system to obtain the fine GPS17 molecular typing results. This result can provide more accurate breast cancer molecular typing for clinicians, assist them in developing more effective treatment plans, and ultimately improve the prognosis of patients. The number of gene markers in the present application is small, the detection technology is mature, the model deployment is convenient, the analysis process is automated, and it has good industrialization and clinical transformation prospects.

[0020] The above merely describes preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art, according to the technical solution and inventive concept of the present application, makes equivalent replacement or change within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.

Claims

1. A combination of genetic markers for breast cancer molecular typing, characterized in that, The 17 gene markers consist of CLDN19, DSG1, FABP7, FGFBP1, KLK7, KLK8, MAPK4, MAPT-AS1, MAPT-IT1, MIA, PTPRZ1, ROPN1, ROPN1B, SCN2B, SERPINA11, SOX11, SRARP, and the expression data of this combination is a necessary and sufficient feature set to implement the GPS17 molecular typing system defined by the present application.

2. A test kit for breast cancer molecular typing, characterized by, The kit comprises reagents for detecting the expression levels of the 17 gene markers of claim 1, and is configured to exclusively acquire the expression data of the 17 gene markers.

3. The test kit according to claim 2, characterized in that, The kit is based on nucleic acid detection technology, including but not limited to fluorescent quantitative PCR technology or high-throughput sequencing technology.

4. A breast cancer molecular subtyping system characterized in that, The kit comprises: a data input module for receiving the expression data of the 17 gene markers of claim 1; a typing calculation module embedded with a classification model taking the expression data of the 17 gene markers as necessary input features, for outputting the breast cancer molecular typing result based on the GPS17 molecular typing system defined by the present application.

5. The breast cancer molecular typing system according to claim 4, wherein, The classification model is a machine learning model, including but not limited to a random forest model or a support vector machine.

6. Use of the gene marker combination of claim 1 in breast cancer molecular typing.

7. Use of the detection kit of claim 2 or 3 in breast cancer molecular typing.

8. Use of the breast cancer molecular typing system of claim 4 or 5 in breast cancer molecular typing. 9.An electronic device comprising a memory and a processor, the memory storing a computer program, wherein, The processor executes the computer program to realize the functions of the breast cancer molecular typing system of claim 4 or 5.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the functions of the breast cancer molecular typing system of claim 4 or 5.

Citation Information

Patent Citations

  • Multi-omics data integration and classification method, system and equipment based on hierarchical attention

    CN121483380A