Method for noninvasive prediction of breast cancer molecular typing and application thereof
By using a machine learning model based on cfDNA features, the problems of high invasiveness and low accuracy of existing breast cancer subtyping methods have been solved. This method achieves non-invasive and accurate prediction of breast cancer molecular subtypes, improving subtyping efficiency and accuracy, and is suitable for the formulation of personalized treatment strategies.
Patent Information
- Application Number
- CN202610043911.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-14
- Publication Date
- 2026-02-10
AI Technical Summary
Existing breast cancer classification methods are either highly invasive or have low predictive accuracy, making them difficult to widely apply in large-scale screening and recurrence monitoring.
A breast cancer subtyping prediction method based on cfDNA feature extraction and machine learning model is proposed. By constructing a machine learning model, the signal intensity of cfDNA features such as LTR repeat sequence regions is utilized, combined with multi-dimensional molecular feature information, to achieve non-invasive prediction of breast cancer molecular subtyping.
It achieves accurate prediction of the three major molecular subtypes of breast cancer, improves the accuracy of subtype prediction and model robustness, and is suitable for patients with dynamic monitoring and those for whom tumor tissue is difficult to obtain, thus assisting in the development of personalized treatment strategies.
Smart Images

Figure CN121506252A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tumor molecular diagnostics, specifically to a non-invasive method for predicting molecular subtyping of breast cancer and its application. Background Technology
[0002] Breast cancer is the most common malignant tumor among women worldwide and one of the leading causes of cancer death. Breast cancer exhibits high heterogeneity, with different subtypes showing significant differences in pathogenesis, treatment response, and prognosis. Currently, the primary clinical approach is the immunohistochemical (IHC) trityping method, classifying breast cancer into ER+ / PR+, HER2+, and triple-negative breast cancer (TNBC) based on the expression of estrogen receptor (ER), progesterone receptor (PR), and human epidermal growth factor receptor 2 (HER2). This method is simple, practical, and highly instructive. ER+ / PR+ breast cancer is commonly treated with endocrine therapy, HER2+ breast cancer can be treated with targeted therapies, while TNBC, due to the lack of target therapies, is difficult to treat and has the worst prognosis. This classification has significant value in clinical diagnosis and treatment.
[0003] Currently, commonly used typing methods rely on histopathological examination, immunohistochemical analysis, and gene expression profiling. These methods are typically invasive, costly, or time-consuming, limiting their application in large-scale screening and recurrence monitoring. cfDNA, a cell-free DNA molecule released into peripheral blood during tumor cell apoptosis or necrosis, has become a novel biomarker for early tumor detection and typing due to its non-invasiveness and high reproducibility. However, the low concentration of cfDNA in blood and its poor signal-to-noise ratio mean that extracting effective typing information and achieving accurate prediction remains a technological challenge. Summary of the Invention
[0004] To address the issues of high invasiveness or low prediction accuracy in existing methods, this invention provides a breast cancer subtyping prediction method based on cfDNA feature extraction and machine learning models.
[0005] In a first aspect of the present invention, a method for constructing a non-invasive predictive model for molecular subtyping of breast cancer is provided, the method comprising constructing a machine learning model based on sample cfDNA features to obtain a non-invasive predictive model for molecular subtyping of breast cancer.
[0006] In another preferred embodiment, the cfDNA feature includes the signal intensity of the LTR repeat sequence region.
[0007] In another preferred embodiment, the following steps are included: (S1) Provides sequencing results of cfDNA in the sample; (S2) The sequencing results are preprocessed and quality controlled to obtain sequencing data that meets the quality control conditions; (S3) Extract cfDNA features from sequencing data that meet quality control conditions, wherein the cfDNA features include: signal intensity of LTR repeat sequence regions; (S4) Construct a machine learning model based on the cfDNA features to obtain a non-invasive prediction model for molecular subtyping of breast cancer; The signal intensity of the LTR repeat sequence region refers to the ratio of the number of cfDNA fragments whose 5' end or midpoint falls into the LTR repeat sequence region to the total number of cfDNA fragments.
[0008] In another preferred embodiment, the sample comprises a healthy sample.
[0009] In another preferred embodiment, the sample comprises a breast cancer sample.
[0010] In another preferred embodiment, the breast cancer sample comprises breast cancer samples of the ER+ / PR+HER2- subtype.
[0011] In another preferred embodiment, the breast cancer sample comprises a HER2+ subtype breast cancer sample.
[0012] In another preferred embodiment, the breast cancer sample comprises a breast cancer sample with a TNBC subtype.
[0013] In another preferred embodiment, the sample comprises: a blood sample, a urine sample, a fecal sample, a tear sample, or a plasma sample.
[0014] In another preferred embodiment, the sample is a plasma sample.
[0015] In another preferred embodiment, step (S1) includes the following sub-steps: (S1-1) Provide samples; (S1-2) Extract cfDNA from the sample; (S1-3) Sequencing the cfDNA to obtain the sequencing results of the cfDNA in the sample.
[0016] In another preferred embodiment, the extraction includes extraction using magnetic beads or extraction using a dedicated cfDNA extraction kit.
[0017] In another preferred embodiment, the sequencing is ultra-low coverage whole-genome sequencing.
[0018] In another preferred embodiment, the sequencing is performed using a 150 bp paired-end sequencing method.
[0019] In another preferred embodiment, the average sequencing depth is 2×-10×, preferably 2×-5×, and more preferably 2×.
[0020] In another preferred embodiment, the sequencing is performed via the Illumina NovaSeq X Plus platform.
[0021] In another preferred embodiment, the preprocessing and quality control includes: removing low-quality reads, removing connector contamination sequences, removing abnormal fragments, or a combination thereof.
[0022] In another preferred embodiment, the quality control includes quality control of sequencing depth, coverage, and GC content deviation.
[0023] In another preferred embodiment, the preprocessing and quality control includes: quality assessment, removal of low-quality reads, removal of adapter sequences, removal of contaminating fragments, alignment to a human reference genome, removal of PCR duplications, sorting, indexing, identification and removal of repetitive sequences, or combinations thereof.
[0024] In another preferred embodiment, the human reference genome is the GRCh37 version of the human reference genome.
[0025] In another preferred embodiment, the reference genome is divided into non-overlapping window regions of 5 Mb in size.
[0026] In another preferred embodiment, preprocessing and quality control include performing quality control and alignment processing on the sequencing results using one or more tools selected from the group consisting of FastQC, Cutadapt, Ktrim, BWA-MEM, Samtools, Picard tools, or combinations thereof.
[0027] In another preferred embodiment, the breast cancer molecular subtyping includes: ER+ / PR+HER2- type, HER2+ type, and TNBC type.
[0028] In another preferred embodiment, the cfDNA feature further includes one or more of the following dimensions of cfDNA features: fragmentation feature, coverage feature, genetic feature, and epigenetic feature.
[0029] In another preferred embodiment, the cfDNA features further include one or more of the following dimensions of cfDNA features: cfDNA fragment distribution, cfDNA sequence, and cfDNA methylation signal.
[0030] In another preferred embodiment, the total number of extracted cfDNA features is 1, preferably 1-3, and most preferably 1-8.
[0031] In another preferred embodiment, the total number of extracted cfDNA features is 1-200.
[0032] In another preferred embodiment, the total number of extracted cfDNA features is 200-9999.
[0033] In another preferred embodiment, the cfDNA feature further includes one or more cfDNA features selected from the group consisting of: fixed-window fragment enrichment features, base preference features, transposon element region coverage features, cfDNA fragment length distribution, motif sequence features, nucleosome localization features, copy number variation (CNV) features, and methylation motif distribution features.
[0034] In another preferred embodiment, the transducer element region includes: a SINE transducer element region, a LINE transducer element region, or a combination thereof.
[0035] In another preferred embodiment, the LTR repeat sequence region comprises an endogenous retrovirus family sequence region.
[0036] In another preferred embodiment, the endogenous retrovirus family sequence region includes: ERV1, ERVK, ERVL, ERVL-MaLR, or a combination thereof.
[0037] In another preferred embodiment, the SINE transposon element region comprises: an Alu element, a MIR element, a tRNA-derived SINE sequence region, or a combination thereof.
[0038] In another preferred embodiment, the LINE transpose element region includes a long, dispersed core element sequence region.
[0039] In another preferred embodiment, the long, dispersed nuclear element sequence region includes: L1 element, L2 element, CR1, or a combination thereof.
[0040] In another preferred embodiment, the cfDNA fragment length distribution refers to the ratio S / L of the number of short fragment length cfDNA and the number of long fragment length cfDNA.
[0041] In another preferred embodiment, the short fragment length cfDNA is cfDNA with a fragment length of 100-150 bp.
[0042] In another preferred embodiment, the long fragment length cfDNA is cfDNA with a fragment length of 151-220 bp.
[0043] In another preferred embodiment, the cfDNA fragment length distribution excludes the ENCODE blacklist region and the UCSC gaptrack region.
[0044] In another preferred embodiment, the motif sequence feature refers to the frequency of occurrence of the 6-mer motif corresponding to each cfDNA fragment in the sample; The 6-mer motif comprises three bases upstream and three bases downstream of the 5' end of the cfDNA fragment.
[0045] In another preferred embodiment, the copy number variation (CNV) feature refers to a CNV pattern covering the entire cfDNA genome.
[0046] In another preferred embodiment, the CNV map is a CNV map covering the entire genome constructed using a sliding window of 5 Mb units.
[0047] In another preferred embodiment, the nucleosome localization feature refers to the coverage density of cfDNA fragments within ±1000 bp of the center of each CTCF high-confidence binding site.
[0048] In another preferred embodiment, a symmetrical analysis window region is constructed within ±1000 bp of the center of each CTCF high-confidence binding site, and this region is divided into bins of fixed size to accurately count the coverage density of cfDNA fragments, thereby generating a meta-coverage map.
[0049] In another preferred embodiment, the fixed-size bin has a length of 10 bp.
[0050] In another preferred embodiment, the fragment enrichment feature of the fixed window refers to the coverage density or count density of cfDNA fragments in each fixed-size window.
[0051] In another preferred embodiment, the fixed-size window has a length of 10 bp.
[0052] In another preferred embodiment, the methylation motif distribution characteristic refers to the frequency of the 5' end of the cfDNA fragment appearing in a specific flanking region of a methylation-sensitive motif; The specific flanking region refers to the three bases upstream of the methylation-sensitive motif and the three bases downstream of the methylation-sensitive motif.
[0053] In another preferred embodiment, the 5' end of the cfDNA comprises three bases upstream and three bases downstream of the 5' end of the cfDNA fragment.
[0054] In another preferred embodiment, the base preference feature includes single base frequencies upstream and downstream of the 5' end of the cfDNA fragment; and / or Frequency of dibase combinations upstream and downstream of the 5' end of the cfDNA fragment.
[0055] In another preferred embodiment, the machine learning model includes: a generalized linear model (GLM), a deep neural network model (DNN), or a lightweight gradient booster model (LightGBM).
[0056] In another preferred embodiment, the machine learning model is a generalized linear model (GLM).
[0057] In another preferred embodiment, step (S4) further includes the following sub-steps: (S4-1) The cfDNA features are normalized and standardized to obtain standardized feature data; (S4-2) A machine learning model is constructed based on the standardized feature data. The model is trained and its performance is evaluated, thereby providing a non-invasive prediction model for molecular subtyping of breast cancer.
[0058] In another preferred embodiment, 10-fold cross-validation is used for model training and performance evaluation.
[0059] In another preferred embodiment, the performance evaluation metrics include: AUC (area under the curve), sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), or a combination thereof.
[0060] In a second aspect of the invention, a non-invasive predictive model for molecular subtyping of breast cancer is provided, the non-invasive predictive model being constructed using the construction method described in the first aspect of the invention.
[0061] In a third aspect of the invention, the use of the non-invasive prediction model described in the second aspect of the invention is provided for preparing a non-invasive prediction system for predicting molecular subtyping of breast cancer.
[0062] In a fourth aspect of the invention, a non-invasive prediction system for molecular subtyping of breast cancer is provided, the system comprising the following modules: (Z1) Data input module, which is configured to input the sequencing results of cfDNA in the sample to be analyzed; (Z2) Preprocessing and quality control module, wherein the preprocessing and quality control module is configured to preprocess and control the sequencing results to obtain sequencing data that meets the quality control conditions; (Z3) Feature extraction and selection module, the feature extraction and selection module is configured to: extract cfDNA features from sequencing data that meet quality control conditions, wherein the cfDNA features include the signal intensity of LTR repeat sequence regions; (Z4) Breast cancer molecular typing evaluation module, wherein the breast cancer molecular typing evaluation module is configured to: predict the breast cancer molecular typing result of the sample to be analyzed based on the cfDNA characteristics by means of the non-invasive prediction model of breast cancer molecular typing described in the second aspect of the present invention; (Z5) Output module, which is configured to output the molecular subtyping results of the breast cancer of the sample to be analyzed.
[0063] In another preferred embodiment, the signal intensity of the LTR repeat sequence region refers to the ratio of the number of cfDNA fragments whose 5' end or midpoint falls into the LTR repeat sequence region to the total number of cfDNA fragments.
[0064] In another preferred embodiment, the cfDNA feature further includes one or more cfDNA features selected from the group consisting of: fixed-window fragment enrichment features, base preference features, transposon element region coverage features, cfDNA fragment length distribution, motif sequence features, nucleosome localization features, copy number variation (CNV) features, and methylation motif distribution features.
[0065] In another preferred embodiment, module (Z1) comprises the following sub-modules: (Z1-1) Sample providing module, the sample providing module being configured to: provide samples; (Z1-2) cfDNA extraction module, wherein the cfDNA extraction module is configured to extract cfDNA from the sample; (Z1-3) Sequencing module, the sequencing module being configured to sequence the cfDNA to obtain the sequencing results of cfDNA in samples of each breast cancer molecular subtype.
[0066] In another preferred embodiment, the sample is a breast cancer sample.
[0067] In another preferred embodiment, the sample is a late-stage breast cancer sample.
[0068] In another preferred embodiment, the sample is a plasma sample.
[0069] It should be understood that, within the scope of this invention, the above-described technical features of this invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be described in detail here. Attached Figure Description
[0070] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0071] Figure 1 This document illustrates a flowchart of the breast cancer molecular subtyping prediction method provided in an embodiment of this application. The left figure shows the framework of the steps for constructing a non-invasive prediction model for breast cancer molecular subtyping, while the right figure shows a detailed flowchart of the research and development process.
[0072] Figure 2 The results show the predictions for the ER+|PR+HER2- genotyping training set based on multiple models.
[0073] Figure 3 The results show the predictions for the ER+|PR+HER2- genotyping validation set based on multiple models.
[0074] Figure 4 The results show the predictions for the HER2+ genotyping training set based on multiple models.
[0075] Figure 5 The results show the predictions for the HER2+ genotyping validation set based on multiple models.
[0076] Figure 6 The results show the predictions for the TNBC subtyping training set based on multiple models.
[0077] Figure 7 The results show the TNBC typing validation set predictions based on multiple models.
[0078] Figure 8 The training results of the GLM model based on the signal intensity of the LTR repeat sequence region are shown.
[0079] Figure 9 The performance of the GLM model based on the signal intensity of the LTR repeat sequence region is shown on the validation set.
[0080] Figure 10 The results show the predictions of the GLM model training set based on fragment length distribution, motif sequence, and nucleosome localization features.
[0081] Figure 11 The results of the GLM model on the validation set are shown.
[0082] Figure 12 The graph shows a comparison of ROC curves for the TuFEst-MS model training dataset; Figure 13 The graph shows a comparison of ROC curves for the TuFEst-MS model validation dataset; Figure 14 Box plots showing significant differences in fragment features between molecular subtypes of breast cancer were displayed, showing that the ER+|PR+ subtype had higher fragment features than the HER2+ subtype and TNBC subtype; Figure 15 Box plots showing significant differences in fragment features between molecular subtypes of breast cancer were displayed, showing that the HER2+ subtype was higher than the ER+|PR+ subtype and the TNBC subtype in fragment features; Figure 16Box plots showing significant differences in fragment features between molecular subtypes of breast cancer were displayed, showing that the TNBC subtype was higher than the ER+|PR+ subtype and the HER2+ subtype in fragment features; Figure 17 The results show that the predicted scores of the real ER+ / PR+ / HER2- genotype samples in the TuFEst-MS prediction model are significantly higher than those of the HER2+ and TNBC genotypes (*p < 0.05, *p < 0.001). The box-shaped lines in the figure represent the score distribution, and the median is represented by a horizontal line.
[0083] Figure 18 The results show that the predicted score of the real HER2+ genotype samples in the TuFEst-MS prediction model is significantly higher than that of the ER+ / PR+ / HER2- and TNBC genotypes (*p < 0.05, *p < 0.001). The box bars in the figure represent the score distribution, and the median is represented by a horizontal line.
[0084] Figure 19 The results show that the predicted score in the TuFEst-MS model for real TNBC genotyping samples is significantly higher than that for ER+ / PR+ / HER2- and HER2+ genotyping (*p < 0.05, *p < 0.001). The box-shaped lines in the figure represent the score distribution, and the median is represented by a horizontal line.
[0085] Figure 20 The Sankey diagram shows the matching between the prediction results of the TuFEst-MS breast cancer subtyping model and the true labels for patients with advanced breast cancer. Detailed Implementation
[0086] Through extensive and in-depth research, the inventors unexpectedly discovered a model for predicting molecular subtypes of breast cancer for the first time. Specifically, by using the signal intensity of the LTR repeat sequence region as the extracted cfDNA feature to construct a machine learning model, accurate predictions of the three major molecular subtypes of breast cancer (ER+ / PR+HER2-, HER2+, and TNBC) can be achieved, with AUC values exceeding 0.78. This performance is essentially equivalent to the prediction results of machine learning models constructed using three common cfDNA features (fragment length distribution, motif sequence, and nucleosome localization features), which have AUC values exceeding 0.81. This invention was developed based on this discovery.
[0087] This invention discloses a non-invasive method for predicting the molecular subtypes of breast cancer based on multi-feature cfDNA molecular omics and its applications. This method uses cfDNA samples as a foundation, integrates multi-dimensional molecular feature information, and employs a supervised machine learning algorithm to construct a classification model, achieving accurate prediction of the three major molecular subtypes of breast cancer. This invention achieves non-invasive detection and is suitable for patient groups requiring dynamic monitoring or where tumor tissue is difficult to obtain. By integrating multiple cfDNA features such as methylation modification, gene mutation, fragment length distribution, and nucleosome localization, it comprehensively characterizes the molecular features related to breast cancer, significantly improving the accuracy of subtype prediction and the robustness of the model. This method can effectively distinguish the three molecular subtypes of breast cancer patients. Considering the significant differences between different subtypes in survival prognosis, immune response, and treatment sensitivity, this method is expected to assist clinicians in developing personalized treatment strategies, improving efficacy, and reducing breast cancer mortality. Therefore, it is expected to serve as an important supplementary means to tissue biopsy, especially playing a key role in identifying HER2 low expression status and making personalized treatment decisions.
[0088] the term To facilitate a clearer understanding of this disclosure, certain terms are first defined. As used herein, unless otherwise expressly specified herein, each of the following terms shall have the meaning given below. Other definitions are set forth throughout the application.
[0089] As used herein, the term “and / or” refers to and covers any and all possible combinations of one or more of the related listed items.
[0090] As used herein, the terms “comprising,” “including,” and “containing” are used interchangeably and include not only closed definitions but also semi-closed and open definitions. In other words, the terms include “consisting of” and “substantially consisting of”.
[0091] As used in this article, the term "coverage density" refers to the density of sequencing reads within a specific region of the genome, usually measured by the average sequencing depth per unit length.
[0092] As used in this article, the term "count density" refers to the number of sequencing reads mapped per unit length of genome or per unit region, used to measure the level of gene or transcript expression in a specific region.
[0093] As used in this paper, the terms “nucleosome localization features”, “CTCF binding site region coverage features”, and “CTCF binding site region coverage variation” are used interchangeably.
[0094] As used in this article, the terms "motif" and "motif" are used interchangeably, both referring to short segments in biological sequences (DNA, RNA, or proteins) that are conserved, have specific functions or structural significance, and are the core functional units or characteristic markers of biomolecules.
[0095] In a specific embodiment, the present invention provides a method for non-invasive prediction of breast cancer molecular subtyping based on multi-feature cfDNA molecular omics and its application, comprising the following steps: (1) Collect plasma samples from the subjects to be tested, extract cfDNA, and perform high-throughput sequencing to obtain sequencing data; (2) The sequencing data is preprocessed and feature extracted to extract multiple cfDNA features related to breast cancer, including but not limited to: Coverage of the area surrounding the CTCF binding site Signal intensity in the LTR repeat sequence region The proportions of methylation (R1), fragmentation (R2), and the ratio of methylation (R1) to fragmentation (R2) in a specified genomic window region, Y1. cfDNA fragment size distribution characteristics Breakpoint motif frequency characteristics, etc.; (3) The extracted cfDNA feature data are normalized and standardized, and dimensionality is reduced by feature selection method; (4) The standardized feature data is divided into a training set and a validation set (287 examples in the training set and 189 examples in the validation set), and 10-fold cross-validation is used for model training and evaluation. (5) Construct multiple breast cancer subtyping prediction models based on the training set, including but not limited to the generalized linear model (GLM), the LightGBM model, the XGBoost model, and the multilayer perceptron (MLP) model; (6) Train the model and evaluate its performance in breast cancer subtyping prediction. Evaluation metrics include AUC, sensitivity, specificity, positive predictive value (PPV) and negative predictive value (NPV). (7) Determine the optimal model for predicting breast cancer subtypes, preferably the GLM model.
[0096] In another preferred embodiment, the data used for model training and validation are derived from plasma cfDNA samples from patients with clinically and pathologically confirmed breast cancer.
[0097] In another preferred embodiment, the processing of the cfDNA sample includes the following steps: (1) Plasma was separated from 8–10 mL of peripheral blood using a double centrifugation method; (2) Extract cfDNA using a magnetic bead cfDNA extraction kit; (3) Construct the library and perform sequencing on the Illumina platform; (4) Use tools such as FastQC, Cutadapt, Ktrim, BWA-MEM, Samtools and Picard to perform quality control and comparison processing on the WGS raw data.
[0098] In another preferred embodiment, the breast cancer classification is a three-tier system, including: ER+ / PR+HER2-, HER2+, and triple-negative breast cancer (TNBC).
[0099] In another preferred embodiment, the total number of cfDNA features extracted in step (2) is 9999.
[0100] In another preferred embodiment, the cfDNA fragment size distribution characteristics include: cfDNA fragment size and the ratio of the number of small cfDNA fragments to large cfDNA fragments based on the genome partitioning window (S / L ratio) (the ratio of 100–150 bp fragments to 151–220 bp fragments), wherein the ENCODE blacklist region and the UCSC gap track region are excluded.
[0101] In another preferred embodiment, the cfDNA breakpoint motif features are 4-base and 6-base motif frequencies, normalized to a vector summing to 1.
[0102] In another preferred embodiment, feature selection is performed using the Lasso regression method.
[0103] In another preferred embodiment, the high-throughput sequencing is performed using the Illumina NovaSeq X Plus platform, with a sequencing type of 150 bp paired-end sequencing and a sequencing depth of approximately 2×.
[0104] In another preferred embodiment, the training and evaluation process employs a strategy that combines bootstrapping with 10-fold cross-validation.
[0105] In another preferred example, during model evaluation, the AUC for predicting ER+|PR+HER2- breast cancer was 0.939, the AUC for HER2+ breast cancer was 0.925, and the AUC for TNBC breast cancer was 0.893.
[0106] In a specific embodiment, the present invention provides a method for predicting molecular subtypes of breast cancer, comprising the following steps: (1) Obtain the first data of the subject, which includes the clinical information data of the subject; at the same time, collect the peripheral blood sample of the subject, separate plasma from the peripheral blood sample and extract cfDNA, and perform high-throughput sequencing on the cfDNA to obtain the raw sequencing data; (1-a) The peripheral blood sample collection volume is 8–10 mL. After collection, it is stored in a dedicated cfDNA blood collection tube (e.g., Ardent BioMed, model BY10240301). The blood sample is subjected to double centrifugation at 1,600×g and 16,000×g to separate the plasma. It is then frozen at –80°C and transported to the central laboratory for unified processing using dry ice.
[0107] (1-b) cfDNA extraction was performed using a magnetic bead extraction kit (e.g., Vazyme, model N913). The extracted products were quantified using a Qubit analyzer and used for library construction. The library was constructed using an Illumina platform library construction kit and subjected to 150bp paired-end sequencing with an average sequencing depth of 2×.
[0108] (2) Perform quality control processing on the raw sequencing data, including removing low-quality reads, adapter contamination sequences and abnormal fragments; (3) Based on the alignment results, extract the following types of feature information from the data, including but not limited to: cfDNA fragment length distribution features, copy number variation (CNV) information, sequence pattern features such as base preference, coverage and enrichment features of cfDNA in different types of transposon element regions throughout the genome, etc. The cfDNA features include but are not limited to the following types: 1) Segment length distribution; 2) Fragment enrichment features of a fixed window; 3) Motif sequence features; 4) Base preference characteristics; 5) Coverage characteristics of the rotating element area; 6) Changes in the coverage of the CTCF binding site region; 7) Distribution characteristics of methylation motifs; 8) CNV features built based on WGS.
[0109] (4) Based on the LASSO regression method, the 9999-dimensional cfDNA features initially extracted were used to identify features closely related to breast cancer subtyping and retain a subset of key features for subsequent modeling.
[0110] (5) The model training and evaluation adopts a combination of bootstrapping and 10-fold cross-validation. In cross-validation, the training set is randomly divided into 10 subsets. In each round, 9 of them are used for training and the remaining one is used for testing. This process is repeated 10 times to ensure that each subset is tested once.
[0111] (6) The training process uses only samples from healthy controls and cancer patients. The model is trained based on features such as motif frequency, and a "predicted score" is generated to predict the probability of cancer. All validation data remains untouched during the training phase. The cancer score ranges from 0 to 1, with a higher score indicating a greater likelihood of having cancer. In a specific implementation, the model's judgment rules use a prediction score of 0.5 as the classification threshold: when the prediction score is greater than 0.5, the model judges the patient as positive; when the prediction score is less than or equal to 0.5, it is judged as negative.
[0112] (7) Construct multiple breast cancer subtyping prediction models based on the training set; (8) Train the above models on the training set respectively, and evaluate their performance on the training set and the test set; The model described in step (7) includes, but is not limited to, the following machine learning algorithms: Generalized Linear Model (GLM), Support Vector Machine (SVM), Random Forest (RF), Gradient Boosting Machine (GBM), Extreme Learning Machine (ELM), eXtreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM), Multilayer Perceptron (MLP), and Deep Neural Network (DNN), etc.
[0113] The evaluation metrics in step (8) include, but are not limited to: area under the curve (AUC), sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV).
[0114] The main advantages of this invention include: (a) This invention collects plasma cfDNA and integrates its mutation, fragment, methylation and other multi-omics features, and constructs a breast cancer subtyping prediction model by combining machine learning algorithms. This model can accurately identify three types of breast cancer: ER+ / PR+, HER2+ and TNBC, significantly improving subtyping efficiency and accuracy, and significantly enhancing the subtyping prediction accuracy and model robustness.
[0115] (b) The method of the present invention is non-invasive, allows for repeated sampling, facilitates dynamic monitoring, has high detection throughput and relatively low cost, and has the beneficial effects of being non-invasive, highly sensitive and clinically adaptable. It can be applied to patient groups that require dynamic monitoring or have difficulty obtaining tumor tissue; it is suitable for various application scenarios such as clinical screening, auxiliary diagnosis and telemedicine, and has strong practical value and promotion prospects.
[0116] (c) The method of this invention can effectively distinguish three molecular subtypes of breast cancer patients. Considering the significant differences between different subtypes in terms of survival prognosis, immune response, and treatment sensitivity, this method is expected to assist clinicians in developing personalized treatment strategies, improving efficacy, and reducing breast cancer mortality.
[0117] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Experimental methods in the following embodiments, unless otherwise specified, are generally performed under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or as recommended by the manufacturer. Unless otherwise stated, percentages and parts are weight percentages and parts by weight.
[0118] Example 1: Experimental method.
[0119] In this embodiment, the steps for constructing a non-invasive prediction model for breast cancer molecular subtyping are as follows: Figure 1 As shown in the left figure, the methodology and workflow during the research and development process are as follows: Figure 1 As shown in the middle right figure, specifically, it includes the following steps: (1) Recruitment of research subjects and sample collection In this embodiment, a total of 529 breast cancer patients were included in the study. All patients signed written informed consent forms. The study was reviewed and approved by the ethics committee and complies with the relevant requirements of the Declaration of Helsinki.
[0120] Of the participants, 317 had breast cancer classified as ER+ / PR+HER2-, 145 had HER2+, and 67 had TNBC. Exclusion criteria included: pregnancy or lactation, history of malignancy within the past five years, suspected but undiagnosed malignancy within the past year, history of blood transfusion within the past 30 days, active mastitis, or other medical conditions unsuitable for participation in this study.
[0121] During preoperative examinations or routine physical checkups, 8-10 mL of peripheral blood was collected from each subject using a dedicated cfDNA blood collection tube. Immediately after collection, the blood was centrifuged in two steps at 4°C: first, at 1,600×g for 10 minutes, and second, at 16,000×g for 10 minutes, to obtain cell-free plasma. The resulting plasma samples were frozen at -80°C and transported to the central laboratory using dry ice for subsequent analysis.
[0122] (2) cfDNA extraction, library construction and sequencing One mL plasma sample was collected, and cfDNA was extracted using the VAMNE MagUltra Cell-Free DNA Extraction Kit (model N913) provided by Vazyme. The concentration of the obtained cfDNA product was determined using a Qubit 4.0 real-time fluorescence analyzer (Thermo Fisher Scientific). The extracted cfDNA was stored at -80°C.
[0123] 5-20 ng cfDNA was used for library construction using the VAHTS Universal DNA Library Construction Kit (Illumina V3, model N610) and barcoded with the VAHTS Illumina Multiplex Primer Kit 5 (model N322). Library quality was assessed using an Agilent 2100 Bioanalyzer to ensure that the fragment size distribution was suitable for subsequent sequencing requirements.
[0124] Sequencing was performed using the Illumina NovaSeq X Plus platform, employing 150 bp paired-end sequencing with an average sequencing depth of 2×, which is classified as ultra-low-pass whole-genome sequencing (WGS).
[0125] (3) Quality control and comparative analysis of sequencing data The raw sequencing data were first quality-assessed using FastQC (v0.12.1), and then low-quality reads, adapter sequences, and contaminant fragments were removed using Cutadapt (v4.5) and Ktrim (v1.4.1). The cleaned reads were aligned to the human reference genome GRCh37 version using BWA-MEM (v0.7.17), retaining only the high-quality reads with unique alignments and PCR duplicates removed.
[0126] The alignment files were subsequently sorted and indexed using Samtools (v1.9), and Picard (v2.18.29) was used for repetitive sequence identification and removal. The cfDNA fragment length distribution was analyzed based on the alignment information. The reference genome was divided into non-overlapping 5 Mb windows for subsequent fragment distribution and CNV feature extraction analysis.
[0127] In another processing step, the sequencing data underwent further Trimmomatic (v0.39) processing, including adapter sequence splicing, removal of low-quality terminal bases (below Q20), removal of regions with an average quality value below Q15 within the sliding window (4 bp), and rejection of reads shorter than 50 bp. Finally, GATK (v4.4.0.0) was used to analyze the obtained alignment data, extracting SNV and indel information for specific target genes, and annotating them using ANNOVAR.
[0128] (4) cfDNA feature construction and normalization Based on the preprocessed data described above, various cfDNA-derived features were constructed, including but not limited to: 1) Fragment Length Feature: A feature based on the distribution of cfDNA fragment lengths was constructed to reflect the differences in apoptosis patterns between tumor and normal tissues. Specifically, sequencing fragments were divided into short fragments (S, 100-150 bp) and long fragments (L, 151-220 bp) according to length intervals. The distribution of each fragment across the entire genome was statistically analyzed, and their ratios were calculated. Simultaneously, to enhance the ability to capture local genomic structural differences, a 100 kb sliding window was introduced across the entire genome. The S / L ratio within each window was calculated, and all window data were standardized to form a fragment length ratio feature vector. This feature is stable, easy to obtain, and effectively reflects the shortening trend of tumor-derived cfDNA fragments, making it one of the important structural features in the modeling system of this invention.
[0129] 2) Motif Sequence Features: A 6-mer motif distribution feature based on the 3 bases before and after the 5' end of each cfDNA fragment (6 bp in total) was defined, aiming to serve as a fragmentomic biomarker for breast cancer subtyping. Specifically, the upstream and downstream 3 bases of each cfDNA fragment were extracted from the 5' end, and these 6-mer sequences were used as motif units. For each sample, the frequency of 6-mer motifs corresponding to all cfDNA fragments was counted, constructing a frequency matrix with the sample as the row and each possible 6-mer motif as the column, where each row represents a sample and each column represents a motif feature dimension. The matrix was then row-normalized: the frequency of all motifs in each row was divided by the total number of fragments in that row, ensuring that the sum of all motif frequencies in each row was always 1, thus eliminating the influence of the total amount of cfDNA in different samples. This feature characterizes the local sequence preference of cfDNA breakpoints and can be used as input into machine learning models for breast cancer subtyping, possessing the advantages of well-defined dimensions, unified normalization standards, and applicability to different sequencing depths and sample sizes.
[0130] 3) Transposon element region coverage characteristics (signal intensity including LTR repetitive sequence regions): Based on the enrichment intensity characteristics of cfDNA fragments in three major transposon element regions (SINE, LINE, LTR), this reflects the differences in fragment release driven by epigenetic abnormalities in breast cancer tissue. Based on annotation information of various TE (transposon element) families in the human genome, the position of each element in the whole genome is extracted. Whether the 5' end or midpoint of the sample cfDNA fragment falls into the target region is statistically analyzed, and normalized by the total number of fragments to construct fragmentomics feature values reflecting the enrichment intensity of SINE, LINE, and LTR regions. Studies have shown that the enrichment intensity of LINE and LTR regions in cfDNA from breast cancer patients is significantly higher than that in non-tumor individuals, suggesting that these regions may be preferred regions for cfDNA release from breast cancer tissue. Among them, SINEs (such as Alu) are short, scattered nuclear elements approximately 100-700 bp in length, distributed in functionally active regions of the genome such as promoters and regulatory regions; LINEs are autonomous retrotransposons, up to 6-7 kb in length, accounting for about 20% of the human genome, and are in a methylated silent state in normal tissues; LTRs are long terminal repeat sequences, originating from endogenous retroviral remnants, and their activation indicates chromatin derepressive state. In tumor cells, these transposon element regions undergo epigenetic reprogramming phenomena such as chromatin opening and demethylation, leading to abnormal enrichment or release of cfDNA fragments in these regions. This feature can be combined with other structural variation or sequence pattern fragmentomics indicators to significantly improve the accuracy of breast cancer subtype identification and risk stratification, and is suitable for tumor monitoring and early screening systems in liquid biopsy scenarios, providing a reliable indicator for non-invasive detection of tumor-specific epigenetic signals.
[0131] 4) Copy Number Variation (CNV) Feature Based on WGS: CNV features are an important means of characterizing tumor genomic structural abnormalities through copy number variation analysis of the entire genome. A sliding window strategy of 5 Mb units is used to scan the entire cfDNA genome sequence, and the coverage depth of cfDNA fragments within each window is statistically analyzed to construct a CNV map covering the entire genome. This feature can sensitively reflect common structural variations in tumor tissues, such as chromosomal amplification or deletion, revealing potential gene dosage changes and providing important evidence for breast cancer subtyping. By comparing CNV map features among different patients, copy number abnormality regions closely related to breast cancer subtypes can be further identified, enhancing the model's ability to identify tumor heterogeneity and improving the accuracy and clinical applicability of breast cancer subtyping.
[0132] 5) cfDNA Coverage Characteristics of CTCF Binding Site Regions: Based on the coverage characteristics of cfDNA fragments in the region surrounding CTCF binding sites, this feature identifies tumor-related changes in chromatin three-dimensional structure. CTCF, as a key regulator of chromatin topology, often binds to chromatin boundary regions with sparse nucleosomes and high DNA exposure, which are important protected sites during cfDNA formation. Specifically, high-confidence CTCF binding sites from the ReMap database, validated by high-throughput ChIP-seq experiments, are selected. A symmetrical analysis window is constructed within ±1000 bp of the center of each site, and this region is divided into multiple fixed-size bins (10 bp). The coverage density of cfDNA fragments is accurately calculated, generating a meta-coverage map as an important indicator reflecting local nucleosome arrangement and chromatin accessibility. Studies have shown that normal cells exhibit highly conserved and regular cfDNA distribution patterns in CTCF binding regions, while tumor cells often show significant changes in cfDNA coverage in these regions due to chromatin remodeling, loss of CTCF binding, or abnormal topology. This feature can stably and quantitatively capture these changes, exhibiting good tissue specificity and state sensitivity. In the fragment omics feature set and classification model constructed in this invention, CTCF region coverage is included as one of the core discriminant features, which helps to significantly improve the discriminant ability of tumor-derived cfDNA and the generalization performance of typing prediction.
[0133] 6) Fixed-window fragment enrichment features (including fragmentation ratio): Fixed-window fragment enrichment features involve dividing the entire genome into fixed-size windows (10 bp) and statistically analyzing the coverage or count density of cfDNA fragments in each window, thereby reflecting the distribution pattern of cfDNA in different regions of the genome. This feature can be used to identify enrichment or deletion signals in tumor-related genomic regions and capture the spatial distribution differences of cfDNA fragments released by tumor cells at the chromosomal level. The enriched signals, after standardization, can be used to further construct discriminative models for breast cancer subtyping, improving the model's sensitivity to implicit genomic features in tumor subtyping.
[0134] 7) Methylation Motif Distribution Characteristics: An indirect epigenetic feature based on the distribution of cfDNA methylation motifs is proposed to reflect the tumor tissue-specific chromatin state and methylation abnormalities. In the human genome, certain specific nucleotide sequences (motifs), such as CGCG, CCGG, and GCGC, are highly correlated with DNA methylation and are often enriched in promoters, CpG islands, and epigenetic regulatory regions. By statistically analyzing the positional relationship between the 5' end of cfDNA fragments and these methylation-sensitive motifs, sequences before and after the breakpoints are extracted, and 6-mer or other length motif distribution frequency matrices are constructed. By comparing the differences in breakpoint density near these motifs (within 3 bp) in different samples, the methylation level and chromatin accessibility changes in the corresponding regions of the tissue from which the cfDNA originated are indirectly reflected. The study found that breakpoints near motifs in hypomethylated regions of tumor samples are significantly enriched, exhibiting a unique fragment distribution pattern, suggesting tumor-related epigenetic reprogramming. This feature does not rely on traditional methylation sequencing methods and has the advantages of being non-invasive, having high resolution, and having stable signals. It can be used as an important indicator of epigenetic status in cfDNA fragmentomics analysis and can be combined with other structural or regional features for precise subtype identification and risk stratification of various tumors such as breast cancer.
[0135] 8) Base Preference Feature: A fragmentomics feature based on base preference at cfDNA break sites is proposed to characterize the sequence selectivity differences between breast cancer patients and non-tumor individuals during cfDNA fragment cleavage or degradation. This feature analyzes the composition of nucleotides before and after the 5' end of the cfDNA fragment, extracting the proportion of single bases and the frequency of dibase combinations at the break site and its adjacent region (6-mer), constructing a feature vector reflecting base preference. These ratios reflect the preference of cfDNA breaks for different sequence environments, indirectly reflecting the chromatin state, cleavage mechanism, and nucleosome localization patterns of the tissue from which the fragment originates. Studies have found that cfDNA fragments released from tumor tissues often exhibit specificity in break sequence preference, such as a high proportion of T / C preference or a low frequency of C / G combinations, suggesting that they differ from normal cell apoptosis pathways or chromatin structures. This base preference feature does not rely on methylation or chromatin data, has good stability, is easy to obtain, and is applicable to standard cfDNA WGS data, showing broad diagnostic and discriminative potential in various applications such as breast cancer subtyping, tumor detection, and tissue tracing.
[0136] 9) Blacklist area filtering: exclude low comparison areas in ENCODE Blacklist and UCSC gap track to reduce noise.
[0137] (5) Model training and cross-validation After feature extraction, LASSO regression was first used for feature selection to obtain a subset of key cfDNA features ranging from 200 to 9999 dimensions. To reduce the risk of model overfitting, 10-fold cross-validation combined with bootstrapping was used for model training and performance evaluation during the training process.
[0138] The models include: Generalized Linear Model (GLM), Support Vector Machine (SVM), Random Forest Classifier (RFC), Gradient Boosting Machine (GBM), Extreme Learning Machine (ELM), eXtreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM), Multilayer Perceptron (MLP), and Deep Neural Network (DNN), etc.
[0139] In the 10-fold cross-validation process, the training set was randomly divided into 10 subsets, with 9 subsets used for training and the remaining subset used for testing. This process was repeated 10 times, ensuring that each subset was used as the test set once. Only training cohort samples from healthy individuals and cancer patients were used. A classifier was trained using features such as motif frequencies in a machine learning algorithm to generate a model predicting the cancer score for each sample. Notably, all validation datasets were not used during model training. Cancer scores ranged from 0 to 1, with higher scores indicating a higher probability of having cancer. After evaluation, the best-performing model was selected for downstream analysis.
[0140] These models were then applied to a validation dataset to generate cancer prediction scores for each validation sample and evaluate model performance. For evaluation, the AUC values of different models in the validation cohort were compared, as well as the sensitivity / specificity at a fixed specificity threshold in the internal validation cohort.
[0141] (5) Model evaluation and breast cancer subtype discrimination After training, the model is evaluated on both the test set and the independent validation set. Evaluation metrics include: AUC (area under the curve), sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV).
[0142] (6) Model comparison and final model determination After model training was completed, the performance of various candidate models was compared. The comparison process was based on the same dataset, and the discrimination effect of each model was evaluated in three molecular subtypes of breast cancer (ER+ / PR+HER2-, HER2+, and TNBC). Through a comprehensive evaluation of the discrimination performance of each model, the generalized linear model was finally selected as the preferred embodiment of the present invention, and the final discrimination model built based on this model was named the TuFEst-MS model.
[0143] Experimental Example 2: Breast Cancer Subtyping Prediction Based on Different Machine Learning Models
[0144] Fragment length distribution features, motif sequence features, methylation motif distribution features, base preference features, transposon element region coverage features, nucleosome localization features, copy number variation (CNV) features, and fixed-window fragment enrichment features were extracted. Different machine learning algorithms—GLM, RFC, GBM, LightGBM, and DNN—were used to construct models, and the results are as follows: Figures 2-7 As shown.
[0145] Figure 2 and Figure 3 The results are shown on the training and validation sets of various models for the ER+|PR+HER2-fractal. It can be seen that the GLM model has the highest AUC value, all greater than 0.9, and the AUC values of the training and validation sets are very close. This indicates that the model performs well on the ER+|PR+HER2-fractal training set and successfully "generalizes" this performance to the ER+|PR+HER2-fractal validation set.
[0146] Figure 4 and Figure 5 The results are for the training and validation sets of various HER2+ subtyping models. It can be seen that the GLM model has the highest AUC value, all greater than 0.9, and the AUC values of the training and validation sets are similar, indicating that the model performs well on the HER2+ subtyping training set and successfully "generalizes" this performance to the HER2+ subtyping validation set.
[0147] Figure 6 and Figure 7 The results are for the training and validation sets of various TNBC classification models. It can be seen that the GLM model has the highest AUC value, all greater than 0.89, and the AUC values of the training and validation sets are very close. This indicates that the model performs well on the TNBC classification training set and successfully "generalizes" this performance to the TNBC classification validation set.
[0148] Therefore, all of these models can be used to distinguish between the three types of breast cancer, with GLM showing the best results.
[0149] Experimental Example 3: Feature Contribution Analysis of Different cfDNA Features Extracted
[0150] The training results and performance on the validation set of the GLM model based on the signal intensity of the LTR repeat sequence region are as follows: Figure 8 and Figure 9 As shown in the figure, the AUC values of the training and validation sets for the ER+|PR+, HER2+, and TNBC subtypes are all above 0.78, and the AUC values of the training and validation sets are very close. This indicates that the model performs well on the TNBC subtype training set and successfully "generalizes" this performance to the TNBC subtype validation set.
[0151] The training set prediction results and validation set prediction results of the GLM model based on fragment length distribution, motif sequence, and nucleosome localization features are as follows: Figure 10 and Figure 11 As shown, the AUC values of the ER+|PR+, HER2+, and TNBC subtypes on both the training and validation sets are above 0.81. This is essentially equivalent to the performance of the GLM model based solely on the signal intensity of the LTR repeat sequence region.
[0152] The training set prediction results and validation set prediction results of the GLM model based on the aforementioned eight cfDNA features are as follows: Figure 12 and Figure 13 As shown, the AUC values of the training and validation sets for the ER+|PR+, HER2+, and TNBC subtypes are all above 0.89.
[0153] further, Figure 14 , Figure 15 , Figure 16 The results showed that the ER+|PR+ subtype, HER2+ subtype, and TNBC subtype of breast cancer were highly distinguishable. Based on validation results from real samples ( Figure 17 , Figure 18 , Figure 19 This further confirms that the model of the present invention has a high degree of differentiation between different subtypes of breast cancer, indicating that the model can be used for the subtyping and accurate diagnosis of breast cancer.
[0154] Experiment 4: Detection of molecular subtyping prediction in 21 patients with advanced breast cancer.
[0155] (1) Information processing of late-stage patient samples Twenty-one patients with advanced breast cancer from multiple clinical centers were selected. Among them, 7 patients had ER+ / PR+ breast cancer, 3 patients had HER2+ breast cancer, and 11 patients had TNBC breast cancer. All samples were advanced breast cancer cases, clinically stage IV. Diagnosis was based on tissue biopsy and imaging examinations. Some patients had distant metastases (such as bone metastases and liver metastases). Basic patient characteristics included gender, age, ID number, enrollment number, primary site, clinical stage (such as T, N, M status), and clinical immunohistochemical typing (ER / PR / HER2 status).
[0156] (2) External validation design of the model To evaluate the model's generalization ability across different datasets, the aforementioned 21 patients were used as an independent external validation set, not participating in model training or cross-validation. For each patient's cfDNA data, a consistent feature set was extracted according to the method described in Example 1, and input into the trained TuFEst-MS model to predict its molecular subtype (ER+ / PR+HER2-, HER2+, and TNBC).
[0157] (3) Comparison of prediction results with clinical labels The model's predicted subtyping results were compared one by one with the actual clinical subtyping labels. The model accurately predicted the corresponding subtyping in most samples, validating its applicability and stability in advanced breast cancer samples.
[0158] (4) Visualization: Expression of Sankey diagram results To visually demonstrate the consistency between model predictions and actual subtyping, as well as the sample distribution structure, a Sankey diagram is used for visualization.
[0159] like Figure 20 As shown, the left-hand nodes represent the patient's actual clinical subtype label, and the right-hand nodes represent the model's predicted subtype result. The line connecting the two represents the matching status for each patient. The line width represents the number of samples in the corresponding subtype category. It can be clearly observed that most samples are consistent with the actual label in the model's prediction, demonstrating the model's accuracy and clinical translation potential.
[0160] Twenty-one patients with advanced breast cancer diagnosed by tissue biopsy were included in the evaluation. The model predicted the molecular subtype of their metastatic lesions with an overall accuracy of 85.7%.
[0161] In eight patients whose primary and metastatic lesion classifications were inconsistent, the classification model still successfully identified the true classifications of seven metastatic lesions, achieving an accuracy rate of 87.5%. This demonstrates that cfDNA fragmentomic features can effectively capture clonal evolution or molecular expression differences during breast cancer progression. These results suggest that, compared to traditional classification methods relying on tissue biopsy, non-invasive detection methods based on cfDNA have stronger global reflective capabilities, making them particularly suitable for patients in whom obtaining metastatic tissue is difficult or inconvenient, providing crucial reference for precise treatment planning.
[0162] The discrepancies in predictions in the three patients were all related to the determination of HER2 expression status. Specifically, one patient with a metastatic lesion confirmed as HER2-positive (IHC 3+) by immunohistochemistry was predicted as HER2-negative by the model; the other two patients with HER2 IHC 2+ and FISH-negative (defined as HER2-negative or low-expression according to the ASCO / CAP guidelines) were predicted as HER2-positive by the model. This phenomenon may be influenced by several factors: First, HER2 expression exhibits significant spatial heterogeneity; expression levels may vary considerably between different metastatic sites and even between different regions within the same lesion, while tissue biopsy can only reflect a single local state. Second, HER2 expression has a dynamic change over time; with disease progression and the cumulative effects of treatment interventions, HER2 status may change, and cfDNA, as a comprehensive signal released from multiple sites, may reflect the changing trends in clonal succession earlier. Third, the cfDNA signal essentially reflects the overall burden released by tumor cells throughout the body; when multiple subclones coexist, liquid biopsy may detect a dominant signal different from the histopathological results. Furthermore, the current HER2 pathological assessment criteria still have some controversy regarding the definition of the HER2 low-expression subtype that is IHC 2+ and FISH negative, and the consistency between different institutions still needs to be improved. cfDNA fragmentomics may show higher sensitivity in this type of "blurred" subtype.
[0163] In summary, these results not only confirm the effectiveness of the TuFEst-MS model in predicting metastatic lesion classification, but also reveal the unique advantages of cfDNA fragment omics technology in reflecting tumor heterogeneity and dynamic changes. In the future, it is expected to serve as an important supplementary tool to tissue biopsy, especially in identifying HER2 low expression status and making personalized treatment decisions.
[0164] All documents mentioned in this invention are incorporated herein by reference as if each document were individually incorporated by reference. Furthermore, it should be understood that after reading the foregoing teachings of this invention, those skilled in the art can make various alterations or modifications to this invention, and these equivalent forms also fall within the scope defined by the appended claims.
Claims
1. A method for constructing a non-invasive predictive model for molecular subtyping of breast cancer, characterized in that, Includes the following steps: (S1) Provides sequencing results of cfDNA in the sample; (S2) The sequencing results are preprocessed and quality controlled to obtain sequencing data that meets the quality control conditions; (S3) Extract cfDNA features from sequencing data that meet quality control conditions, wherein the cfDNA features include: signal intensity of LTR repeat sequence regions; (S4) Construct a machine learning model based on the cfDNA features to obtain a non-invasive prediction model for molecular subtyping of breast cancer; The signal intensity of the LTR repeat sequence region refers to the ratio of the number of cfDNA fragments whose 5' end or midpoint falls into the LTR repeat sequence region to the total number of cfDNA fragments.
2. The method as described in claim 1, characterized in that, The molecular subtypes of breast cancer include: ER+ / PR+HER2-, HER2+, and TNBC.
3. The method as described in claim 1, characterized in that, The cfDNA features also include one or more cfDNA features selected from the group consisting of: fixed-window fragment enrichment features, base preference features, transposon element region coverage features, cfDNA fragment length distribution, motif sequence features, nucleosome localization features, copy number variation (CNV) features, and methylation motif distribution features.
4. The method as described in claim 3, characterized in that, The cfDNA fragment length distribution refers to the ratio S / L of the number of short-fragment cfDNA and the number of long-fragment cfDNA.
5. The method as described in claim 3, characterized in that, The motif sequence characteristics refer to the frequency of the 6-mer motif corresponding to each cfDNA fragment in the sample; The 6-mer motif comprises three bases upstream and three bases downstream of the 5' end of the cfDNA fragment.
6. The method as described in claim 3, characterized in that, The base preference feature includes the single base frequencies upstream and downstream of the 5' end of the cfDNA fragment; and / or Frequency of dibase combinations upstream and downstream of the 5' end of the cfDNA fragment.
7. The method as described in claim 1, characterized in that, The machine learning models include: Generalized Linear Model (GLM), Deep Neural Network Model (DNN), or Lightweight Gradient Boosting Machine Model (LightGBM).
8. A non-invasive predictive model for molecular subtyping of breast cancer, characterized in that, The non-invasive prediction model is constructed using the construction method described in claim 1.
9. The use of the non-invasive prediction model according to claim 8, characterized in that, This is used to prepare a non-invasive prediction system for predicting molecular subtypes of breast cancer.
10. A non-invasive predictive system for molecular subtyping of breast cancer, characterized in that, The system includes the following modules: (Z1) Data input module, which is configured to input the sequencing results of cfDNA in the sample to be analyzed; (Z2) Preprocessing and quality control module, wherein the preprocessing and quality control module is configured to preprocess and control the sequencing results to obtain sequencing data that meets the quality control conditions; (Z3) Feature extraction and selection module, the feature extraction and selection module is configured to: extract cfDNA features from sequencing data that meet quality control conditions, wherein the cfDNA features include: signal intensity of LTR repeat sequence regions; (Z4) Breast cancer molecular typing evaluation module, wherein the breast cancer molecular typing evaluation module is configured to: predict the breast cancer molecular typing result of the sample to be analyzed based on the cfDNA characteristics using the non-invasive prediction model of breast cancer molecular typing as described in claim 8; (Z5) Output module, which is configured to output the molecular subtyping results of the breast cancer of the sample to be analyzed.
Citation Information
Patent Citations
Non-invasive cancer early screening system based on cfDNA omics characteristics
CN113160889A
Breast cancer molecular typing method, device and system based on unsupervised learning
CN113643269A
Cell-free DNA methylation test for breast cancer
WO2024112946A1
Classification of breast tumors using DNA methylation from liquid biopsy
WO2025019254A1