Breast cancer lymph node metastasis prediction method and application thereof

By extracting the signal intensity of LTR repeat sequence regions from plasma cfDNA and using machine learning models, a method for predicting lymph node metastasis in breast cancer was constructed. This method solves the problems of high invasiveness or low accuracy in existing technologies, and achieves high-precision non-invasive prediction and personalized treatment guidance.

CN121506510APending Publication Date: 2026-02-10OMIXSCIENCE(HANGZHOU) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610043910.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-14
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Current technologies for assessing lymph node metastasis in breast cancer rely on imaging examinations and histopathology, which are either highly invasive or have low accuracy, making it difficult to achieve non-invasive and accurate prediction.

Method used

By extracting the signal intensity of LTR repeat sequence regions from plasma cfDNA as features, a method for predicting lymph node metastasis in breast cancer was constructed by combining it with a machine learning model. This method includes sample sequencing, preprocessing, quality control, and feature extraction, and a non-invasive prediction model was built.

Benefits of technology

It achieves high-precision non-invasive prediction of lymph node metastasis in breast cancer, with an AUC value of over 0.87, which is significantly better than traditional methods and is suitable for dynamic monitoring and individualized treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506510A_ABST
    Figure CN121506510A_ABST
Patent Text Reader

Abstract

The invention relates to the field of tumor molecular diagnosis, in particular to a breast cancer lymph node metastasis prediction method and application thereof. Specifically, the invention provides a construction method of a breast cancer lymph node metastasis prediction model by taking the signal intensity of an LTR repetitive sequence region as one of extracted cfDNA features, and provides an application of the model in prediction of the breast cancer lymph node metastasis condition, so that non-invasive detection of breast cancer lymph node metastasis is realized, and the prediction accuracy of breast cancer lymph node metastasis is improved. The method is especially suitable for patient groups needing dynamic monitoring or difficult to obtain tumor tissues.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of tumor molecular diagnosis, in particular to a breast cancer lymph node metastasis prediction method and application thereof. BACKGROUND

[0002] Breast cancer is the most common malignant tumor in women worldwide and one of the leading causes of cancer death. Breast cancer is highly heterogeneous, and its clinical outcome depends not only on molecular subtypes but also on tumor metastatic behavior. Lymph node metastasis is one of the most important metastatic modes of breast cancer and is a key indicator affecting patient staging, prognosis, and treatment decisions. Generally, the status of axillary lymph node metastasis is directly related to the choice of surgical approach, the development of chemotherapy and radiotherapy regimens, and the overall survival expectation of patients.

[0003] Currently, the assessment of breast cancer lymph node metastasis mainly relies on imaging examination and histopathological examination. However, imaging examination has limited sensitivity and specificity, which may lead to missed or misdiagnosed cases; while histopathological examination, although highly accurate, is invasive, increasing the risk of trauma and postoperative complications for patients. Therefore, how to achieve non-invasive and accurate prediction of breast cancer lymph node metastasis is an important technical challenge in clinical diagnosis and treatment.

[0004] In recent years, circulating free DNA (cfDNA), as a kind of free nucleic acid molecule derived from tumor cell apoptosis or necrosis and released into peripheral blood, has been widely used in early detection, recurrence monitoring, and efficacy evaluation of tumors due to its non-invasiveness, high repeatability, and ability to reflect tumor dynamic changes. Studies have shown that the fragmentomic characteristics, methylation patterns, and mutation information of specific sites of cfDNA may be closely related to tumor metastatic behavior. However, due to the low concentration of cfDNA in blood and the fact that tumor-derived signals are often masked by normal cell background, how to effectively extract molecular features related to lymph node metastasis and construct an efficient prediction model remains a technical challenge that needs to be addressed. SUMMARY

[0005] The present application provides a breast cancer lymph node metastasis prediction method based on cfDNA feature extraction and machine learning model, which solves the problem of high invasiveness or low prediction accuracy in existing methods.

[0006] In the first aspect of the present application, a method for constructing a non-invasive prediction model of breast cancer lymph node metastasis is provided, comprising the following steps: (S1) providing sequencing results of cfDNA in the sample; (S2) preprocessing and quality control of the sequencing results to obtain sequencing data meeting the quality control conditions; (S3) extracting cfDNA features from the sequencing data satisfying the quality control condition, wherein the cfDNA features comprise signal intensity of LTR repeat sequence region; (S4) constructing a machine learning model based on the cfDNA features, thereby obtaining a non-invasive prediction model of breast cancer lymph node metastasis; In some embodiments, the signal intensity of the LTR repeat sequence region refers to the ratio of the number of cfDNA fragments with 5' end or midpoint falling into the LTR repeat sequence region to the total number of cfDNA fragments in the sample cfDNA.

[0007] In another preferred embodiment, the sample comprises a breast cancer sample.

[0008] In another preferred embodiment, the sample comprises a breast cancer sample that has occurred lymph node metastasis.

[0009] In another preferred embodiment, the sample comprises a breast cancer sample that has not occurred lymph node metastasis.

[0010] In another preferred embodiment, the sample comprises a plasma sample, a blood sample, a urine sample, a fecal sample, a tear sample, or a combination thereof.

[0011] In another preferred embodiment, the sample is a plasma sample.

[0012] In another preferred embodiment, step (S1) comprises the following sub-steps: (S1-1) providing a sample; (S1-2) extracting cfDNA from the sample; (S1-3) sequencing the cfDNA, thereby obtaining sequencing results of cfDNA in the sample.

[0013] In another preferred embodiment, the extraction comprises using magnetic bead method or using cfDNA extraction special kit.

[0014] In another preferred embodiment, the sequencing is ultra-low coverage whole genome sequencing.

[0015] In another preferred embodiment, the sequencing adopts 150 bp paired-end sequencing mode.

[0016] In another preferred embodiment, the average sequencing depth of the sequencing is 2x.

[0017] In another preferred embodiment, the sequencing is performed by Illumina NovaSeq X Plus platform.

[0018] In another preferred embodiment, in step (S1), the sequencing results are WGS raw data.

[0019] In another preferred embodiment, the preprocessing and quality control comprises removing low quality reads, removing adapter contamination sequences, or a combination thereof.

[0020] In another preferred embodiment, the quality control comprises quality control on sequencing depth, coverage, GC content bias.

[0021] In another preferred embodiment, the preprocessing and quality control comprises removing low quality reads, removing adapter sequences, aligning to human reference genome, sorting, indexing, removing PCR duplicates, or a combination thereof.

[0022] In another preferred embodiment, the human reference genome is human reference genome GRCh37 version.

[0023] In another preferred embodiment, the reference genome is divided into 5 Mb size or 100 kb non-overlapping window regions.

[0024] In another preferred embodiment, preprocessing and quality control comprises quality control and alignment processing of the sequencing results using one or more tools selected from the group consisting of FastQC, Cutadapt, Trimmomatic, BWA-MEM, Samtools, Picard tools, or a combination thereof.

[0025] In another preferred embodiment, the breast cancer lymph node metastasis status comprises metastasis has occurred, metastasis has not occurred.

[0026] In another preferred embodiment, the cfDNA features further comprise one or more cfDNA features selected from the group consisting of fixed window fragment enrichment features, base bias features, transposable element region coverage signal features, cfDNA fragment length distribution features, motif features of cfDNA break site sequences, nucleosome positioning features, copy number variation (CNV) features.

[0027] In another preferred embodiment, the transposable element region comprises SINE transposable element region, LINE transposable element region, or a combination thereof.

[0028] In another preferred embodiment, the LTR repeat sequence region comprises endogenous retrovirus family sequence region.

[0029] In another preferred embodiment, the endogenous retrovirus family sequence region comprises ERV1, ERVK, ERVL, ERVL-MaLR, or a combination thereof.

[0030] In another preferred embodiment, the SINE transposable element region comprises Alu element, MIR element, tRNA-derived SINE sequence region, or a combination thereof.

[0031] In another preferred embodiment, the LINE transpose element region includes a long, dispersed core element sequence region.

[0032] In another preferred embodiment, the long, dispersed nuclear element sequence region includes: L1 element, L2 element, CR1, or a combination thereof.

[0033] In another preferred embodiment, the cfDNA fragment length distribution characteristic refers to the ratio S / L of the number of short fragment length cfDNA and the number of long fragment length cfDNA; The short fragment length cfDNA is cfDNA with a fragment length of 100-150 bp, and the long fragment length cfDNA is cfDNA with a fragment length of 151-220 bp.

[0034] In another preferred embodiment, the cfDNA fragment length distribution characteristics exclude the ENCODE blacklist region and the UCSCgap track region.

[0035] In another preferred embodiment, the motif feature of the cfDNA break site sequence refers to the frequency of the 6-mer motif corresponding to each cfDNA fragment in the sample; The 6-mer motif comprises three bases upstream and three bases downstream of the 5' end of the cfDNA fragment.

[0036] In another preferred embodiment, the copy number variation (CNV) feature refers to a CNV pattern covering the entire cfDNA genome.

[0037] In another preferred embodiment, the CNV map is a CNV map covering the entire genome constructed using a sliding window of 5 Mb units.

[0038] In another preferred embodiment, the nucleosome localization feature refers to the coverage density of cfDNA fragments within the CTCF binding region.

[0039] In another preferred embodiment, the fixed-size bin has a length of 10 bp.

[0040] In another preferred embodiment, the fragment enrichment feature of the fixed window refers to the coverage density or count density of cfDNA fragments in each fixed-size window.

[0041] In another preferred embodiment, the fixed-size window has a length of 10 bp.

[0042] In another preferred embodiment, the base preference feature includes: single base frequencies upstream and downstream of the 5' end of the cfDNA fragment; and / or Frequency of dibase combinations upstream and downstream of the 5' end of the cfDNA fragment.

[0043] In another preferred embodiment, the machine learning model includes: a generalized linear model (GLM), a support vector machine (SVM), a multilayer perceptron (MLP), and a deep neural network (DNN).

[0044] In another preferred embodiment, the machine learning model is a generalized linear model (GLM). In another preferred embodiment, step (S4) further includes the following sub-steps: (S4-1) The cfDNA features are normalized and standardized to obtain standardized feature data; (S4-2) A machine learning model is constructed based on the standardized feature data. The model is trained and its performance is evaluated, thereby providing a non-invasive prediction model for lymph node metastasis in breast cancer.

[0045] In another preferred embodiment, 10-fold cross-validation is used for model training and performance evaluation.

[0046] In another preferred embodiment, the performance evaluation metrics include: AUC (area under the curve), sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), or a combination thereof.

[0047] In a second aspect of the invention, a non-invasive predictive model for lymph node metastasis in breast cancer is provided, the non-invasive predictive model being constructed using the construction method described in the first aspect of the invention.

[0048] In a third aspect of the invention, the use of the non-invasive prediction model described in the second aspect of the invention is provided for preparing a non-invasive prediction system for predicting lymph node metastasis in breast cancer.

[0049] In a fourth aspect of the invention, a non-invasive prediction system for lymph node metastasis in breast cancer is provided, the system comprising the following modules: (Z1) Data input module, which is configured to input the sequencing results of cfDNA in the sample to be analyzed; (Z2) Preprocessing and quality control module, wherein the preprocessing and quality control module is configured to preprocess and control the sequencing results to obtain sequencing data that meets the quality control conditions; (Z3) Feature extraction and selection module, the feature extraction and selection module is configured to: extract cfDNA features from sequencing data that meet quality control conditions, wherein the cfDNA features include: signal intensity of LTR repeat sequence regions; (Z4) Breast cancer lymph node metastasis evaluation module, wherein the breast cancer lymph node metastasis evaluation module is configured to: predict the breast cancer lymph node metastasis status of the sample to be analyzed based on the cfDNA characteristics using the non-invasive prediction model described in the second aspect of the present invention; (Z5) Output module, which is configured to output the lymph node metastasis status of the sample to be analyzed.

[0050] In another preferred embodiment, the signal intensity of the LTR repeat sequence region refers to the ratio of the number of cfDNA fragments whose 5' end or midpoint falls into the LTR repeat sequence region to the total number of cfDNA fragments.

[0051] In another preferred embodiment, the cfDNA feature further includes one or more cfDNA features selected from the group consisting of: fixed window fragment enrichment features, base preference features, transposon element region coverage signal features, cfDNA fragment length distribution features, cfDNA break site sequence motif features, nucleosome localization features, and copy number variation (CNV) features.

[0052] In another preferred embodiment, module (Z1) comprises the following sub-modules: (Z1-1) Sample providing module, the sample providing module being configured to: provide samples; (Z1-2) cfDNA extraction module, wherein the cfDNA extraction module is configured to extract cfDNA from the sample; (Z1-3) Sequencing module, the sequencing module being configured to sequence the cfDNA to obtain the sequencing results of cfDNA in each type of breast cancer lymph node metastasis sample.

[0053] In another preferred embodiment, the sample is a breast cancer sample.

[0054] In another preferred embodiment, the sample is a plasma sample.

[0055] In another preferred embodiment, the status of breast cancer lymph node metastasis includes: metastasis has occurred, or no metastasis has occurred.

[0056] It should be understood that, within the scope of this invention, the above-described technical features of this invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be described in detail here. Attached Figure Description

[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0058] Figure 1 This document illustrates a flowchart of a method for predicting lymph node metastasis in breast cancer according to an embodiment of this application. The left figure shows the framework of steps for constructing a non-invasive prediction model for breast cancer lymph node metastasis, while the right figure shows a detailed flowchart of the research and development process.

[0059] Figure 2 The graph shows a comparison of ROC curves for different model training sets.

[0060] Figure 3 The graph shows a comparison of ROC curves for different model validation sets.

[0061] Figure 4 The training set results of the GLM model based on the signal intensity of the LTR repeat sequence region are shown.

[0062] Figure 5 The performance of the GLM model based on the signal intensity of the LTR repeat sequence region is shown on the validation set.

[0063] Figure 6 The ROC curve of the TuFEst-LNM model training dataset is shown.

[0064] Figure 7 The confusion matrix between the predictions and actual results on the TuFEst-LNM model training dataset is shown.

[0065] Figure 8 The ROC curve of the TuFEst-LNM model validation set is shown.

[0066] Figure 9 The confusion matrix between the predictions and actual results of the TuFEst-LNM model validation set is shown.

[0067] Figure 10 The ROC curve of the external independent validation set of the TuFEst-LNM model is shown.

[0068] Figure 11 The confusion matrix between the predictions and actual results of the TuFEst-LNM model on the external independent validation set is shown.

[0069] Figure 12This demonstrates the matching between TuFEst-LNM predictions and actual results from patients in an external independent validation set. Detailed Implementation

[0070] Through extensive and in-depth research, the inventors unexpectedly discovered a high-precision model for predicting lymph node metastasis in breast cancer patients. Specifically, by using the signal intensity of the LTR repeat sequence region alone as an extracted cfDNA feature to construct a machine learning model, a rapid, efficient, and non-invasive prediction of whether lymph node metastasis has occurred in breast cancer was achieved, with AUC values ​​consistently exceeding 0.83. Furthermore, after optimizing the machine learning model (using the GLM model) and adding extracted cfDNA features, the AUC values ​​consistently exceeded 0.87. This invention was completed based on these findings.

[0071] Specifically, the inventors combined multi-dimensional molecular omics features to construct a machine learning model capable of highly accurate prediction of lymph node metastasis status in breast cancer patients. Even in patients where there are clinical discrepancies between the primary lesion and metastatic status, this model can still accurately identify the true lymph node metastasis, with a prediction accuracy significantly superior to traditional imaging and single molecular marker detection methods. Therefore, the technology proposed in this invention holds promise as an important supplementary tool to pathological diagnosis and tissue biopsy, playing a crucial role, especially in the early detection of potential metastatic risks, guiding surgical plan selection, and developing individualized adjuvant therapy strategies.

[0072] This invention discloses a method for predicting lymph node metastasis in breast cancer based on multi-feature cfDNA molecular omics and its application. This method uses plasma cfDNA as the analysis object, integrating cfDNA fragmentomics and molecular omics features such as fragmentation ratio, fragment length distribution, nucleosome localization, repetitive sequence signal intensity, and breakpoint motif frequency. A supervised machine learning algorithm is used to construct a predictive model, achieving accurate identification of lymph node metastasis risk in breast cancer patients. This invention achieves non-invasive detection and is suitable for patient groups requiring dynamic monitoring or those where histological sampling is difficult.

[0073] By integrating different types of cfDNA features, this method characterizes the molecular features associated with lymph node metastasis in breast cancer, significantly improving prediction accuracy and model robustness. This method can effectively distinguish between patients at risk of lymph node metastasis and those without. Considering that lymph node metastasis is a significant risk factor for breast cancer progression, recurrence, and distant metastasis, this method holds promise for assisting clinicians in stratified management, optimizing treatment decisions, and improving the overall survival and prognosis of breast cancer patients.

[0074] the term To facilitate a clearer understanding of this disclosure, certain terms are first defined. As used herein, unless otherwise expressly specified herein, each of the following terms shall have the meaning given below. Other definitions are set forth throughout the application.

[0075] As used herein, the term “and / or” refers to and covers any and all possible combinations of one or more of the related listed items.

[0076] As used herein, the terms “comprising,” “including,” and “containing” are used interchangeably and include not only closed definitions but also semi-closed and open definitions. In other words, the terms include “consisting of” and “substantially consisting of”.

[0077] As used in this article, the term "coverage density" refers to the density of sequencing reads within a specific region of the genome, usually measured by the average sequencing depth per unit length.

[0078] As used in this article, the term "count density" refers to the number of sequencing reads mapped per unit length of genome or per unit region, used to measure the level of gene or transcript expression in a specific region.

[0079] As used in this article, the terms "motif" and "motif" are used interchangeably, both referring to short segments in biological sequences (DNA, RNA, or proteins) that are conserved, have specific functions or structural significance, and are the core functional units or characteristic markers of biomolecules.

[0080] As used in this article, LTR (Long Terminal Repeat) is a non-coding repetitive sequence at both ends of a retroviral genome. It contains regulatory elements such as promoters and enhancers, but does not participate in protein coding. Its structure consists of a 5' end and a 3' end LTR. LTRs are long terminal repeat sequences derived from endogenous retroviral remnants, and their activation indicates chromatin derepression.

[0081] In a specific embodiment, the present invention provides a method for predicting lymph node metastasis in breast cancer based on multi-feature cfDNA molecular omics and its application, the method comprising the following steps: (1) Collect plasma samples from the subjects to be tested, extract cfDNA, and perform high-throughput sequencing to obtain sequencing data; (2) The sequencing data is preprocessed and feature extracted to obtain multi-dimensional cfDNA molecular features related to lymph node metastasis in breast cancer, including but not limited to: cfDNA fragment length distribution characteristics; Fragment enrichment features of a fixed window; Motif characteristics of cfDNA breakpoint sequences; Base preference characteristics; The transposable element regions (SINE transposable element region, LINE transposable element region, LTR repeat sequence region) cover signal characteristics; Nucleosome localization characteristics; Copy number variation (CNV) characteristics; (3) The extracted cfDNA features were normalized and standardized, and dimensionality was reduced by feature selection method to obtain the feature subset most relevant to lymph node metastasis of breast cancer. (4) Divide the standardized feature data into training set and validation set, and use 10-fold cross-validation to train the model and evaluate its performance; (5) Construct multiple breast cancer lymph node metastasis prediction models based on the training set, including but not limited to GLM model, SVM model, MLP model or DNN model; (6) Train the model and evaluate its performance in predicting lymph node metastasis in breast cancer. Evaluation metrics include AUC, sensitivity, specificity, positive predictive value (PPV) and negative predictive value (NPV). (7) Determine the optimal model for predicting the risk of lymph node metastasis in breast cancer patients, wherein the optimal model is preferably the GLM model.

[0082] In another preferred embodiment, the data used for model training and validation are derived from plasma cfDNA samples from patients with clinically and pathologically confirmed breast cancer.

[0083] In another preferred embodiment, the processing steps for the cfDNA sample include: (1) Collect 8–10 mL of blood from the peripheral blood of the subject and separate it into plasma by centrifugation twice; (2) cfDNA was purified using a commercially available cfDNA extraction kit based on magnetic beads. (3) Construct sequencing libraries and perform sequencing experiments on the Illumina platform; (4) Use software tools such as FastQC, Cutadapt, Ktrim, BWA-MEM, Samtools and Picard to perform quality control and comparison processing on the raw data of whole genome sequencing (WGS).

[0084] In another preferred example, breast cancer cases included in the study were recorded according to clinicopathological type, but the study focused on the differences in lymph node metastasis.

[0085] In another preferred example, a total of 9,999 candidate molecular features were extracted from cfDNA sequencing data for subsequent modeling.

[0086] In another preferred embodiment, the cfDNA fragment size distribution features include: the overall fragment size distribution pattern, and the ratio of small to large fragments (S / L ratio) based on the genomic window division, where small fragments are defined as 100–150 bp and large fragments are defined as 151–220 bp, and the ENCODE blacklist region and UCSC gap track region are excluded in the analysis.

[0087] In another preferred embodiment, the sequence preference of cfDNA break sites is characterized by statistically analyzing the frequencies of 4-base and 6-base motifs and normalizing them into a vector that sums to 1.

[0088] In another preferred embodiment, the feature selection method employs Lasso regression to screen for key cfDNA features that are highly correlated with lymph node metastasis.

[0089] In another preferred embodiment, the key cfDNA feature highly associated with lymph node metastasis includes the signal intensity of the LTR repeat sequence region.

[0090] In another preferred embodiment, the high-throughput sequencing experiment was performed using the Illumina NovaSeq X Plus platform, with a sequencing strategy of 150 bp paired-end sequencing and a sequencing depth of approximately 2 × coverage.

[0091] In another preferred embodiment, bootstrapping and 10-fold cross-validation strategies are combined during model training and evaluation to ensure the robustness and generalization ability of the prediction results.

[0092] In another preferred embodiment, the model performs as follows in predicting lymph node metastasis in breast cancer patients: the AUC reaches 0.9 in the training set data and the AUC value is 0.874 in the validation set data, and the overall prediction accuracy remains at a high level.

[0093] In a specific embodiment, the present invention provides a method for predicting lymph node metastasis in breast cancer. The development process of the method of the present invention includes the following steps: (1) Obtain clinical information data of the subjects and collect peripheral blood samples from the subjects; separate plasma and extract cfDNA from the blood samples, and then perform high-throughput sequencing to obtain raw sequencing data; (2) Perform quality control on the sequencing data, including removing low-quality sequences, adapter contamination and abnormal fragments, to ensure the reliability of downstream analysis; (3) Based on the comparison results, extract cfDNA feature information related to the risk of lymph node metastasis in breast cancer, the features including but not limited to: 1) Characteristics of cfDNA fragment length distribution; 2) Fragment enrichment features of a fixed window; 3) Motif characteristics of cfDNA breakpoint sequences; 4) Base preference characteristics; 5) Signal characteristics covered by the transpose element area; 6) Nucleosome localization characteristics; 7) Copy number variation (CNV) characteristics; (4) The 9999-dimensional cfDNA features initially extracted were processed by feature selection methods such as LASSO regression to screen a subset of key features highly correlated with lymph node metastasis in breast cancer for subsequent modeling; wherein, the subset of key features highly correlated with lymph node metastasis in breast cancer; (5) The model training and validation adopts a strategy of bootstrapping combined with 10-fold cross-validation. The training dataset is randomly divided into 10 subsets. Each time, 9 subsets are selected for training and the remaining 1 is used for testing. This process is repeated 10 times to ensure that each subset is used as a test set once, thereby improving the model's stability and generalization ability. (6) During the training process, samples from patients with positive lymph node metastasis and patients without metastasis were used to construct a prediction model through cfDNA molecular characteristics (such as motif frequency, fragmentation ratio, etc.) and generate a "predicted score" for the risk of lymph node metastasis. The score range is 0-1, and the higher the value, the greater the possibility of metastasis. In the model's judgment rules, the predicted score is set with a classification threshold of 0.5: when the predicted score is greater than 0.5, the model judges the patient as "positive for lymph node metastasis"; when the predicted score is less than or equal to 0.5, it is judged as "no lymph node metastasis". (7) Construct various breast cancer lymph node metastasis prediction models based on training datasets. The models include, but are not limited to: generalized linear model (GLM), support vector machine (SVM), random forest (RF), gradient booster (GBM), extreme gradient booster model (XGBoost), lightweight gradient booster (LightGBM), extreme learning machine (ELM), multilayer perceptron (MLP), and deep neural network (DNN), etc. (8) Train the model and evaluate its performance on the training set, validation set and external independent validation set. The evaluation metrics include, but are not limited to, area under the curve (AUC), sensitivity, specificity, positive predictive value (PPV) and negative predictive value (NPV).

[0094] In another preferred embodiment, the machine learning model includes: a generalized linear model (GLM), a support vector machine (SVM), a multilayer perceptron (MLP), and a deep neural network (DNN).

[0095] In another preferred embodiment, the clinical information data includes: age, tumor stage, molecular subtype, pathological lymph node status, or a combination thereof.

[0096] In another preferred embodiment, step (1) further includes: (1-a) Collect 8–10 mL of peripheral blood samples from the subjects and place them in a dedicated cfDNA preservation tube (e.g., Ardent BioMed, model BY10240301). The plasma is obtained by two-step centrifugation at 1,600×g and 16,000×g. The plasma samples are stored at –80 °C and transported to the central laboratory for unified processing using dry ice. (1-b)cfDNA was extracted using the magnetic bead method (e.g., Vazyme, model N913). The extract was used for library construction after being detected by Qubit fluorescence quantitative quantification. The library preparation was based on the Illumina platform library preparation kit, and the sequencing method was 150 bp paired-end sequencing, with an average sequencing depth of about 2×, thus obtaining the raw sequencing data.

[0097] The main advantages of this invention include: (a) This invention constructs a breast cancer lymph node metastasis prediction model by collecting plasma cfDNA and integrating its fragment omics, copy number variation (CNV), and sequence characteristics, combined with machine learning algorithms. This model can accurately identify the risk of lymph node metastasis in patients, significantly improving prediction efficiency and accuracy compared to traditional imaging and single molecular marker detection, while maintaining good robustness and generalization ability.

[0098] (b) The method of the present invention is non-invasive and allows for repeatable sampling, enabling dynamic monitoring of lymph node metastasis in breast cancer patients while avoiding the invasive risks of surgical excision or pathological biopsy. This method offers high throughput and low cost, making it feasible and convenient for large-scale clinical application and suitable for various clinical scenarios such as preoperative risk assessment, auxiliary diagnosis, and long-term follow-up monitoring.

[0099] (c) The method of the present invention can effectively distinguish between patients with lymph node metastasis risk and those without metastasis. Considering that lymph node metastasis is an important risk factor for breast cancer progression, recurrence, and distant metastasis, this method is expected to help clinicians perform stratified management before surgery, optimize individualized treatment strategies, avoid overtreatment or undertreatment, and thus improve the overall survival rate and long-term prognosis of breast cancer patients.

[0100] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Experimental methods in the following embodiments, unless otherwise specified, are generally performed under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or as recommended by the manufacturer. Unless otherwise stated, percentages and parts are weight percentages and parts by weight.

[0101] Example 1: Experimental method.

[0102] (1) Recruitment of research subjects and sample collection In this embodiment, a total of 404 patients with pathologically confirmed breast cancer were included. All subjects signed written informed consent forms, and the study protocol was approved by the ethics committee and complies with the relevant requirements of the Declaration of Helsinki.

[0103] Of the patients, 115 were confirmed by surgical pathology to have lymph node metastasis, and 289 were not. Exclusion criteria included: pregnancy or lactation, coexisting malignant tumors, suspected but undiagnosed tumors within the past year, history of blood transfusion within 30 days, active inflammation, or serious comorbidities that made them unsuitable for participation in the study.

[0104] During preoperative examinations or routine physical checkups, 8–10 mL of peripheral blood is collected using a dedicated cfDNA blood collection tube. Immediately after collection, the blood is centrifuged in two steps at 4 °C (first step: 1,600 × g for 10 min; second step: 16,000 × g for 10 min) to obtain cell-free plasma. The resulting plasma samples are frozen at -80 °C and transported to the central laboratory on dry ice for further analysis.

[0105] (2) cfDNA extraction, library construction, and sequencing One mL of plasma was collected, and cfDNA was extracted using a commercial magnetic bead cfDNA extraction kit. The concentration and quality were measured using a Qubit real-time fluorescence analyzer. The extracted cfDNA from each sample was stored at -80 °C until library construction.

[0106] 5–20 ng cfDNA were collected and libraries were constructed using a universal library construction kit, followed by barcoding. Library quality was assessed using fragment size distribution analysis with an Agilent 2100. Sequencing was performed on the Illumina NovaSeq platform using a 150 bp paired-end sequencing strategy, with an average sequencing depth of ultra-low-pass (2×) whole-genome sequencing.

[0107] (3) Sequencing data preprocessing and alignment The raw sequencing data were first quality controlled using FastQC, followed by adapter sequence removal and low-quality fragment cutting using Cutadapt and Trimmomatic. The cleaned high-quality reads were aligned to the human reference genome (GRCh37) using BWA-MEM, and then sorted, indexed, and subjected to PCR duplicate removal using Samtools and Picard.

[0108] The reference genome was divided into fixed windows (5 Mb or 100 kb sliding windows) for subsequent fragmentomics and copy number variation feature extraction.

[0109] (4) cfDNA feature construction and normalization 1) cfDNA fragment length characteristics: The distribution and ratio of short fragments (S, 100–150 bp) and long fragments (L, 151–220 bp) were statistically analyzed. The S / L ratio was calculated and standardized using a sliding window to reflect the fragment shortening trend of tumor-derived cfDNA in metastatic patients.

[0110] 2) Motif characteristics of cfDNA break sites: Three bases before and after the 5' end of the cfDNA fragment were extracted to form 6-mer motifs. A break sequence frequency matrix was constructed. Specifically, for each sample, the frequency of 6-mer motifs corresponding to all cfDNA fragments was counted, constructing a frequency matrix with the sample as the row and each possible 6-mer motif as the column. Each row represents one sample, and each column represents one motif feature dimension. The matrix was then row-normalized: the frequency of all motifs in each row was divided by the total number of fragments in that row, ensuring that the sum of all motif frequencies in each row was always 1, thus eliminating the influence of the total amount of cfDNA in different samples. Metastatic patients showed a significant preference in certain sequence environments, suggesting an association with abnormal chromatin remodeling and nucleosome localization.

[0111] 3) Transposon element region coverage characteristics: The coverage intensity of cfDNA in transposon element regions such as LINE, SINE, and LTR was statistically analyzed. Based on the annotation information of various TE (transposon element) families in the human genome, the position of each element in the whole genome was extracted. Whether the 5' end or midpoint of the sample cfDNA fragment falls into the target region was statistically analyzed, and normalized with the total number of fragments to construct fragmentomics feature values ​​reflecting the enrichment intensity of SINE, LINE, and LTR regions. The study found that the fragment signal in the LTR and LINE regions was significantly increased in patients with positive lymph node metastasis, suggesting that this feature can serve as a marker of metastasis risk.

[0112] 4) Copy number variation (CNV) characteristics: A 5 Mb window was used to scan the entire genome for fragment coverage and draw a CNV map. Extensive chromosomal amplification / deletion is common in metastatic patients and can be used as an auxiliary diagnostic indicator.

[0113] 5) Nucleosome localization characteristics: The distribution pattern of cfDNA in the CTCF binding region (within ±1 kb of the center of each high-confidence CTCF binding site) was statistically analyzed to construct a meta-coverage map. Specifically, high-confidence CTCF binding sites confirmed by high-throughput ChIP-seq experiments from the ReMap database were selected. A symmetrical analysis window was constructed within ±1000 bp of the center of each site, and this region was divided into multiple fixed-size bins (10 bp). The coverage density of cfDNA fragments was accurately calculated to generate a meta-coverage map, which serves as an important indicator reflecting the local nucleosome arrangement and chromatin accessibility. Metastatic positive patients showed significant abnormalities, reflecting changes in the three-dimensional structure of chromatin.

[0114] 6) Fixed window fragment enrichment features: The whole genome coverage is statistically analyzed in units of 10 bp to identify chromosomal regions with abnormal fragment enrichment or deletion in patients with lymph node metastasis.

[0115] 7) Base preference characteristics: Analyze the composition of nucleotides before and after the 5' end of the cfDNA fragment, extract the proportion of single bases and the frequency of two-base combinations at the breakpoint and its adjacent region (6-mer), and extract the sequence preference pattern by calculating the proportion or combination frequency of A, T, C, and G bases to reveal the differences in cfDNA fragmentation mechanisms.

[0116] (5) Model training and cross-validation Based on the aforementioned feature matrix, LASSO regression was first used for feature selection to obtain a subset of key transition discriminant features (200–8000 dimensions). Subsequently, various machine learning algorithms (GLM, SVM, MLP, DNN, etc.) were used for model training.

[0117] During training, 10-fold cross-validation and bootstrapping were used to reduce the risk of overfitting. The model output is a "metastasis risk score" of 0–1, with a higher score indicating a greater likelihood of lymph node metastasis.

[0118] (6) Model evaluation and determination of the best model The performance of candidate models was compared on the training and validation sets. Evaluation metrics included AUC, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV). The results are as follows: Figure 2 , Figure 3 As shown.

[0119] The overall results show that the generalized linear model (GLM) performs best in terms of robustness and interpretability, and was selected as the final model, named TuFEst-LNM (Tumor Fragmentomics-based Estimator for Lymph Node Metastasis).

[0120] This model can significantly improve the non-invasive prediction ability of lymph node metastasis risk in breast cancer patients, providing a powerful tool for preoperative staging and individualized treatment planning.

[0121] (7) Model performance verification and feature contribution analysis Furthermore, to verify the performance stability and feature contribution of the TuFEst-LNM model, the applicant conducted model training and evaluation based on different feature sets. First, the model was constructed using only LTR-type features and trained using GLM. The results showed that the area under the receiver operating characteristic (AUC) curve of the model on the training set exceeded 0.83, indicating that LTR repeat sequence region features have high discriminative ability in predicting lymph node metastasis in breast cancer. Figure 4 and Figure 5 As shown.

[0122] Based on this, the applicant incorporated LTR features along with fragment enrichment features within a fixed window, base preference features, cfDNA fragment length distribution features, motif features of cfDNA breakpoint sequences, nucleosome localization features, and copy number variation features into the training process to construct a comprehensive prediction model based on all features. The training set results are as follows: Figure 6 and Figure 7 As shown, the results of the validation set are as follows: Figure 8 and Figure 9 As shown in the figure. The results show that the comprehensive feature model achieves an AUC of 0.90 on the training set and an AUC of 0.874 on the independent validation set, which is significantly higher than the model based solely on LTR features.

[0123] The above results demonstrate that LTR-type features make a significant contribution to the prediction of lymph node metastasis in breast cancer, and their combined use with other features can further improve the overall performance and generalization ability of the model, thus validating the robustness and clinical application potential of the TuFEst-LNM model.

[0124] Experimental Example 2: Detection of lymph node metastasis prediction in 125 breast cancer patients.

[0125] (1) Information processing of breast cancer patients In a preferred embodiment of the present invention, a total of 125 pathologically confirmed breast cancer patients were enrolled from multiple clinical centers, including 40 patients with clinically diagnosed lymph node metastasis and 85 patients with clinically diagnosed non-metastatic lymph nodes. All subjects completed informed consent signing, and plasma samples were collected and cfDNA was extracted.

[0126] The patients' basic clinical information includes: gender, clinical center of origin, age, enrollment number, clinical stage (T stage, N stage, M stage), and lymph node metastasis status (metastasis / non-metastasis). This information is used for clinical labeling and subsequent model evaluation.

[0127] (2) External validation design of the model To verify the generalization ability of the TuFEst-LNM prediction model described in this invention on different datasets, the aforementioned 125 patients were used as an independent external validation set. This dataset was not used for any model training or cross-validation to ensure the objectivity and reliability of the validation results.

[0128] For each patient's cfDNA data, feature extraction was performed according to the method described in Example 1 to obtain a feature set containing multi-dimensional indicators such as fragment length distribution features, fragment omics region coverage features, base preference features, and GC content range features. The feature set was then input into the trained TuFEst-LNM prediction model to obtain the lymph node metastasis prediction results for each patient.

[0129] (3) Comparison of prediction results with clinical labels The predictions output by the TuFEst-LNM model were compared one by one with the clinicopathological labels. The results are as follows: Figure 10 and Figure 11 As shown: Of the 40 patients who actually developed lymph node metastasis, the model correctly predicted 24 cases as having lymph node metastasis and 16 cases as not having metastasis. Of the 85 patients who did not actually have lymph node metastasis, the model correctly predicted that 83 cases had no metastasis, while only 2 cases were predicted to have metastasis.

[0130] Statistical calculations showed that the overall prediction accuracy reached 85.6%, indicating that the model has high applicability and stability in external clinical samples.

[0131] (4) Visualization of results To visually demonstrate the consistency between model predictions and actual clinical classifications, the above results were analyzed using visualization. For example... Figure 12 As shown: 1) The pie chart on the left shows that among 40 patients with actual lymph node metastasis, 24 were accurately predicted to have metastasis and 16 were predicted not to have metastasis. 2) The pie chart on the right shows that among the 85 patients who did not actually metastasize, 83 were accurately predicted as non-metastatic, and only 2 were predicted as metastatic.

[0132] The visualization clearly demonstrates the sensitivity and specificity of the model's predictions.

[0133] (5) Technical effects and application value The verification results of this embodiment show that the TuFEst-LNM model described in this invention can accurately predict lymph node metastasis in breast cancer patients based on cfDNA fragment omics characteristics. Compared with traditional methods relying on imaging or tissue biopsy, this invention has the following advantages: 1) Non-invasive: Predictive results can be obtained through plasma cfDNA, avoiding invasive sampling; 2) High efficiency: The prediction process only requires data input and calculation, making it suitable for rapid clinical assessment; 3) Generalization: Validation in independent, multicenter clinical samples shows that the model has good stability and generalizability; 4) Clinical value: It can assist doctors in conducting individualized risk assessments, and is particularly suitable for subtyping identification and treatment decision-making in cases of low HER2 expression.

[0134] In summary, this embodiment demonstrates the reliability and clinical application potential of the predictive model of the present invention in the identification of lymph node metastasis in breast cancer, laying the foundation for its subsequent large-scale promotion and translational application.

[0135] All documents mentioned in this invention are incorporated herein by reference as if each document were individually incorporated by reference. Furthermore, it should be understood that after reading the foregoing teachings of this invention, those skilled in the art can make various alterations or modifications to this invention, and these equivalent forms also fall within the scope defined by the appended claims.

Claims

1. A method for constructing a non-invasive predictive model for lymph node metastasis in breast cancer, characterized in that, Includes the following steps: (S1) Provides sequencing results of cfDNA in the sample; (S2) The sequencing results are preprocessed and quality controlled to obtain sequencing data that meets the quality control conditions; (S3) Extract cfDNA features from sequencing data that meet quality control conditions, wherein the cfDNA features include the signal intensity of LTR repeat sequence regions; (S4) A machine learning model is constructed based on the cfDNA features to obtain a non-invasive prediction model for lymph node metastasis in breast cancer. The signal intensity of the LTR repeat sequence region refers to the ratio of the number of cfDNA fragments whose 5' end or midpoint falls into the LTR repeat sequence region to the total number of cfDNA fragments.

2. The method as described in claim 1, characterized in that, The status of lymph node metastasis in breast cancer includes: metastasis has occurred and metastasis has not occurred.

3. The method as described in claim 1, characterized in that, The cfDNA features also include one or more cfDNA features selected from the group consisting of: fixed window fragment enrichment features, base preference features, transposon element region coverage signal features, cfDNA fragment length distribution features, motif features of cfDNA break site sequences, nucleosome localization features, and copy number variation features.

4. The method as described in claim 3, characterized in that, The cfDNA fragment length distribution characteristic refers to the ratio S / L of the number of short fragment cfDNA and the number of long fragment cfDNA; The short fragment length cfDNA is cfDNA with a fragment length of 100-150 bp, and the long fragment length cfDNA is cfDNA with a fragment length of 151-220 bp.

5. The method as described in claim 3, characterized in that, The motif characteristics of the cfDNA break site sequence refer to the frequency of the 6-mer motif corresponding to each cfDNA fragment in the sample. The 6-mer motif comprises three bases upstream and three bases downstream of the 5' end of the cfDNA fragment.

6. The method as described in claim 3, characterized in that, The base preference features include: single base frequencies upstream and downstream of the 5' end of the cfDNA fragment; and / or Frequency of dibase combinations upstream and downstream of the 5' end of the cfDNA fragment.

7. The method as described in claim 1, characterized in that, The machine learning models include: Generalized Linear Model (GLM), Support Vector Machine (SVM), Multilayer Perceptron (MLP), and Deep Neural Network (DNN).

8. A non-invasive predictive model for lymph node metastasis in breast cancer, characterized in that, The non-invasive prediction model is constructed using the construction method described in claim 1.

9. The use of the non-invasive prediction model according to claim 8, characterized in that, This is used to prepare a non-invasive prediction system for predicting lymph node metastasis in breast cancer.

10. A non-invasive prediction system for lymph node metastasis in breast cancer, characterized in that, The system includes the following modules: (Z1) Data input module, which is configured to input the sequencing results of cfDNA in the sample to be analyzed; (Z2) Preprocessing and quality control module, wherein the preprocessing and quality control module is configured to preprocess and control the sequencing results to obtain sequencing data that meets the quality control conditions; (Z3) Feature extraction and selection module, the feature extraction and selection module is configured to: extract cfDNA features from sequencing data that meet quality control conditions, wherein the cfDNA features include: signal intensity of LTR repeat sequence regions; (Z4) Breast cancer lymph node metastasis evaluation module, wherein the breast cancer lymph node metastasis evaluation module is configured to: predict the breast cancer lymph node metastasis status of the sample to be analyzed based on the cfDNA characteristics using the non-invasive prediction model described in claim 8; (Z5) Output module, which is configured to output the lymph node metastasis status of the sample to be analyzed.

Citation Information

Patent Citations

  • Breast cancer lymph node metastasis prediction system based on gene spectrum

    CN120998502A

  • Compositions and Methods for Detection, Prognosis and Treatment of Breast Cancer

    US20090118175A1

  • Methods of diagnosis and therapeutic targeting of clinically intractable malignant tumors

    US20180057890A1