Plasma lipid isomer biomarker for predicting neoadjuvant therapy response of breast cancer and application

Through the prediction model constructed by high-throughput mass spectrometry technology and LightGBM algorithm, seven plasma lipid isomer biomarkers were screened out, solving the problem of inaccurate prediction of neoadjuvant treatment response in breast cancer, achieving efficient and accurate prediction of treatment response, and improving the individualized guidance ability of treatment.

CN120404966APending Publication Date: 2025-08-01CANCER INST & HOSPITAL CHINESE ACADEMY OF MEDICAL SCI +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510342871.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the response to neoadjuvant therapy in breast cancer through plasma lipomics methods, especially due to the lack of high-precision lipid isomer analysis tools, resulting in inaccurate prediction of treatment responses, which may lead to undertreatment or overtreatment.

Method used

The prediction model was constructed using high-throughput mass spectrometry technology and LightGBM algorithm. By analyzing the characteristic values of lipid isomers in the patient's plasma, seven plasma lipid isomer biomarkers including PC 15:0_18:1(n-10) and PC 15:0_18:1(n-7) were screened out, and a prediction model for neoadjuvant treatment response in breast cancer was combined with machine learning algorithms.

Benefits of technology

Efficient and accurate prediction of responses to neoadjuvant therapy for breast cancer is achieved. The AUC of the model on the test set and the external validation set is 0.957 and 0.936, respectively, providing scientific basis to guide individualized treatment and reduce the risk of under-treatment or over-treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120404966A_ABST
    Figure CN120404966A_ABST
Patent Text Reader

Abstract

The invention relates to a plasma lipid isomer biomarker for predicting neoadjuvant therapy response of breast cancer and application. The biomarker is PC 15: 018: 1 (n-10), PC 15: 018: 1 (n-7), PC 14: 018: 1 (n-9) / PC 14: 018: 1 (n-7), PE 18: 1 (n-9) 18: 2 / PE 18: 1 (n-7) 18: 2, PE 18: 1 (n-9) 18: 2, PC 16: 018: 1 (n-9) / PC 16: 018: 1 (n-7), and PE 18: 1 (n-7) 18: 2. In addition, a prediction model of breast cancer NAT treatment response is also constructed, the problems of tumor heterogeneity and difficulty in pathological biopsy sampling are solved, and the prediction model shows excellent prediction capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a plasma lipid isomer biomarker for predicting the response of breast cancer to neoadjuvant therapy and its application, belonging to the field of biomedical technology. Background Art

[0002] Tumor heterogeneity is a key problem in the diagnosis and treatment of breast cancer, increasing the complexity of breast cancer treatment options and the uncertainty of treatment efficacy. Early diagnosis, efficacy prediction and prognosis improvement of breast cancer have always been the research focuses of clinical researchers.

[0003] Neoadjuvant Therapy (NAT) refers to the systemic drug treatment before surgery, which has become a very important part of the comprehensive treatment strategy for breast cancer. Its advantages include downstaging the primary tumor, increasing the breast-conserving surgery rate, predicting tumor treatment sensitivity, and evaluating the treatment benefit of patients. The Pathological Complete Response (pCR) rate is a powerful indicator for evaluating the response of NAT and an important basis for adjusting the adjuvant treatment plan. Previous studies have shown that patients who achieve pCR after neoadjuvant therapy have better disease-free survival (DFS) and overall survival (OS) than those with residual tumors. Therefore, pCR is used as a surrogate endpoint for EFS in prospective neoadjuvant chemotherapy clinical studies. However, 40-60% of patients still have residual lesions after NAT, and some patients may also be over-treated, suffering from treatment-related adverse reactions and the risk of death. Therefore, it is urgent to explore and discover biomarkers that can predict the response to NAT, so as to more accurately guide the individualized treatment of breast cancer and avoid under-treatment or over-treatment.

[0004] Lipids are important energy donors for life activities, and also the structural basis of biological membranes and participate in biological information transmission, playing an important role in the growth, invasion and metastasis of tumors. Lipid metabolism plays a crucial role in the process of cancer progression, involving profound remodeling to meet the higher requirements of rapidly proliferating cancer cells for metabolism. In breast cancer, increased lipogenesis and enhanced lipolysis pathways provide the necessary components for cell division, membrane formation and energy production, thus supporting tumor growth. Lipid metabolism remodeling promotes tumor progression and metastasis by affecting the tumor microenvironment. Lipidomics is the systematic study of lipid characteristics and metabolic pathways, and has become a powerful tool for cancer research. The lipidome is directly related to disease phenotypes and is an important bridge between genomics, proteomics, etc. and disease models. By analyzing lipidomics data in tumor tissues, lipid biomarkers or lipid metabolic pathways related to treatment response can be discovered. Changes in lipid metabolism characteristics can predict treatment response, suggesting that changes in lipid metabolism characteristics may be potential biomarkers for predicting pCR, providing more comprehensive information and more effective treatment strategies for the precise diagnosis and treatment of breast cancer.

[0005] However, the structures of lipids are extremely complex, with more than 40,000 different structures. Variations in multiple dimensions such as the head group, carbon chain, C=C position, sn-position, chain modifications, and stereoconfiguration lead to rich and subtle structural differences and corresponding differences in biological functions, which have important impacts on life activities. The main limitation of current traditional lipidomics methods is the inability to resolve the detailed structural characteristics of lipid species, such as the position specificity of carbon-carbon double bonds (C=C) widely present in biological systems and the sn-position of fatty acyl or alkyl chains. These differences play important roles in regulating the biological behaviors of tumor cells. There is still a need for analytical techniques suitable for large-scale resolution of lipid molecular structures and spatial distribution information for comprehensive qualitative and quantitative analysis of lipids in cells, tissues, or organisms. The latest advances in tandem mass spectrometry (MS) techniques, including ultraviolet photodissociation, ozone-induced dissociation, and photochemical derivatization, have addressed this limitation and enabled high-resolution structural characterization of lipid isomers. These innovations have revealed the unique functional roles of specific lipid isomers. For example, position isomers of C=C bonds have been shown to affect key biophysical properties such as membrane permeability, thickness, and cholesterol-phospholipid interactions, all of which affect cell flexibility and may contribute to increased adaptability of cancer cells. Despite these advances, the diagnostic potential of lipid isomers remains largely unexploited due to the lack of powerful analytical workflows and scalable large-scale biomarker validation tools. Developing a comprehensive platform for identifying and characterizing lipid isomer biomarkers is crucial for promoting their application in the precision diagnosis and treatment of breast cancer.

[0006] Previous studies have confirmed the value of using lipidomics to detect breast cancer tumor tissues for predicting treatment efficacy and patient prognosis. Given that previous plasma lipidomics analysis of breast cancer patients was limited to secondary mass spectrometry, with limited biomarkers detected and huge individual differences, there is still a need for in-depth structural identification of plasma lipid molecules and ultimately the construction of a more efficient prediction model for neoadjuvant therapy efficacy of breast cancer based on high-precision structural lipidomics. Summary of the Invention

[0007] In view of the above problems, the present invention provides a plasma lipid isomer biomarker for predicting the response to neoadjuvant therapy (NAT) of breast cancer and its application. The present invention constructs a NAT treatment response prediction model based on the LightGBM algorithm. This model can accurately predict the response probability of patients to neoadjuvant therapy by analyzing the lipid isomer characteristic values in the plasma of patients, thus providing a scientific basis for clinical practice. The technical solution of the present invention is as follows: The present invention uses high-throughput mass spectrometry technology to detect and analyze lipid fine-structure metabolites in plasma samples of breast cancer patients undergoing neoadjuvant chemotherapy, and screens and obtains plasma lipid isomer biomarkers for predicting the response of breast cancer to neoadjuvant therapy. The plasma lipid isomer biomarkers are as follows: PC 15:0_18:1(n-10), PC 15:0_18:1(n-7), PC 14:0_18:1(n-9) / PC 14:0_18:1(n-7), PE 18:1(n-9)_18:2 / PE 18:1(n-7)_18:2, PE 18:1(n-9)_18:2, PC 16:0_18:1(n-9) / PC 16:0_18:1(n-7), PE 18:1(n-7)_18:2.

[0008] Furthermore, the present invention also provides a prediction model for the response of breast cancer to neoadjuvant therapy. The prediction model is constructed using the machine learning algorithm LightGBM, and the response of breast cancer to neoadjuvant therapy is predicted through the above-mentioned plasma lipid isomer biomarkers.

[0009] Furthermore, the construction method of the prediction model is as follows: (1) The data set is divided into a training set and a test set. Among them, 70% is the training set and 30% is the test set; on the training set, the model will perform feature selection and train the LightGBM model based on the training data; in each iteration, the model gradually constructs a new decision tree based on the residual information predicted previously to correct the prediction result; (2) Through 5-fold cross-validation, hyperparameters such as the number of decision trees, learning rate, and maximum depth are optimized; (3) Use the test set to evaluate the final model; (4) Substitute the expression level values of the 7 plasma lipid isomer biomarkers described in claim 1 of the patient into the model to output a probability value p. If p exceeds a specific cutoff threshold, it can be determined as a positive diagnosis.

[0010] Even further, the construction method of the prediction model is specifically as follows: Step 1: Data set division Divide the data set into a training set and a test set. Use 70% of the data as the training set and the remaining 30% of the data as the test set; the training set is used for training and tuning the model, while the test set is used to evaluate the final performance of the model.

[0011] Step 2: Feature selection In the training set, the maximum information coefficient (MIC) feature selection method is used to perform secondary screening on the lipidomics data to obtain candidate lipid features; for clinical variables, variable transformation and segmentation are performed, and then mutual information screening is performed again to obtain clinical features; the lipid features and clinical features are combined to form the input matrix features.

[0012] Step 3: Model training The LightGBM algorithm is used to train the training set; during the training process, the model first initializes a baseline prediction value; then, in each iteration, by calculating the residual or gradient information of the previous model, a new decision tree is constructed to correct the model's prediction; each time a new decision tree is added, it helps to gradually reduce the prediction error and finally form a comprehensive score; this comprehensive score is converted into a probability value through the Logistic (Sigmoid) function, indicating the likelihood that the patient will have a good response to neoadjuvant therapy (NAT).

[0013] Step 4: Cross-validation To avoid overfitting and improve the generalization ability of the model, 5-fold cross-validation is used to evaluate the model.

[0014] The steps of the cross-validation are as follows: the dataset is divided into 5 subsets, and each time 4 subsets are used to train the model, and the remaining 1 subset is used for validation; this is repeated 5 times to ensure that the model can perform stably on different data subsets and avoid biases caused by specific data partitions.

[0015] Step 5: Hyperparameter tuning During the cross-validation process, the hyperparameters of the LightGBM model are adjusted to optimize the model performance. The main hyperparameters to be tuned include the number of decision trees, learning rate, maximum depth, etc. Through methods such as grid search or random search, the optimal combination of hyperparameters is selected to improve the prediction accuracy and generalization ability of the model.

[0016] Step 6: Final evaluation and output After hyperparameter tuning and cross-validation are completed, the test set is used to evaluate the final model; the expression level values of the 7 plasma lipid isomer biomarkers as described in claim 1 of the patient are substituted into the model to output the probability value p. If p exceeds a specific cutoff threshold, it can be judged as a positive diagnosis.

[0017] Further, the cutoff threshold of the prediction model of the present invention is 0.536. The present invention uses Youden's index to define the optimal critical value. The calculation formula of Youden's index is: J = Sensitivity + Specificity - 1. During the calculation process, by selecting the cutoff value corresponding to the maximum Youden's index, the model is ensured to have the best sensitivity and specificity.

[0018] The present invention has the following advantages compared with the prior art: 1. The present invention overcomes the problems of tumor heterogeneity and the difficulty of obtaining pathological biopsy samples, and provides a new type of liquid biopsy, namely a stable lipid fine structure metabolic biomarker of breast cancer NAT plasma; at the same time, a prediction model for the treatment response of breast cancer NAT is established.

[0019] 2. The plasma lipid isomer biomarker of the present invention has good stability and repeatability, and can be monitored in the plasma of different patients and at different time points, providing a reliable and durable biomarker for the prediction of the treatment response of breast cancer NAT.

[0020] 3. The present invention constructs a prediction model for the treatment response of breast cancer NAT based on plasma lipid metabolism characteristics. By using the advanced machine learning algorithm LightGBM, the model can efficiently and accurately predict the response of patients to neoadjuvant therapy. Among them, for the prediction model of the treatment response of breast cancer NAT constructed by the present invention, the area under the curve (AUC) of the test set is 0.957, and the AUC of the external validation set is 0.936. Description of the Drawings

[0021] Figure 1 It is a volcano plot of differential metabolites, showing the lipid characteristic differences between the pathological complete remission and non-remission groups of breast cancer patients; among them, the horizontal axis is log2 (fold change), indicating the fold change of lipid characteristics between the two groups, and the vertical axis is -log 10 (significance of difference).

[0022] Figure 2 It is an ROC curve graph of different prediction models, used to evaluate the effect of candidate biomarkers on the classification model; Figure 2 The AUC values of the LightGBM, random forest (RF), multi-layer perceptron (MLP), logistic regression (LR), decision tree, and K-nearest neighbor (KNN) models are shown; among them, the horizontal axis is the false positive rate (FPR), that is, 1 - specificity, and the vertical axis is the true positive rate (TPR), that is, sensitivity.

[0023] Figure 3It is a flowchart for the construction of a prediction model for the response of neoadjuvant therapy for breast cancer based on the LightGBM algorithm.

[0024] Figure 4 It is a graph of the lipid isomer biomarker of the present invention and its importance score (using SHAP value).

[0025] Figure 5 It is a receiver operating characteristic curve (ROC curve) graph based on the LightGBM model. Detailed implementation manners

[0026] The technical solutions of the present invention are described in detail below through specific implementation manners. It should be understood that the following specific implementation manners are only exemplary. Any modification or change as long as it does not deviate from the design of the technical solutions of the present invention should be within the scope of the claims of the present invention. The present invention is described in detail below with reference to the embodiments.

[0027] Embodiment: A plasma lipid isomer biomarker for predicting the response of neoadjuvant therapy for breast cancer, and the construction of a prediction model for the response of breast cancer NAT therapy (1) Collection of clinical samples: Based on the optimization of preoperative chemotherapy regimens for different subtypes of breast cancer and the exploration of related biomarkers in the Capital Clinical Characteristics Application Research project (Trial registration: ClinicalTrials.gov NCT02041338) and the screening and application of lipid biomarkers in body fluids. From January 2016 to January 2021, patients with biopsy-confirmed primary invasive breast cancer without distant metastasis were recruited, and the clinical stage diagnosis was from stage IIa to IIIc. All patients received 4-6 cycles of NAT before breast cancer surgery, and the treatment regimens and time arrangements were guided by the National Comprehensive Cancer Network (NCCN) guidelines. The NAT treatment regimens mainly included anthracyclines and taxanes. In addition, HER2-positive patients also received targeted therapy with trastuzumab (8 mg / kg loading dose, followed by 6 mg / kg continuous dose) and pertuzumab (840 mg loading dose, followed by 420 mg continuous dose).

[0028] During the neoadjuvant treatment process, physical examinations and imaging examinations (breast MRI or ultrasound) were performed every 2 cycles to evaluate the clinical efficacy. The efficacy was judged according to the Response Evaluation Criteria in Solid Tumors (RECIST) version 1.1. The clinical efficacy could be divided into complete remission (CR), partial remission (PR), stable disease (SD), and progressive disease (PD).

[0029] (2)Immunohistochemistry and pathological response assessment The status of estrogen receptor (ER), progesterone receptor (PR), HER2, and Ki67 index was determined by immunohistochemistry (IHC). If less than 1% of tumor cells showed nuclear staining, it was classified as ER / PR negative; if 1% or more of tumor cells showed nuclear staining, it was classified as ER / PR positive. The cut-off value of Ki67 was set at 14%. HER2 negative was defined as an IHC score of 0 or 1, and HER2 positive was defined as an IHC score of 3. For tumors with an IHC score of 2, fluorescence in situ hybridization (FISH) was performed to confirm the HER2 status by evaluating the HER2 / CEP17 ratio and HER2 copy number. The National Cancer Center performed standard histopathological analysis to evaluate the pCR of NAT. The specimens were fixed in 10% neutral buffered formalin, processed overnight with a standard tissue processor, cut into 5-mm sections, and stained with an automated staining system. Pathologic complete response (pCR) was defined as no residual invasive carcinoma, allowing for the presence of residual carcinoma in situ, and no invasive disease in the ipsilateral sentinel node or axillary dissection lymph nodes (yPT0 / isN0).

[0030] (3)Peripheral blood specimen collection and processing Consecutive peripheral blood samples of breast cancer patients were collected before NACT treatment (C0) and after two cycles of treatment (C2). A total of 4 ml of fresh peripheral blood was taken from the patients at baseline (C0) and after two treatment cycles (C2) and placed in EDTA anticoagulant tubes. After plasma collection, it was gently mixed immediately and subjected to preliminary separation within two hours. The plasma separation process included centrifuging the blood collection tubes at 4 °C and 1600 g for 10 minutes. After centrifugation, the upper plasma was aliquoted into multiple 2.0 mL centrifuge tubes and centrifuged at 16,000 g at 4 °C for 10 minutes. The resulting supernatant was transferred to a new centrifuge tube as the final plasma sample and stored at -80 °C. For plasma samples, 50 μL of plasma was diluted with 1 mL of deionized water in a 10 mL centrifuge tube. Methanol (1 mL) and chloroform (2 mL) were added, and the mixture was vortexed for 10 minutes and then centrifuged at 13,000 rpm for 15 minutes. The extraction process was repeated, and the combined chloroform layer was dried under a nitrogen stream. The resulting lipid extract was reconstituted in 1 mL of methanol and stored at -20 °C until analysis.

[0031] (4) High-precision lipidomics (LC-PB-MS / MS) detection and analysis including isomer analysis A total of 210 people were included in the study, and their plasma samples before and after NAT treatment were subjected to structural lipidomics analysis. The multidimensional structure included MS1, fatty acyl chain level, sn-position, and C=C position level. The LC-PB-MS platform system consisted of a Nexera SIL-30ACFV UPLC system (Shimadzu, Tokyo, Japan), a 9030 Q-TOF mass spectrometer (Shimadzu, Tokyo, Japan), and a flow microreactor OMEGA (PURSPEC, Technology Inc.). The liquid chromatograph was equipped with a degasser, two pumps, an autosampler, and a column oven. Separation was carried out on a BEH UPLC HILIC column (100 mm × 2.1 mm, 1.6 µm, Waters, Milford, MA, USA). The column temperature was set at 30 °C. The mobile phase included A: ACN / acetone isopropanol (50 / 48 / 2, v / v / v) and B: ammonium acetate aqueous solution (10 mM). Gradient elution separation was performed at a flow rate of 0.3 mL / min (starting with 90% of A, decreasing to 85% at 2.4 minutes, to 80% at 3.2 minutes, remaining at 80% from 3.2 - 5 minutes, decreasing to 70% at 5.1 minutes and maintaining this ratio until 6 minutes, increasing to 90% at 6.1 minutes and maintaining this ratio until 10 minutes). Mass spectrometry analysis was performed on a 9030 Q-TOF mass spectrometer (Shimadzu, Tokyo, Japan), and mass spectra in the m / z range of 100 - 900 were recorded in positive electrospray mode. Other parameters were as follows: ESI voltage, 4500 V; curtain gas, 25 psi; CAD gas, 8; interface heater temperature, 350 °C; nebulizing gas 1 and gas 2, 55 psi; declustering potential, 80 V. The post-column PB reaction was carried out using a flow microreactor OMEGA (PURSPEC, Technology Corporation). The inlet and outlet of the microreactor were connected to the LC column and the ESI source through PEEK tubes (inner diameter 0.01 inches, outer diameter 1 / 16 inches, length 80 cm). A low-pressure mercury lamp (model 80 - 1057 - 01, BHK Corporation) emitting centered at 254 nm was equipped.

[0032] (5) Lipidomics analysis and biomarker collection The LC-MS, LC-MS / MS, and LC-PB-MS / MS data were analyzed using the in-house software LipidOA. The raw data were converted into an open file format (.mzmL) using Lab solutions software (Shimadzu, Tokyo, Japan), which can be read by LipidOA. For relative quantification at the total component level, the MS1 spectra of specific lipids and the corresponding internal standard SPLASH were imported into LipidOA, and type I isotope correction was performed. LipidOA reads the LC-MS / MS data in negative ion mode and provides lipid identification at the fatty acyl (alkyl) chain level and sn position level. LipidOA reads the LC-PB-MS / MS data in positive ion mode and provides C=C position identification. Lipid class searches included PC, PE, PG, PI, PS, CER, SM, LPC, and LPE. The precursor ion and MS / MS fragment mass tolerances were set at 10 ppm, respectively. Identification at the lipid chain level, fatty acyl (alkyl) chain level, Sn position level, and C=C position was performed based on the retention times specific to each class. Lipid isomer markers were obtained through fine-structure analysis of lipid isomers in this step and were used in subsequent statistical analyses to predict the response to neoadjuvant treatment for breast cancer. The lipid isomer biomarkers were PC 15:0_18:1(n-10), PC 15:0_18:1(n-7), PC 14:0_18:1(n-9) / PC 14:0_18:1(n-7), PE 18:1(n-9)_18:2 / PE 18:1(n-7)_18:2, PE 18:1(n-9)_18:2, PC 16:0_18:1(n-9) / PC 16:0_18:1(n-7), PE 18:1(n-7)_18:2.

[0033] (6)Establishment of the lipidomic characteristic spectrum of breast cancer Using high-precision relative quantitative lipidomics, the peripheral fine lipidomic profiles at multiple levels, such as lipid composition, subclass, chain length, saturation, and isomers, were dynamically monitored to complete the identification of differential substances and statistical analysis.

[0034] Data preprocessing: Missing values are always represented as "NA". For data without missing values, this step is skipped. Zero values are processed to remove peaks with a zero-value ratio greater than 50% (default setting) in all sample groups, as peaks with a large zero-value ratio are mostly noise peaks. For the acquisition of lipidomics data, especially large-scale lipidomics data, the drift of metabolite signal intensity over time is a major interfering factor. During the data acquisition process (within-batch and between-batch), unnecessary variations in the measurement results of metabolic peaks are inevitable, and the reasons include sample handling and preparation, liquid chromatography column degradation, matrix effects, mass spectrometer contamination, and non-linear instrument drift during long-term operation.

[0035] (7) Pre-screening by lipid metabolism characteristic statistics method To mitigate these effects and address the missing value problem, the present invention compared different estimation methods to understand their effectiveness in correcting these variations. Probabilistic principal component analysis (PPCA) was finally selected because this method is robust in handling missing data and can adjust the drift of signal intensity. Relative quantification of isomers was achieved by calculating the intensity ratio of diagnostic ions corresponding to C=C positional isomers. We performed differential metabolite analysis on all samples in the internal validation cohort. Differential metabolites were determined using the Mann-Whitney U test, and the statistical significance of each lipid characteristic was defined as a p-value less than 0.05 and an F-fold change greater than 1.2. According to the p-value, the most discriminative spectral features were further determined by the t-test, such as Figure 1 shown. All statistical analyses of clinical data were performed using the Chi-square test or Fisher's exact test, and the Welch Two Sample t test. Clinical information was evaluated using the R software package "compareGroups" (v4.2.2).

[0036] (8) Secondary screening by integrated machine learning and NAT treatment response prediction modeling ① Selection of candidate biomarkers First, the research objects are randomly grouped into a training set and a test set. In the training set, the maximum information coefficient (MIC) feature selection method is used to conduct a secondary screening of the lipidomics data, and a total of 76 candidate lipid features are obtained. For the 12 clinical variables, variable transformation and segmentation are performed, and then mutual information screening is carried out again. Finally, 2 clinical features are retained. The above 76 lipid features are combined with 2 clinical features to form an input matrix feature with a total of 78 features. First, these features are combined according to the individual number to form a feature matrix, and the actual response of each patient in neoadjuvant therapy ("significant remission" or "no significant remission") is marked.

[0037] ② Validation and evaluation of the optimal potential biomarker combination model To complete the validation of candidate biomarkers and evaluate the effect of candidate biomarkers on the classification model, three machine learning models are used: Logistic Regression (LR), Random Forest (RF), and Support Vector Machine (SVM). A 5-fold cross-validation is performed on the model constructed by the biomarker combination. Through ROC curve analysis, the performance of candidate biomarkers in classifying different sample groups is determined. On the ROC graph, the horizontal axis represents the false positive rate (FPR), which is 1 - Specificity, indicating the proportion of true negative samples misjudged as positive; the vertical axis represents the true positive rate (TPR), also known as Sensitivity, indicating the proportion of true positive samples correctly judged as positive. The single model index is judged by the area under the ROC curve (AUC value) for specificity and sensitivity. Figure 2 The results of the AUC of different prediction models are shown. Among them, the closer the area under the ROC curve is to 1, the greater its clinical diagnostic efficiency, and the higher the indicators of specificity and sensitivity, the better. In the results of this analysis, the AUC value of the LightGBM model is 0.957, showing the most excellent performance and indicating its best performance in classification tasks.

[0038] ③ Diagnostic model construction According to the above model results, a diagnostic model is constructed using the LightGBM algorithm. LightGBM is a commonly used ensemble learning model based on Gradient Boosting Decision Trees (GBDT), which can predict the probability of an event occurring and has good performance in dealing with high-dimensional features and diverse data. In the present invention, the biomarkers preliminarily screened in the previous analysis are directly used as the input features of the model, and combined with several clinical variables of the patients, a prediction model for the response of neoadjuvant therapy (NAT) for breast cancer is constructed. On the training set, LightGBM will first initialize a baseline prediction value, and then in each iteration, based on the residual or gradient information of the previous model, a new decision tree is constructed to correct the prediction. The model will gradually accumulate the prediction results of each tree and finally form a comprehensive score. This comprehensive score will be transformed through the Logistic (Sigmoid) function to output a probability value within the range of [0,1], which is used to represent the likelihood that the patient will have a good response to neoadjuvant therapy (NAT). To optimize the model performance, 5-fold cross-validation is used to tune the hyperparameters of the number of decision trees, learning rate, and maximum depth to ensure the generalization ability of the model and improve the diagnostic accuracy. And the cutoff threshold is selected according to the results of the training set. The construction process of the model is as Figure 3 shown. The lipid isomer biomarkers and their importance scores (using SHAP values) of the present invention are as Figure 4 shown. Among them, the SHAP values of PC 15:0_18:1(n-10) and PC 15:0_18:1(n-7) are the highest, with SHAP values of 0.812 and 0.431 respectively, indicating that these lipid features contribute greatly to the prediction ability of the model.

[0039] The present invention utilizes the determined 7 biomarkers and clinical features, brings the expression level values of the 7 biomarkers and the patient's clinical information in the test set into the above - constructed model to output the probability value p. If p exceeds a specific cutoff threshold, it can be judged as a positive diagnosis (that is, a good response to NAT treatment). To obtain this cut - off value, we also use the Youden's index to define the optimal critical value for diagnosis determination. The Youden's index is a commonly used indicator to measure the overall diagnostic effectiveness. When equal weights are given to sensitivity and specificity, the critical value corresponding to the maximum Youden's index is the best critical point for the biomarker's discrimination ability. Because at this time the sum of sensitivity and specificity is the largest, the optimal critical value can have good sensitivity and specificity at the same time. In fact, the optimal critical value is not necessarily unique, and there may be multiple ones, which can be selected according to different requirements for sensitivity and specificity. High sensitivity is often applied to: diagnosing diseases with severe conditions but good curative effects to prevent missed diagnosis; diseases that may be caused by multiple diseases to exclude the possibility of a certain disease; general surveys or regular health checks to screen for a certain disease. High specificity is often used in: diagnosing when the probability of a patient having a certain disease is relatively high for confirmation; diseases with severe conditions but poor curative effects and prognoses to prevent misdiagnosis; when the radical treatment method of a disease causes great harm and confirmation is needed to avoid unnecessary harm to the patient. The optimal critical value calculated using the Youden's index and the corresponding test set indicators are shown in Table 1.

[0040] Table 1 Optimal critical value and corresponding test set indicators

[0041] Evaluation of the diagnostic ability of the model: Based on the constructed LightGBM model, the receiver operating characteristic (ROC) analysis is performed using the training set and the test set respectively. The results are as Figure 5 shown. The AUC value of the training set is 0.998, and the AUC value of the test set is 0.957. When the discrimination standard critical value of the binary classification model is set to 0.536, the accuracy of this model in the test set can reach 0.909, with a specificity of 0.870 and a sensitivity of 1.000. It can be seen that the present invention shows good predictive performance for judging the response of patients to NAT.

[0042] Experimental example 1: In this experiment, lipid metabolomics analysis was performed on plasma samples from 119 breast cancer patients to screen for candidate lipidomic features, and an integrated machine learning algorithm was used to construct a prediction model for the response to neoadjuvant therapy (NAT) in breast cancer. Among these samples, several key lipid metabolites were identified through baseline analysis. Through an integrated machine learning algorithm (such as LightGBM), 7 significant biomarkers were screened out: PC 15:0_18:1(n-10), PC 15:0_18:1(n-7), PC 14:0_18:1(n-9) / PC 14:0_18:1(n-7), PE 18:1(n-9)_18:2 / PE 18:1(n-7)_18:2, PE 18:1(n-9)_18:2, PC 16:0_18:1(n-9) / PC 16:0_18:1(n-7), PE 18:1(n-7)_18:2; the 7 biomarkers were detected using the prediction model constructed by the present invention. Finally, the AUC of the model in the validation set was 0.957. When the cutoff threshold was determined to be 0.536, the accuracy of the model was 0.909, the specificity was 0.870, and the sensitivity was 1.000.

[0043] Test Example 2: To verify the results in the first experiment, 91 external validation data samples were further used in this experiment for model evaluation. These external samples were from different clinical sources and were used to test the generalization ability and diagnostic effect of the model. The lipidomic features in the external validation set underwent the same preprocessing process and were compared with the 7 biomarkers selected in Test Example 1: PC 15:0_18:1(n-10), PC 15:0_18:1(n-7), PC 14:0_18:1(n-9) / PC 14:0_18:1(n-7), PE18:1(n-9)_18:2 / PE 18:1(n-7)_18:2, PE 18:1(n-9)_18:2, PC 16:0_18:1(n-9) / PC 16:0_18:1(n-7), PE 18:1(n-7)_18:2. The 91 external samples were predicted using the LightGBM model constructed in Test Example 1, and the diagnostic performance of the model on the external validation set was calculated. The AUC of the model on the external validation set was 0.936, the accuracy was 0.867, the specificity was 0.897, and the sensitivity was 0.800.

Claims

1. A plasma lipid isomer biomarker for predicting the response of breast cancer to neoadjuvant therapy, characterized in that, The plasma lipid isomer biomarkers are: PC 15:0_18:1(n-10), PC 15:0_18:1(n-7), PC 14:0_18:1(n-9) / PC 14:0_18:1(n-7), PE 18:1(n-9)_18:2 / PE 18:1(n-7)_18:2, PE 18:1(n-9)_18:2, PC 16:0_18:1(n-9) / PC 16:0_18:1(n-7), PE 18:1(n-7)_18:

2.

2. A prediction model for the response of neoadjuvant treatment of breast cancer, characterized in that a prediction model is constructed using the machine learning algorithm LightGBM, and the response of neoadjuvant treatment of breast cancer is predicted through the plasma lipid isomer biomarkers as described in claim 1.

3. The method for constructing the prediction model according to claim 2, wherein, The steps of the construction method are as follows: (1) The data set is divided into a training set and a test set, where 70% is the training set and 30% is the test set; on the training set, the model will perform feature selection and train the LightGBM model based on the training data; in each iteration, the model gradually constructs a new decision tree based on the residual information predicted previously to correct the prediction result; (2) Through 5-fold cross-validation, hyperparameters such as the number of decision trees, learning rate, and maximum depth are optimized; (3) The test set is used to evaluate the final model; (4) The expression level values of the 7 plasma lipid isomer biomarkers of the patient as described in claim 1 are substituted into the model to output a probability value p. If p exceeds a specific cutoff threshold, it can be judged as a positive diagnosis.

4. The construction method according to claim 3, characterized in that The specific steps of the construction method are as follows: Step 1: Data set division The data set is divided into a training set and a test set. 70% of the data is used as the training set, and the remaining 30% of the data is used as the test set; the training set is used for the training and tuning of the model, while the test set is used to evaluate the final performance of the model; Step 2: Feature selection In the training set, the maximum information coefficient (MIC) feature selection method is used to perform secondary screening on the lipidomics data to obtain candidate lipid features; For clinical variables, variable transformation and segmentation are performed, and then mutual information screening is performed again to obtain clinical features; the lipid features and clinical features are combined to form an input matrix feature; Step 3: Model training The LightGBM algorithm is used to train the training set; during the training process, the model first initializes a baseline prediction value; then, in each iteration, by calculating the residual or gradient information of the previous model, a new decision tree is constructed to correct the prediction of the model; the addition of each new decision tree helps to gradually reduce the prediction error and finally form a comprehensive score; this comprehensive score is converted into a probability value through the Logistic (Sigmoid) function, indicating the possibility that the patient has a good response to neoadjuvant treatment (NAT); Step 4: Cross-validation To avoid overfitting and improve the generalization ability of the model, 5-fold cross-validation is used to evaluate the model; Step Five: Hyperparameter Tuning During the cross-validation process, the performance of the model is optimized by adjusting the hyperparameters of the LightGBM model; the hyperparameters to be tuned include the number of decision trees, the learning rate, and the maximum depth; Step Six: Final Evaluation and Output After hyperparameter tuning and cross-validation are completed, the test set is used to evaluate the final model; the expression level values of the 7 plasma lipid isomer biomarkers as described in Claim 1 of the patient are substituted into the model to output the probability value p, and if p exceeds a specific cutoff threshold, it can be judged as a positive diagnosis.

5. The construction method according to claim 4, wherein In Step Four described above, the steps of cross-validation are: the dataset is divided into 5 subsets, and each time 4 subsets are used to train the model, and the remaining 1 subset is used for validation; this is repeated 5 times to ensure that the model can perform stably on different data subsets and avoid biases caused by specific data partitions.

6. The construction method according to claim 4, characterized in that In Step Five described above, the optimal hyperparameter combination is selected through grid search or random search methods.

7. The prediction model obtained by the construction method according to any one of claims 3-6, characterized in that The cutoff threshold of the prediction model is 0.536.