Prediction method, device and equipment for reaching pCR through neoadjuvant chemotherapy of triple negative breast cancer and medium

By acquiring peripheral blood count data and clinicopathological features, the XGB model is used to predict the pCR of neoadjuvant chemotherapy for triple-negative breast cancer, which solves the problems of insufficient timeliness and accuracy in the prediction of existing technologies and provides more timely and accurate prediction support.

CN121506249APending Publication Date: 2026-02-10TANGSHAN PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511651377.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

In existing technologies, the prediction of pathological complete response (pCR) for neoadjuvant chemotherapy in triple-negative breast cancer is not timely and has low accuracy, relying on invasive methods and subjective morphological changes for assessment.

Method used

By acquiring peripheral blood count data (neutrophil, lymphocyte, platelet, and monocyte counts) and clinicopathological features (T stage, N stage, and Ki-67) of the subjects to be tested, the XGB model is used to capture the nonlinear relationships and interaction effects between the data, calculate the immune inflammation score, and predict pCR.

Benefits of technology

It enables non-invasive procedures, provides more timely and accurate pCR predictions, improves prediction accuracy, and helps doctors develop or adjust treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506249A_ABST
    Figure CN121506249A_ABST
Patent Text Reader

Abstract

The invention discloses a method, a device, equipment and a medium for predicting the reaching of pCR of neoadjuvant chemotherapy of triple negative breast cancer, and relates to the technical field of medical data processing, and the method comprises the following steps: obtaining clinical pathological characteristics of a to-be-detected object and peripheral blood counting data before neoadjuvant chemotherapy; the peripheral blood count data comprises neutrophil count, lymphocyte count, platelet count and mononuclear cell count; the clinical pathological characteristics comprise a T staging characteristic, an N staging characteristic and a Ki-67 characteristic; determining a target biomarker related to pathologic complete remission pCR prediction based on the peripheral blood count data, and calculating an immune inflammation score according to the target biomarker; the clinical pathological features and the immune inflammation scores are processed through a trained XGB model, and a pCR prediction result of the to-be-detected object is obtained; and the XGB model is used for capturing a nonlinear relationship and an interaction effect among the data. According to the invention, the problem of poor detection timeliness can be avoided, and the pCR prediction accuracy is improved at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical data processing technology, and in particular to a method, device, equipment and medium for predicting pCR (progression complete response) in neoadjuvant chemotherapy for triple-negative breast cancer. Background Technology

[0002] In the field of life science research, triple-negative breast cancer (TNBC) is characterized by its high invasiveness and recurrence risk. As a malignant tumor with a high incidence rate, early diagnosis and precision treatment are crucial for improving patient prognosis. Neoadjuvant chemotherapy (NAC) is the standard treatment for TNBC, which can not only shrink the primary tumor size and reduce the tumor stage, but also achieve pathological complete response (pCR) through postoperative pathological evaluation. To avoid the toxic side effects of ineffective chemotherapy and the risk of delayed surgery, and to achieve individualized precision treatment for breast cancer, it is particularly important to study how to predict whether patients achieve the key efficacy indicator pCR after neoadjuvant chemotherapy.

[0003] Currently, related technologies rely on invasive methods such as surgical pathological examination or imaging monitoring to assess pathological complete remission (pCR) after NAC. However, this approach cannot provide predictive information before or in the early stages of treatment, resulting in poor predictive timeliness. Furthermore, because it relies solely on subjective morphological changes for assessment, it is rather one-sided and leads to low predictive accuracy. Summary of the Invention

[0004] The purpose of this application is to provide a method, device, equipment, and medium for predicting pCR (progression complete response) in neoadjuvant chemotherapy for triple-negative breast cancer, in order to solve the technical problems of poor timeliness and low accuracy of traditional methods.

[0005] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a method for predicting pCR (progression-free complete response) in triple-negative breast cancer achieved with neoadjuvant chemotherapy, including: Acquire clinicopathological features and peripheral blood count data of the subject before neoadjuvant chemotherapy; the peripheral blood count data includes: neutrophil count, lymphocyte count, platelet count, and monocyte count; the clinicopathological features include: T stage features, N stage features, and Ki-67 features; Based on the peripheral blood count data, target biomarkers associated with pCR prediction of pathological complete remission were identified, and an immune inflammation score was calculated based on the target biomarkers; The clinicopathological features and the immune inflammation score are processed by a trained XGB model to obtain the pCR prediction results of the subject to be tested; the pCR prediction results are used to characterize the probability of achieving pathological complete remission after neoadjuvant chemotherapy; the XGB model is used to capture the nonlinear relationships and interaction effects between the data.

[0006] Secondly, this application provides a predictive device for achieving pCR (progression complete response) with neoadjuvant chemotherapy in triple-negative breast cancer, the device comprising: The acquisition module is used to acquire the clinicopathological features of the subject to be tested and peripheral blood count data before neoadjuvant chemotherapy; the peripheral blood count data includes: neutrophil count, lymphocyte count, platelet count and monocyte count; the clinicopathological features include: T stage features, N stage features and Ki-67 features; The determination module is used to determine target biomarkers related to pathological complete remission (pCR) prediction based on the peripheral blood count data, and to calculate an immune inflammation score based on the target biomarkers; The prediction module is used to process the clinicopathological features and the immune inflammation score through a trained XGB model to obtain the pCR prediction result of the subject to be tested; the pCR prediction result is used to characterize the probability value of achieving pathological complete remission after neoadjuvant chemotherapy; the XGB model is used to capture the nonlinear relationship and interaction effect between the data.

[0007] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method for predicting pCR (progression complete response) in neoadjuvant chemotherapy for triple-negative breast cancer as described above.

[0008] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for predicting pCR (progression complete response) in neoadjuvant chemotherapy for triple-negative breast cancer as described above.

[0009] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a method, device, equipment, and medium for predicting pathological complete remission (pCR) after neoadjuvant chemotherapy in triple-negative breast cancer. The method includes: acquiring clinicopathological features and peripheral blood count data of the subject before neoadjuvant chemotherapy; peripheral blood count data includes neutrophil count, lymphocyte count, platelet count, and monocyte count; clinicopathological features include T stage features, N stage features, and Ki-67 features; determining target biomarkers related to pCR prediction based on peripheral blood count data, and calculating an immune inflammation score based on the target biomarkers; processing the clinicopathological features and immune inflammation score using a trained XGB model to obtain the pCR prediction result for the subject; the pCR prediction result is used to characterize the probability of achieving pathological complete remission after neoadjuvant chemotherapy; the XGB model is used to capture the nonlinear relationships and interaction effects between the data. Compared to existing technologies, this approach acquires peripheral blood count data (neutrophil, lymphocyte, platelet, and monocyte counts) and clinicopathological features (T stage, N stage, Ki-67) of the subject before NAC (necrotic inflammatory response) without invasive procedures. This solves the problems of traditional methods' inability to predict NAC in advance and poor timeliness, helping doctors to formulate or adjust treatment plans as early as possible, thus providing more comprehensive data for subsequent pCR (progressive response) prediction. Furthermore, it integrates two-dimensional data, including peripheral blood count data and tumor pathological features, rather than relying on subjective morphological changes, avoiding the limitations of traditional assessments that are one-sided. This allows for a comprehensive and accurate identification of target biomarkers, leading to more accurate immune inflammation scores calculated based on these biomarkers. The scores and clinicopathological features are then directly input into a trained XGB model. By capturing the nonlinear relationships and interaction effects between data, the accuracy of pCR prediction is significantly improved, providing more timely, comprehensive, and accurate decision support for NAC treatment in TNBC patients. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram of the application environment for a method to predict pCR (progression complete response) in neoadjuvant chemotherapy for triple-negative breast cancer according to an embodiment of this application. Figure 2 A flowchart illustrating a method for predicting pCR (progression complete response) in neoadjuvant chemotherapy for triple-negative breast cancer according to an embodiment of this application; Figure 3 A flowchart illustrating a method for generating an interactive analysis report according to an embodiment of this application; Figure 4 A schematic diagram of the functional modules of a predictive device for achieving pCR in neoadjuvant chemotherapy for triple-negative breast cancer, provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0013] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0014] As mentioned in the background section, triple-negative breast cancer (TNBC) is an aggressive subtype of breast cancer that is negative for estrogen receptor (ER), progesterone receptor (PR), and human epidermal growth factor receptor 2 (HER2), accounting for approximately 10%–15% of all breast cancer cases. TNBC is characterized by high proliferative activity, a high risk of early recurrence and metastasis, and poor long-term prognosis. Furthermore, it lacks clearly defined targeted therapies, thus chemotherapy remains the primary systemic treatment. Currently, related techniques rely on invasive methods such as surgical pathology examination or imaging monitoring to assess pathological complete remission (pCR) after non-invasive breast cancer (NAC). However, this approach cannot provide predictive information before or early in treatment, resulting in poor predictive timeliness. Moreover, because it relies solely on subjective morphological changes for assessment, it is rather one-sided and leads to low predictive accuracy.

[0015] To address the aforementioned shortcomings, this application provides a method for predicting pCR (progression-free complete response) in neoadjuvant chemotherapy for triple-negative breast cancer. Compared to existing technologies, this method acquires peripheral blood count data (neutrophil, lymphocyte, platelet, and monocyte counts) and clinicopathological features (T stage, N stage, Ki-67) of the subject before NAC (necrotic inflammatory response). This method requires no invasive procedures and solves the problems of traditional methods' inability to predict in advance and poor timeliness, helping physicians to formulate or adjust treatment plans earlier, thus providing more comprehensive data for subsequent pCR prediction. Furthermore, it integrates two-dimensional data, including peripheral blood count data and tumor pathological features, rather than relying on subjective morphological changes, avoiding the limitations of traditional, one-sided assessments. This allows for comprehensive and accurate identification of target biomarkers, leading to more accurate immune inflammation scores calculated based on these biomarkers. The scores and clinicopathological features are then directly input into a trained XGB model. By capturing nonlinear relationships and interaction effects between data, the accuracy of pCR prediction is significantly improved, providing more timely, comprehensive, and accurate decision support for NAC treatment in TNBC patients.

[0016] This application provides a method for predicting pCR (progression-free complete response) in neoadjuvant chemotherapy for triple-negative breast cancer, which can be applied to, for example... Figure 1 The illustrated application environment for a method to predict pCR (progression-free complete response) with neoadjuvant chemotherapy in triple-negative breast cancer is shown. This environment includes a terminal 102, a server 104, and a data storage system. The terminal 102 communicates with the server 104 via a network. The data storage system stores clinicopathological features and peripheral blood count data acquired by the server 104. The data storage system can be set up independently, integrated into the server 104, or located in the cloud or on another server. The terminal 102 can send the acquired clinicopathological features and peripheral blood count data to the server 104. After receiving the data, the server 104 performs biomarker identification, scoring, and model prediction processing to generate a pCR prediction result. Furthermore, in some embodiments, the method for predicting pCR with neoadjuvant chemotherapy in triple-negative breast cancer can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly perform biomarker identification, scoring, and model prediction processing to generate a pCR prediction result.

[0017] The terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster composed of multiple servers, or it can be a cloud server.

[0018] In one exemplary embodiment, such as Figure 2 As shown, a method for predicting pCR (progression-free complete response) in neoadjuvant chemotherapy for triple-negative breast cancer is provided. This method is executed by a computer device, specifically a terminal or server, or both. In this embodiment, the method is applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps S201 to S203. Wherein: Step S201: Obtain the clinicopathological features of the subject to be tested and peripheral blood count data before neoadjuvant chemotherapy; peripheral blood count data includes: neutrophil count, lymphocyte count, platelet count and monocyte count; clinicopathological features include: T stage features, N stage features and Ki-67 features.

[0019] It should be noted that the subjects to be tested can be patients with breast cancer or TNBC patients. The peripheral blood count data mentioned above refers to data collected within a preset time period before the start of neoadjuvant chemotherapy. This peripheral blood count data can include: neutrophil count, lymphocyte count, platelet count, and monocyte count. The preset time period can be customized according to actual needs, such as one week or one month. By setting a preset time period for peripheral blood count data, it is possible to ensure that the data collection time is within an appropriate range, avoiding collection time that is too early (e.g., one month before NAC) or too late (e.g., after NAC starts), which may lead to distortion of subsequent inflammation scores due to changes in the patient's inflammatory status.

[0020] The computer equipment can acquire data in real time from the electronic medical record system or external medical devices through the input interface, ensuring that data collection is completed within one week before the start of neoadjuvant chemotherapy to reflect the baseline immune inflammation status, thereby improving the accuracy of subsequent calculations.

[0021] As is understandable, T-stage (Tumor Stage) refers to the stage of a tumor, used to describe its size and extent of invasion. It is an indicator of the degree of local tumor progression. The higher the T-stage number, the larger and deeper the tumor. For example, T1 indicates a small tumor, while T3 or T4 indicates a large tumor or one that has invaded the chest wall or skin. T-stage characteristics are negatively correlated with pCR (progression-free response), meaning that the later the T-stage, the lower the probability of pCR. It is an important negative contributing variable in predictive models, reflecting the impact of local tumor burden on the efficacy of chemotherapy.

[0022] N-stage (NodeStage) refers to lymph node staging, used to describe whether a tumor has metastasized to regional lymph nodes and the extent of metastasis, assessing the risk of tumor spread. Higher stages indicate more severe lymph node metastasis; for example, N0 stage has no metastasis, N1 stage has a small amount of metastasis, and N2 or N3 stage has extensive metastasis. N-stage characteristics are negatively correlated with pCR (progressive complete response), meaning the later the N-stage, the lower the probability of pCR. As a negative contributing variable, it reflects the inhibitory effect of tumor spread on chemotherapy response.

[0023] Ki-67 is a protein marker reflecting the proliferative activity of tumor cells, and its index represents the proportion of cancer cells in the division and proliferation phase. Key significance: A higher Ki-67 index indicates more active tumor cell division and stronger invasiveness. Ki-67 characteristics are positively correlated with pCR (progression-free response), meaning a higher Ki-67 index leads to a higher probability of pCR, making it a positive contributing variable. This is because actively proliferating tumor cells are more sensitive to neoadjuvant chemotherapy (NAC) and more likely to achieve complete remission through chemotherapy.

[0024] Optionally, the above-mentioned clinicopathological features refer to the clinical features of the tumor of the subject to be tested. The clinicopathological features and peripheral blood count data of the subject to be tested can be collected in real time from the electronic medical record system through the input interface of the computer device, or obtained from the database or blockchain, or directly imported from external devices. This embodiment does not limit the acquisition method of clinicopathological features and peripheral blood count data.

[0025] In this embodiment, by acquiring the clinicopathological characteristics of the subject to be tested and peripheral blood count data before neoadjuvant chemotherapy, comprehensive and accurate data guidance information can be provided for subsequent pCR prediction, so as to facilitate accurate pCR prediction.

[0026] Step S202: Based on peripheral blood count data, identify target biomarkers associated with pathological complete remission (pCR) prediction, and calculate the immune inflammation score based on the target biomarkers.

[0027] It should be noted that the aforementioned target biomarkers, selected from peripheral blood count data, are specific indicators significantly associated with pathological complete remission (pCR). These biomarkers reflect the impact of the body's inflammatory-immune status on chemotherapy efficacy and are the core basis for calculating the immune-inflammatory score. The immune-inflammatory score (IIS) is a comprehensive value formed by quantifying and integrating the selected target biomarkers according to their association strength with pCR. It is used to centrally reflect the impact of the body's inflammatory-immune status before neoadjuvant chemotherapy on pathological complete remission.

[0028] In one embodiment, target biomarkers associated with pCR prediction in pathological complete remission are identified based on peripheral blood count data, including: Candidate biomarkers were calculated based on peripheral blood count data. Univariate logistic regression analysis was performed on each candidate biomarker to obtain significance P-values. Based on the significance P-values, the LASSO-logistic regression algorithm was used to select the optimal penalty parameter through multiple cross-validations, and target biomarkers associated with pathological complete remission pCR prediction were screened from all candidate biomarkers.

[0029] Specifically, after obtaining peripheral blood count data and clinicopathological features, candidate biomarkers can be calculated based on the peripheral blood count data, including basic indicators such as neutrophils, lymphocytes, platelets, and monocytes. Then, statistical methods are used to screen the candidate biomarkers, eliminating redundant indicators without significant association with pCR and retaining target biomarkers with independent predictive value. Finally, the weights of these target biomarkers are obtained, which can be determined by their association strength with pCR, and an immune inflammation score (IIS) is calculated based on these weights. This IIS transforms scattered blood indicators into a comprehensive quantitative value, not only retaining the core information of each biomarker reflecting the body's inflammation-immune status (e.g., high LMR indicates strong immune function, which is conducive to pCR; high NLR indicates a severe inflammatory response, which may inhibit therapeutic efficacy), but also eliminating the one-sidedness of single indicators through integration, thus more accurately reflecting the interaction between the host's immune status and the tumor, providing reliable inflammatory and immune dimension data support for subsequent pCR prediction. The aforementioned statistical methods can be, for example, the LASSO regression algorithm.

[0030] The aforementioned candidate biomarkers include target biomarkers, which include: neutrophil-to-lymphocyte ratio, lymphocyte-to-monocyte ratio, systemic immune inflammation index, and neutrophil-to-monocyte ratio. Based on peripheral blood count data, candidate biomarkers are calculated, including: Calculate the neutrophil-lymphocyte ratio based on neutrophil and lymphocyte counts; The lymphocyte-to-monocyte ratio was calculated based on the lymphocyte count and monocyte count. Systemic immune inflammation indices were determined based on the neutrophil-lymphocyte ratio and platelet count. The neutrophil-monocyte ratio was calculated based on neutrophil and monocyte counts.

[0031] In this embodiment, the neutrophil-lymphocyte ratio (NLR) reflects the body's inflammation and immune balance and can be obtained by dividing the neutrophil count by the lymphocyte count. The neutrophil-monocyte ratio (NMR) = neutrophil count / monocyte count. The platelet-lymphocyte ratio (PLR) = platelet count ÷ lymphocyte count. The lymphocyte-monocyte ratio (LMR) = lymphocyte count ÷ monocyte count, which reflects the balance between immune cell killing and suppression of tumors. The systemic immune inflammation index (SII) = platelet count. Neutrophil-to-lymphocyte ratio.

[0032] It's important to note that the lymphocyte-monocyte ratio (LMR) reflects immune killing capacity: lymphocytes are the core immune cells in anti-tumor therapy, while monocytes may promote tumor invasion; a higher LMR indicates that immune killing is greater than tumor invasion, making the immune system more sensitive to NAC and increasing the probability of pCR (pro-recovery). The NLR reflects the level of systemic inflammation: neutrophils represent pro-inflammatory responses (pro-tumor), while lymphocytes represent immunity (anti-tumor); a higher NLR indicates that inflammation is greater than immunity, potentially leading to poorer NAC efficacy and a lower probability of pCR. The systemic immune inflammation index integrates "inflammation + coagulation": SII = platelet count × NLR. Platelets promote tumor metastasis, while NLR reflects inflammation; a higher SII indicates a double increase in the risk of inflammation and metastasis, resulting in a lower probability of pCR after NAC. The neutrophil-monocyte ratio refines the "balance of inflammatory cell subtypes": neutrophils promote inflammation, while monocytes promote tumor invasion; a higher NMR indicates an imbalance in the ratio of these two pro-tumor cells, potentially leading to poorer NAC efficacy.

[0033] In this embodiment, after obtaining all candidate biomarkers, univariate logistic regression analysis is performed on each candidate biomarker individually. This involves analyzing the association between each candidate biomarker and "post-NAC pCR" separately, calculating the statistical results for each candidate biomarker, such as the significance P-value. The significance P-value is then compared to a preset threshold. When the preset threshold is 0.05, biomarkers with a significance P-value < 0.05 are retained. P < 0.05 means that the association between the biomarker and pCR is "not accidental," with a probability of over 95% that the association is real. This yields five relevant biomarkers, allowing for further optimization and narrowing down the scope. The LASSO-logistic regression algorithm and 10-fold cross-validation algorithm were executed. For the five biomarkers selected in the first step, the LASSO-logistic regression algorithm was performed using R language, and the optimal penalty parameter (lambda.min) was selected through 10-fold cross-validation. Finally, four target biomarkers with non-zero weight coefficients were retained. These target biomarkers could be NLR, LMR, SII, and NMR, and the corresponding weight coefficients for each target biomarker were calculated, for example, 2.454352383 for LMR and -0.146124084 for NMR. By employing the logistic regression algorithm, the two major technical problems of multicollinearity and overfitting can be solved, ensuring the stability and generalization ability of the immune inflammation score IIS.

[0034] Understandably, in logistic regression algorithms, the weight coefficient is a core quantitative indicator measuring the degree and direction of a biomarker's influence on IIS. It can be understood as the biomarker's contribution to IIS; the magnitude of the weight coefficient represents the degree of influence. The larger the absolute value of the weight coefficient, the more significant the biomarker's contribution to IIS. For example, the LMR coefficient is 2.454, with an absolute value much larger than other biomarkers, meaning that LMR has the strongest influence on IIS among the four biomarkers. The positive or negative sign of the weight coefficient represents the direction of influence. A positive weight coefficient means that the higher the value of the target biomarker, the higher the IIS. For example, a positive LMR coefficient means that the higher the LMR, the higher the IIS. Subsequent SHAP analysis further confirmed the positive association between LMR and other inflammatory markers and pCR. A negative weight coefficient means that the higher the biomarker value, the lower the IIS. For example, an NLR coefficient of -0.888 means that the higher the NLR, the lower the IIS, indirectly suggesting that a high NLR may reduce the probability of pCR.

[0035] The calculation of the immune inflammation score based on the target biomarker includes: Obtain the weighting coefficients corresponding to the target biomarkers; the weighting coefficients are used to characterize the positive or negative contribution of the target biomarkers to the prediction of pathological complete remission (pCR); calculate the immune inflammation score based on the target biomarkers and their corresponding weighting coefficients.

[0036] For example, taking the weighting coefficients for the lymphocyte-monocyte ratio as 2.454352383, the neutrophil-monocyte ratio as -0.146124084, the neutrophil-lymphocyte ratio as -0.88760856, and the systemic immune inflammation index as -0.002578236 as an example, the immune inflammation score (IIS) = (lymphocyte-monocyte ratio × 2.454352383) - (neutrophil-monocyte ratio × 0.146124084) - (neutrophil-lymphocyte ratio × 0.887608565) - (systemic immune inflammation index × 0.002578236).

[0037] This step involves acquiring multiple biomarkers and converting them into an immune inflammation score, which addresses the problem of limited prediction by a single biomarker in existing technologies. By employing a logistic regression algorithm, target biomarkers associated with pathological complete remission (pCR) prediction are accurately screened from all candidate biomarkers, facilitating subsequent pCR prediction based on effective biomarker data. Algorithm optimization ensures the reliability and clinical relevance of the score, further improving the accuracy of pCR prediction.

[0038] Step S203: The clinicopathological features and immune inflammation scores are processed by the trained XGB model to obtain the pCR prediction results of the subjects to be tested; the pCR prediction results are used to characterize the probability of achieving pathological complete remission after neoadjuvant chemotherapy; the XGB model is used to capture the nonlinear relationships and interaction effects between data.

[0039] Understandably, correlation analysis was conducted between the calculated immune inflammation score (IIS) and the clinicopathological characteristics (T stage, N stage, Ki-67) of TNBC patients to verify the clinical significance of IIS and ensure that IIS is not only a mathematical combination of data, but also truly reflects the biological characteristics of the tumor. For example, T stage or N stage represents tumor burden and metastasis risk, and Ki-67 represents tumor proliferative activity, laying the foundation for subsequent integration of these variables to predict pCR.

[0040] The XGB model described above is a tree-structured model. Clinical pathological features and immune inflammation scores are processed through the trained XGB model to obtain the pCR prediction results for the target subject, including: The average pCR probability of all samples in the training set is obtained as the initial baseline value. Based on the hyperparameters in the XGB model, the error of the preceding prediction results is corrected for each decision tree based on the threshold judgment of clinical pathological features and immune inflammation scores. The optimal correction method is found by searching the gradient direction of the error. At the same time, based on the depth of the decision tree and the node splitting rules, the nonlinear relationship and complex interaction effect between data are captured to output the corresponding correction value. The preceding prediction result of the first tree in all decision trees is used as the initial baseline value. The preceding prediction results of subsequent decision trees are the sum of the initial baseline value and the correction values ​​of all preceding trees. The correction values ​​of all decision trees are accumulated and superimposed to obtain the original output score. The original output score is converted into a probability value between 0 and 1 by the Sigmoid function to obtain the pCR prediction result of the object to be detected.

[0041] Specifically, the XGB model described above can use the average pCR probability of all samples in the training set as the initial baseline value, for example, an overall average pCR probability of 35%. This is the starting point for prediction, representing an initial inference about samples in the unknown target object. Then, relying on the hyperparameters optimized during model training, the initial baseline value is iteratively corrected through multiple decision trees. These hyperparameters can include tree depth, learning rate, and tree structure, ensuring fast and accurate mapping of input features to predicted probabilities without manual intervention. The model can be configured with threshold judgment rules, such as whether IIS > 4.2, whether the T-stage feature is T1 / T2, or whether Ki-67 > 30%. The first decision tree compares the error between the initial baseline value and the actual pCR prediction results in the training set based on the threshold judgment rules of the input features (clinical pathology features and immune inflammation scores). By calculating the gradient direction of the error—which indicates how to correct to minimize the error—the optimal correction method is found, gradually approaching the true result, thus outputting the first corrected value. The preceding prediction results are then updated based on this first corrected value. For example, if the actual pCR rate of a certain type of sample is 60%, and the initial baseline value of 35% has an error of 25%, by calculating the gradient direction of the error, the first correction value is output as +20%. At this time, the previous prediction result is updated to 35% + 20% = 55%. Herein lies the actual pCR annotation result in the training set.

[0042] Starting with the second tree, each tree uses the cumulative result of all previous decision trees (e.g., 55%) as the new preceding prediction result. The residual error between this new prediction and the actual prediction result is recalculated; for example, the error between 55% and the actual 60% is 5%. Based on more granular feature thresholds (e.g., "when IIS > 4.2 and T stage is T1, samples with Ki-67 > 30% have a larger error"), nonlinear relationships and complex interaction effects between data are captured, resulting in a targeted correction value (e.g., +5%). This correction is then summed to obtain the corresponding preceding prediction result of 60%. This process is repeated for all decision trees. After all decision trees have completed iterations, all correction values ​​are accumulated and summed to the initial baseline value to obtain a raw output score, such as the score corresponding to 60% in the example above. Finally, the sigmoid function is used to convert this raw output score into a probability value between 0 and 1, which is the predicted result for the subject to achieve pCR after neoadjuvant chemotherapy. Among them, nonlinear relationships, such as IIS, have a weaker effect on pCR in the 3-4 interval than in the 4-5 interval. Complex interaction effects, such as high IIS, only significantly promote pCR when N is N0, and the effect is weakened in N1.

[0043] In this step, the error is continuously corrected through a gradient boosting mechanism, while the decision tree structure is used to accurately capture the nonlinear relationships and complex interaction effects between features. Ultimately, the prediction results not only better match the patterns of clinical data, but also have quantitative interpretability, providing a reliable basis for personalized treatment decisions.

[0044] The above XGB model is constructed through the following steps: The process involves: acquiring historical data of the sample objects; annotating the historical data with pCR annotation results; dividing the historical data into training and validation sets according to a preset ratio; inputting the training set into the initial XGB model for training; optimizing the hyperparameters of the initial XGB model through cross-validation based on the pCR annotation results to obtain the trained model; hyperparameters include: tree depth, learning rate, and regularization parameter; tree depth is used to control the complexity of a single tree, the learning rate is used to control the contribution weight of each tree, and the regularization parameter is used to suppress overfitting; inputting the validation set into the trained model for validation to obtain the XGB model; the XGB model is used to capture the nonlinear relationships and interaction effects between data points.

[0045] It should be noted that during the XGB model construction phase, historical data of the sample subjects can be obtained. These sample subjects can be patients with breast cancer. For example, data from 700 patients can be used as the training set, and data from 300 patients can be used as the internal validation set for model construction and evaluation. First, the core parameters are optimized using the training set data from 700 patients. By inputting the training set into the initial XGB model for training, hyperparameters such as the number and depth of decision trees and the learning rate are iteratively adjusted. This allows the model to learn the combined influence of clinicopathological features (T stage, N stage, Ki-67) and immune inflammation score (IIS) on pCR from the training data. At the same time, regularization mechanisms are used to avoid overfitting, such as limiting the complexity of the trees to prevent the model from losing its generalization ability, thus obtaining the trained model. Subsequently, computer equipment can use the validation set data from 300 patients to perform multi-dimensional evaluations of the trained model, including discriminant capability validation, calibration validation, and clinical net benefit evaluation.

[0046] For discriminative ability validation, the area under the ROC curve (AUC) of the XGB model can be calculated. This metric reflects the model's ability to distinguish between patients who "achieved pCR" and those who "did not achieve pCR." A metric closer to 1 indicates more accurate discrimination; for example, an AUC > 0.95 indicates the model can effectively identify high / low pCR probability populations. For calibration validation, the model's predicted probability can be compared with the actual pCR incidence using calibration curves. For instance, if a 70% probability is predicted, approximately 70% of patients actually achieve pCR, ensuring consistency between the predicted value and the actual result and avoiding the risk of overestimation or underestimation. For clinical net benefit assessment, decision curve analysis (DCA) can quantify the model's clinical value at different thresholds. For example, when physicians use "predicted pCR probability > 50%" as the basis for treatment decisions, can the model reduce unnecessary interventions or missed diagnoses, ultimately proving that the model's practical application brings positive clinical benefits?

[0047] In addition, the aforementioned computer equipment can also be equipped with the ability to process multimodal data integration. In addition to core clinical pathological features and immune inflammation scores, it can also be compatible with real-world data such as patient age, tumor size, and treatment plans. Through preprocessing steps such as unifying data formats and handling missing values, it ensures that the model can still run stably in complex clinical scenarios, thereby improving its applicability and promotion value in actual diagnosis and treatment.

[0048] This application provides a method for predicting pCR (progression-free complete response) in neoadjuvant chemotherapy for triple-negative breast cancer. Compared with existing technologies, this method obtains peripheral blood count data (neutrophil, lymphocyte, platelet, and monocyte counts) and clinicopathological features (T stage, N stage, Ki-67) of the subject before NAC (negative acute cerebral infarction) without invasive procedures. This solves the problems of traditional methods being unable to predict in advance and having poor timeliness, helping doctors to formulate or adjust treatment plans as early as possible, thus providing more comprehensive data for subsequent pCR prediction. Furthermore, it can integrate two-dimensional data such as peripheral blood count data and tumor pathological features, rather than relying on subjective morphological changes, avoiding the shortcomings of traditional assessments that are one-sided. This allows for the comprehensive and accurate identification of target biomarkers, resulting in more accurate immune inflammation scores calculated based on these target biomarkers. The scores and clinicopathological features are then directly input into a trained XGB model. By capturing the nonlinear relationships and interaction effects between data, the accuracy of pCR prediction is significantly improved, providing more timely, comprehensive, and accurate decision support for NAC treatment in TNBC patients.

[0049] In one embodiment, after obtaining the pCR prediction results, a specific implementation method for obtaining an interactive analysis report is also provided; please refer to [link to relevant documentation]. Figure 3 As shown, it includes the following steps: Step S301: Process the pCR prediction results using the SHAP interpretation framework and calculate the SHAP value of each input variable; each input variable includes clinicopathological features and immune inflammation score.

[0050] Step S302: Call the chart generation program to perform correlation mapping processing on the SHAP value and the corresponding input variable, generate interpretation results and display them; the interpretation results include global interpretation charts and local interpretation charts; the global interpretation chart is used to reflect the general trend of the influence of each input variable on the pCR prediction result; the local interpretation chart is used to show the variable influence of a single object to be detected.

[0051] Step S303: Based on the interpretation results and pCR prediction results, obtain an interactive analysis report.

[0052] Specifically, after obtaining the pCR prediction results, the computer equipment uses the SHAP interpretation framework to process the XGB model output. This processing involves calculating the SHAP value for each input variable to quantify the contribution of the immune inflammation score, T stage, N stage, and Ki-67 to the prediction results, and generating global and local interpretation charts. The SHAP (SHapley Additive exPlanations) framework is based on the Shapley value principle in game theory, quantifying the specific impact of each input variable (T stage feature, N stage feature, Ki-67, and immune inflammation score IIS) on the pCR prediction results. For each target object, the computer device calculates the SHAP value of each variable by decomposing the predicted probability output by the model. The SHAP value includes positive and negative values. A positive value indicates that the variable increases the predicted probability. For example, a positive SHAP value for high Ki-67 indicates that it increases the pCR probability. A negative value indicates that the variable inhibits the predicted probability. For example, a negative SHAP value for a later T stage indicates that it decreases the pCR probability. The absolute value reflects the strength of the influence. For example, if the absolute value of the SHAP of IIS is greater than that of a certain stage feature, it indicates that its influence on the prediction result of the target object is more significant.

[0053] After obtaining the SHAP values ​​for each input iteration, a chart generation program can be used to associate the SHAP values ​​with the input variables, visually presenting the influence patterns and generating global or local explanatory charts. Global explanatory charts can be SHAP summary charts or beehive plots. The horizontal axis of the chart displays the distribution of SHAP values ​​for all variables, while the vertical axis reflects the relationship between variable values ​​(such as Ki-67 levels) and SHAP values. For example, the chart can clearly show that "SHAP values ​​are mostly positive and concentrated in the high range when Ki-67 > 30%" (generally improving pCR) and "SHAP values ​​for T4 stages are all negative" (generally suppressing pCR), helping doctors understand the general trend of the variables' influence on the overall population. Local explanatory charts can be waterfall plots or force-directed plots, displaying the SHAP values ​​of each variable in order of influence intensity for a single patient. For example, a waterfall plot for a patient can show that "IIS = 5.2 (SHAP = +0.3) is the main factor improving pCR prediction, while N1 stage (SHAP = -0.15) slightly reduces the prediction probability," allowing doctors to clearly identify the key driving factors for the individual patient's prediction outcome.

[0054] After obtaining the SHAP interpretation results, the output interface can be used to control the display of the SHAP interpretation results and support the visual control of medical devices. For example, the SHAP interpretation results can be integrated with the pCR prediction probability to generate an interactive report, allowing doctors to switch between viewing global patterns and individual details through the interface. Dynamic interaction is also supported to help doctors verify the model conclusions with clinical experience, so as to realize personalized treatment decisions and improve the interpretability and clinical trust of the model.

[0055] Optionally, the computer device can generate control signals based on the pCR prediction probability and interpretation results, and send these control signals to external devices through an output interface. These control signals could be, for example, controlling the display screen to output probability values, SHAP interpretation charts, or triggering alarms to adjust treatment plans, such as optimizing surgical timing or adjusting chemotherapy regimens, to ensure the method is applied in the clinical decision-making process. This step can also include a real-time feedback mechanism, such as automatically generating an alarm signal to control external medical devices if the predicted probability falls below a threshold, and supporting data logging for subsequent model iteration and multi-center validation, thereby promoting the application of precision medicine in triple-negative breast cancer care. The external medical devices could be, for example, a notification system or an electronic medical record update module.

[0056] The technical solution of this application calculates biomarkers such as NLR, LMR, SII, and NMR by collecting peripheral blood count data, constructs an immune inflammation score (IIS) using the LASSO-logistic regression algorithm, and integrates clinical pathological variables such as T stage, N stage, and Ki-67 as input features. It uses a trained XGB model to generate pCR probability values ​​and applies the SHAP interpretation framework to quantify variable contributions, thereby improving prediction accuracy and interpretability and facilitating intuitive acquisition of the patient's pathological condition.

[0057] For example, 6967 breast cancer patients were identified, and 1000 TNBC patients were selected. These patients were then randomly divided into a training set of 700 cases and a validation set of 300 cases at a ratio of 7:3. Peripheral blood count data and clinicopathological data of these patients were obtained one week before the start of NAC. Peripheral blood count data included neutrophils, lymphocytes, platelets, and monocytes. Clinicopathological data may include patient age, BMI, family history, menstrual characteristics, T stage, N stage, histological subtype, grade, HER2, Ki-67, p53, and androgen receptor, etc. Six candidate biomarkers were calculated based on the peripheral blood count data. A univariate logistic regression algorithm was used for screening to obtain a significant P-value. P-values ​​<0.05 were processed to obtain five relevant indicators. Then, LASSO (10-fold cross-validation, lambda.min was used to retain four target biomarkers, including NLR, LMR, SII, and NMR. Then, the weighting coefficients of the four target biomarkers were obtained, and the immune inflammation score was calculated. The immune inflammation score IIS = (LMR × 2.454352383) - (NMR × 0.146124084) - (NLR × 0.887608565) - (SII × 0.002578236).

[0058] Multivariate logistic regression was used to confirm T-staging, N-staging, Ki-67, and IIS as independent predictors. Based on these independent predictors and the immune inflammation score, eight machine learning algorithms (including XGB) were used to train the model. A 10-fold cross-validation algorithm was employed to obtain the trained model. The trained model was evaluated using a validation set to determine the evaluation ROC (AUC>0.95), calibration curve, and DCA. The model with the best performance index was selected as the final XGB model, which can then be used to predict pCR in the subjects under test. Furthermore, the SHAP framework was applied to generate global or local interpretation charts based on the pCR prediction results, showing that inflammatory immune markers and Ki-67 contribute positively to pCR prediction, while T / N contributes negatively. Finally, a web calculator was developed, supporting input of patient data and outputting pCR probabilities. This approach ensures non-invasiveness, efficient data processing, applicability to real-world data, and support for multi-center validation.

[0059] The technical solution presented in this application achieves non-invasive, high-accuracy, and well-calibrated predictive effects, improving the model's net clinical benefit. It supports early identification of high-response patients, optimizes surgical timing, and facilitates personalized treatment decisions, promoting the application of precision medicine in TNBC care. Furthermore, SHAP visualization enhances model transparency and trustworthiness. In the construction and validation of the XGB model, the baseline features of the training set (700 cases) and validation set (300 cases) were balanced, ensuring data comparability. Analysis showed a significant correlation between the immune inflammation score (IIS) and clinicopathological features, with higher IIS in early-stage (T1 / T2, N0) and high Ki-67 patients, suggesting a link between inflammatory immune status and tumor biological behavior. The constructed XGB model exhibits excellent performance, with a high area under the ROC curve (AUC>0.95) indicating strong discriminative ability. Consistent calibration curves demonstrate good agreement between predicted probabilities and actual results. Decision curve analysis (DCA) shows high net clinical benefit, and the visualization of the XGB model's confusion matrix further validates its accurate classification ability. Furthermore, SHAP analysis clearly showed that the inflammatory immune marker (IIS) contributed the most to pCR prediction and had a positive impact, while T / N staging had a negative impact, confirming the value of multi-dimensional integration. This approach is superior to single-indicator prediction and has significant clinical benefits: it can reduce the toxicity caused by ineffective chemotherapy, help doctors determine the timing of surgery in a timely and accurate manner, and because it only relies on peripheral blood and routine pathology data, it supports application in resource-limited environments and can rapidly improve the clinical management efficiency of triple-negative breast cancer (TNBC).

[0060] Based on the same inventive concept, this application also provides a predictive device for achieving pCR (progression complete response) with neoadjuvant chemotherapy for triple-negative breast cancer as described above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of the predictive device for achieving pCR with neoadjuvant chemotherapy for triple-negative breast cancer provided below can be found in the limitations of the predictive method for achieving pCR with neoadjuvant chemotherapy for triple-negative breast cancer described above, and will not be repeated here.

[0061] In one exemplary embodiment, such as Figure 4 As shown, a predictive device for achieving pCR (progression complete response) with neoadjuvant chemotherapy in triple-negative breast cancer is provided. The device includes: The acquisition module 510 is used to acquire the clinicopathological features of the subject to be tested and peripheral blood count data before neoadjuvant chemotherapy; the peripheral blood count data includes: neutrophil count, lymphocyte count, platelet count and monocyte count; the clinicopathological features include: T stage features, N stage features and Ki-67 features; The determination module 520 is used to identify target biomarkers associated with pathological complete remission pCR prediction based on peripheral blood count data, and to calculate an immune inflammation score based on the target biomarkers; The prediction module 530 is used to process clinicopathological features and immune inflammation scores through a trained XGB model to obtain pCR prediction results for the subjects to be tested; the pCR prediction results are used to characterize the probability of achieving pathological complete remission after neoadjuvant chemotherapy; the XGB model is used to capture the nonlinear relationships and interaction effects between the data.

[0062] As an optional implementation, the determining module 520 is specifically used for: Candidate biomarkers were calculated based on peripheral blood count data; Univariate logistic regression analysis was performed on each candidate biomarker to obtain the significance P-value; Based on the significance P-value, the LASSO-logistic regression algorithm was used to select the optimal penalty parameter through multiple cross-validations, and target biomarkers associated with the prediction of pathological complete remission pCR were screened from all candidate biomarkers.

[0063] As an optional implementation, the determining module 520 is further configured to: Calculate the neutrophil-lymphocyte ratio based on neutrophil and lymphocyte counts; The lymphocyte-to-monocyte ratio was calculated based on the lymphocyte count and monocyte count. Systemic immune inflammation indices were determined based on the neutrophil-lymphocyte ratio and platelet count. The neutrophil-monocyte ratio was calculated based on neutrophil and monocyte counts.

[0064] As an optional implementation, the determining module 520 is further configured to: Obtain the weighting coefficients corresponding to the target biomarkers; the weighting coefficients are used to characterize the positive or negative contribution of the target biomarkers to the prediction of complete pathological remission pCR. An immune inflammation score is calculated based on the target biomarker and its corresponding weighting coefficient.

[0065] As an optional implementation, the prediction module 530 is specifically used for: The average pCR probability of all samples in the training set is used as the initial baseline value. Based on the hyperparameters in the XGB model, each decision tree is judged based on thresholds of clinicopathological features and immune inflammation scores to correct the error of the preceding prediction results. The optimal correction method is found by searching the gradient direction of the error. At the same time, based on the depth of the decision tree and the node splitting rules, the nonlinear relationship and complex interaction effects between data are captured to output the corresponding correction value. The preceding prediction result of the first tree in all decision trees is the initial baseline value, and the preceding prediction results of subsequent decision trees other than the first tree are the sum of the initial baseline value and the correction values ​​of all preceding trees. The correction values ​​of all decision trees are accumulated and superimposed to obtain the original output score. The original output score is then converted into a probability value between 0 and 1 using the Sigmoid function to obtain the pCR prediction result of the object to be detected.

[0066] As an optional implementation, the above-described apparatus is further used for: The pCR prediction results were processed using the SHAP interpretation framework to calculate the SHAP values ​​of each input variable, including clinicopathological features and immune inflammation scores. The chart generation program is invoked to perform correlation mapping between SHAP values ​​and corresponding input variables, generate interpretation results, and display them. The interpretation results include global interpretation charts and local interpretation charts. The global interpretation chart is used to reflect the general trend of the influence of each input variable on the pCR prediction results. The local interpretation chart is used to show the variable influence of a single object to be detected. An interactive analysis report is generated based on the interpretation results and pCR prediction results.

[0067] As an optional implementation, the XGB model is constructed through the following steps: Obtain historical data of the sample objects; the historical data contains pCR annotation results. Historical data is divided into training and validation sets according to a preset ratio; The training set is input into the initial XGB model for training. The hyperparameters in the initial XGB model are optimized through cross-validation to obtain the trained model. The hyperparameters include: tree depth, learning rate, and regularization parameter. The tree depth is used to control the complexity of a single tree, the learning rate is used to control the contribution weight of each tree, and the regularization parameter is used to suppress overfitting. The validation set is input into the trained model for validation, resulting in the XGB model.

[0068] The device for predicting pCR (progression-free complete response) in neoadjuvant chemotherapy for triple-negative breast cancer provided in this application embodiment acquires peripheral blood count data (neutrophil, lymphocyte, platelet, and monocyte counts) and clinicopathological features (T stage, N stage, Ki-67) of the subject before NAC (necrotic angina) without invasive operation. This solves the problems of traditional methods being unable to predict in advance and having poor timeliness, helping doctors to formulate or adjust treatment plans as early as possible, thus providing more comprehensive data for subsequent pCR prediction. It can also integrate two-dimensional data such as peripheral blood count data and tumor pathological features, rather than relying on subjective morphological changes, avoiding the shortcomings of traditional assessments that are one-sided, thereby comprehensively and accurately identifying target biomarkers, and making the immune inflammation score calculated based on the target biomarker more accurate. The score and clinicopathological features are then directly input into a trained XGB model, and by capturing the nonlinear relationship and interaction effect between data, the accuracy of pCR prediction is greatly improved, providing more timely, comprehensive, and accurate decision support for NAC treatment of TNBC patients.

[0069] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 5 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores video tag processing data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When executed by the processor, the computer program implements a method for predicting pCR (progression complete response) in neoadjuvant chemotherapy for triple-negative breast cancer.

[0070] The processor in the aforementioned computer device can execute Python / R code, and the memory can store pre-trained AGB model parameters and the formula for calculating the immune inflammation score (IIS). Input or output interfaces are used to connect to the electronic medical record system and the display screen. First, peripheral blood data from the week prior to NAC is collected via the input interface, and target biomarkers such as NLR, LMR, SII, and NMR are calculated. The LASSO coefficient is then used to construct the immune inflammation score, which can be expressed by the following formula: IIS = (LMR × 2.454352383) - (NMR × 0.14612) 4084)-(NLR×0.887608565)-(SII×0.002578236); Then, IIS is integrated with T-staging, N-staging, and Ki-67 as inputs, and the processor runs the XGB model (10x cross-validation optimization) to generate probabilities; Finally, the SHAP module is used to output the interpretation results, such as the positive contribution of inflammatory indicators. The output interface controls external devices, which can generate control signals and respond to the control signals to trigger alarms; The implementation is based on training set (700 cases) and validation set (300 cases) data to ensure applicability to the hospital environment, support for multi-user access and data privacy protection.

[0071] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0072] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0073] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0074] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0075] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0076] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0077] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0078] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0079] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for predicting pCR (progression-free complete response) in neoadjuvant chemotherapy for triple-negative breast cancer, characterized in that, The methods for predicting pCR (progression complete response) with neoadjuvant chemotherapy in triple-negative breast cancer include: Acquire clinicopathological features and peripheral blood count data of the subject before neoadjuvant chemotherapy; the peripheral blood count data includes: neutrophil count, lymphocyte count, platelet count, and monocyte count; the clinicopathological features include: T stage features, N stage features, and Ki-67 features; Based on the peripheral blood count data, target biomarkers associated with pCR prediction of pathological complete remission were identified, and an immune inflammation score was calculated based on the target biomarkers; The clinicopathological features and the immune inflammation score are processed by a trained XGB model to obtain the pCR prediction results of the subject to be tested; the pCR prediction results are used to characterize the probability of achieving pathological complete remission after neoadjuvant chemotherapy; the XGB model is used to capture the nonlinear relationships and interaction effects between the data.

2. The method for predicting pCR (progression complete response) in neoadjuvant chemotherapy for triple-negative breast cancer according to claim 1, characterized in that, Based on the peripheral blood count data, target biomarkers associated with pCR prediction in pathological complete remission were identified, including: Candidate biomarkers are calculated based on the peripheral blood count data; Univariate logistic regression analysis was performed on each candidate biomarker to obtain the significance P-value; Based on the significance P-value, the LASSO-logistic regression algorithm was used to select the optimal penalty parameter through multiple cross-validations, and target biomarkers associated with the prediction of pathological complete remission pCR were screened from all candidate biomarkers.

3. The method for predicting pCR (progression complete response) in neoadjuvant chemotherapy for triple-negative breast cancer according to claim 2, characterized in that, The candidate biomarkers include target biomarkers, which include: neutrophil-lymphocyte ratio, lymphocyte-monocyte ratio, systemic immune inflammation index, and neutrophil-monocyte ratio. Based on the peripheral blood count data, candidate biomarkers are calculated, including: The neutrophil-lymphocyte ratio is calculated based on the neutrophil and lymphocyte counts. The lymphocyte-to-monocyte ratio is calculated based on the lymphocyte count and monocyte count. The systemic immune inflammation index is determined based on the neutrophil-lymphocyte ratio and the platelet count; The neutrophil-monocyte ratio is calculated based on the neutrophil count and the monocyte count.

4. The method for predicting pCR (progression complete response) in neoadjuvant chemotherapy for triple-negative breast cancer according to claim 1, characterized in that, Calculate an immune inflammation score based on the target biomarker, including: Obtain the weighting coefficients corresponding to the target biomarkers; the weighting coefficients are used to characterize the positive or negative contribution of the target biomarkers to the prediction of pathological complete remission pCR; The immune inflammation score is calculated based on the target biomarker and its corresponding weighting coefficient.

5. The method for predicting pCR (progression complete response) in neoadjuvant chemotherapy for triple-negative breast cancer according to claim 1, characterized in that, The XGB model is a tree model; The clinicopathological features and the immune inflammation score are processed using a trained XGB model to obtain the pCR prediction results for the subject to be tested, including: The average pCR probability of all samples in the training set is used as the initial baseline value. Based on the hyperparameters in the XGB model, each decision tree is judged according to the threshold of the clinicopathological features and the immune inflammation score to correct the error of the preceding prediction results. The optimal correction method is found by the gradient direction of the error. At the same time, based on the depth of the decision tree and the node splitting rules, the nonlinear relationship and complex interaction effect between data are captured to output the corresponding correction value. The preceding prediction result of the first tree in all decision trees is the initial baseline value, and the preceding prediction results of subsequent decision trees other than the first tree are the sum of the initial baseline value and the correction values ​​of all preceding trees. The correction values ​​of all decision trees are accumulated and superimposed to obtain the original output score. The original output score is then converted into a probability value between 0 and 1 using the Sigmoid function to obtain the pCR prediction result of the object to be detected.

6. The method for predicting pCR (progression complete response) in neoadjuvant chemotherapy for triple-negative breast cancer according to claim 1, characterized in that, After obtaining the pCR prediction results, the method further includes: The pCR prediction results are processed using the SHAP interpretation framework to calculate the SHAP value of each input variable; the input variables include clinicopathological features and immune inflammation scores. The chart generation program is invoked to perform correlation mapping processing on the SHAP value and the corresponding input variables, generate interpretation results, and display them; the interpretation results include a global interpretation chart and a local interpretation chart; the global interpretation chart is used to reflect the general trend of the influence of each input variable on the pCR prediction result; the local interpretation chart is used to show the variable influence of a single object to be detected; An interactive analysis report is obtained based on the interpretation results and the pCR prediction results.

7. The method for predicting pCR (progression complete response) in neoadjuvant chemotherapy for triple-negative breast cancer according to claim 1, characterized in that, The XGB model is constructed through the following steps: Obtain historical data of the sample object; the historical data contains pCR annotation results. The historical data is divided into a training set and a validation set according to a preset ratio; The training set is input into the initial XGB model for training. Based on the pCR annotation results, the hyperparameters in the initial XGB model are optimized through cross-validation to obtain the trained model. The hyperparameters include: tree depth, learning rate, and regularization parameter. The tree depth is used to control the complexity of a single tree, the learning rate is used to control the contribution weight of each tree, and the regularization parameter is used to suppress overfitting. The validation set is input into the trained model for validation to obtain the XGB model.

8. A device for predicting pCR (progression complete response) in neoadjuvant chemotherapy for triple-negative breast cancer, characterized in that, The predictive device for achieving pCR with neoadjuvant chemotherapy in triple-negative breast cancer includes: The acquisition module is used to acquire the clinicopathological features of the subject to be tested and peripheral blood count data before neoadjuvant chemotherapy; the peripheral blood count data includes: neutrophil count, lymphocyte count, platelet count and monocyte count; the clinicopathological features include: T stage features, N stage features and Ki-67 features; The determination module is used to determine target biomarkers related to pathological complete remission (pCR) prediction based on the peripheral blood count data, and to calculate an immune inflammation score based on the target biomarkers; The prediction module is used to process the clinicopathological features and the immune inflammation score through a trained XGB model to obtain the pCR prediction result of the subject to be tested; the pCR prediction result is used to characterize the probability value of achieving pathological complete remission after neoadjuvant chemotherapy; the XGB model is used to capture the nonlinear relationship and interaction effect between the data.

9. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the steps of the method for predicting pCR (progression complete response) in neoadjuvant chemotherapy for triple-negative breast cancer as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the method for predicting pCR (progression complete response) in neoadjuvant chemotherapy for triple-negative breast cancer as described in any one of claims 1-7.