A Predictive Model for Adjuvant Therapy in Breast Cancer and Its Application

By constructing a cross-platform gene pair prediction model, the problem of cross-platform data inconsistency in the response to neoadjuvant chemotherapy in triple-negative breast cancer was solved, achieving high-precision and high-generalization prediction and providing scientific guidance for personalized treatment plans.

CN121148480BActive Publication Date: 2026-04-03WEIFANG MEDICAL UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies suffer from cross-platform data inconsistency when predicting neoadjuvant chemotherapy response in triple-negative breast cancer patients, resulting in the inability of the model to generalize. Furthermore, traditional markers are difficult to validate across platforms, leading to low diagnostic efficacy and an inability to provide personalized treatment guidance.

Method used

A cross-platform gene pair prediction model was constructed. Effective gene pairs were screened using bulk, microarray, and single-cell sequencing data. Machine learning methods such as RFE and LASSO were combined, and a gene pair feature-guided PCR hierarchical prediction system was constructed through a multi-layer artificial neural network to achieve cross-platform data integration and accurate prediction.

Benefits of technology

It achieves high-precision and high-generalization prediction of response to neoadjuvant therapy in breast cancer, with an AUC of 0.914. It has high sensitivity and specificity, and provides scientific guidance for personalized treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121148480B_ABST
    Figure CN121148480B_ABST
Patent Text Reader

Abstract

This invention discloses a predictive model for adjuvant therapy of breast cancer and its application, belonging to the field of breast cancer technology. This invention utilizes gene-pair collaborative characterization to construct a Gene-Pair Classification System (GPCS), overcoming the shortcomings of traditional models based on differentially expressed genes (DEGs), which suffer from noise interference from single-gene expression abundance, limitations imposed by detection platforms, significantly restricted classification performance, and low predictive accuracy. It eliminates the need for standardization, using the GPCS model to overcome single-gene noise interference, accurately capture inter-gene interactions, and pioneers a collaborative characterization system based on the relative expression levels of gene pairs, compressing feature redundancy and significantly improving classification accuracy. This invention is compatible with various transcriptome chips, bulk transcriptomes, and single-cell transcriptome platforms, exhibiting high diagnostic efficacy and achieving a key breakthrough in predictive response models for triple-negative breast cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of breast cancer technology, specifically to a predictive model for adjuvant therapy of breast cancer and its application. Background Technology

[0002] Breast cancer is the most common cancer worldwide, accounting for 31% of cancers in women. Neoadjuvant chemotherapy (NAC) is the standard treatment for patients with inoperable or extensively resectable breast cancer. Treatment outcomes are categorized as pathologic complete response (PCR) and non-pathologic complete response (non-PCR). Pathologic complete response is defined as the absence of residual invasive cancer cells (ypT0 / is, ypN0) in the histopathological specimens of the breast and axillary lymph nodes after NAC, and it is associated with a good prognosis. Triple-negative breast cancer (TNBC) is a type of breast cancer where immunohistochemical staining of the cancer tissue shows negative results for estrogen receptor (ER), progesterone receptor (PR), and the proto-oncogene Her-2. This type of breast cancer accounts for 10.0%–20.8% of all pathological types of breast cancer and has unique biological behavior and clinicopathological features, with a poorer prognosis than other types. For triple-negative breast cancer, although the success rate of PCR (probability of response to neoadjuvant chemotherapy) has been increasing with the advancement of medical technology, significant heterogeneity in treatment response still exists among individuals. Certain therapies are ineffective for some patients, who not only suffer from the toxic side effects of ineffective treatment but also miss the optimal treatment window due to the inability to adjust treatment methods in a timely manner. Therefore, clinicians need to develop a system that can predict the response to NAC (negative breast cancer) treatment while simultaneously guiding clinicians' decision-making.

[0003] To develop systems for predicting NAC (Neoadjuvant chemotherapy), serological methods have been attempted, but their high cost and limited validation in large populations have prevented widespread clinical application. Recently, with the development of machine learning and neural networks, these technologies have shown great promise in predicting NAC responses. The patent "CN112542247A Method and System for Predicting the Probability of Pathological Complete Remission after Neoadjuvant Chemotherapy in Breast Cancer" points out that models such as Oncotype DX® have low diagnostic efficacy (AUC generally <0.75) in predicting pathological complete remission after neoadjuvant chemotherapy (NAC) and are only applicable to hormone receptor-positive (HR+) patients, ineffective for HER2+ or triple-negative breast cancer. Furthermore, current biomarker exploration studies mostly use bulk transcriptomics, microarray transcriptomics, or single-cell data alone for differential analysis to obtain biomarkers. However, the inconsistency in platform data distribution makes cross-platform validation of biomarkers difficult, neglecting the diversity of biomarker development techniques, model design, and input structures. Therefore, overcoming the differences between different platforms, developing cross-platform applicable models, and improving diagnostic efficacy are key to developing predictive response models for triple-negative breast cancer. Summary of the Invention

[0004] To address the aforementioned limitations of existing technologies, the present invention aims to provide a predictive model for adjuvant therapy in breast cancer and its application. Gene pairs describe the interactions between genes and the complexity of gene expression, reducing interference from external factors and providing more accurate results. The cross-platform gene pair prediction response model for triple-negative breast cancer designed in this invention is specifically tailored for predicting the response to neoadjuvant therapy in TNBC patients. First, gene pairs are constructed using bulk, microarray, and single-cell sequencing data, and effective gene pairs are screened through statistical tests. Then, combining machine learning methods such as RFE (Recursive Feature Elimination) and LASSO (Least Absolute Shrinkage and Selection Operator) with multi-layer artificial neural network training, a PCR hierarchical prediction system guided by gene pair features is constructed. The assignment method accurately indicates whether PCR has been achieved to advance the implementation of subsequent treatment plans.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] In a first aspect, the present invention provides a gene pair set for evaluating the prognostic effects of triple-negative breast cancer, said gene pair set comprising 20 core gene pairs, as follows:

[0007] ARG2_FANCA, BAMBI_EZH2, CASP1_GPSM3, CD36_CXCL13, COL11A1_PLOD2, COL5A1_TWF1, CPA3_IRF5, CSF1R_CXCL13, CSF1R_TYMS, CXCL13_ FAS, CXCL13_ND2, EZH2_HK2, EZH2_PARP4, FADD_ITGA6, FANCA_RORC, FANCA_THBD, FSTL3_KRAS, IL1B_ITGA6, IRF4_ROR2, RPS6KB1_SFRP4.

[0008] Furthermore, the GeneID numbers of the genes in the core gene pair in NCBI are as follows: ARG2: 384, FANCA: 2175, BAMBI: 25805, EZH2: 2146, CASP1: 834, GPSM3: 63940, CD36: 948, CXCL13: 10563, COL11A1: 1301, PLOD2: 5352, COL5A1: 1289, TWF1: 5756, CPA3: 1359, IRF5: 3663, CSF1R: 1436, CXC L13: 10563, CSF1R: 1436, TYMS: 7298, FAS: 355, ND2: 6775094, HK2: 3099, PARP4: 143, FADD: 8772, ITGA6: 3655, FANCA: 21 75, RORC: 6097, THBD: 7065, FSTL3: 10272, KRAS: 3845, IL1B: 3553, IRF4: 3662, ROR2: 4920, RPS6KB1: 6198, SFRP4: 6424.

[0009] A second aspect of the invention provides the application of the said gene pair set in either (i) or (ii) the following: (i) evaluating the prognostic effects of triple-negative breast cancer; (ii) constructing a predictive model for adjuvant therapy in breast cancer.

[0010] In a third aspect, the present invention provides a prediction model for adjuvant therapy of breast cancer, which utilizes the aforementioned gene pair set to construct a multilayer artificial neural network to obtain the prediction model for adjuvant therapy of breast cancer.

[0011] Furthermore, the multilayer artificial neural network includes an input layer, a hidden layer, and an output layer; the input layer of the model takes a vector of length 20, where each dimension is 0 or 1, representing the expression relationship of a gene pair.

[0012] Furthermore, the vector is obtained as follows: the relative magnitude of the expression of two genes within the same gene pair in the same sample is calculated using CEIA_B, and represented by binary encoding 0 or 1;

[0013] The formula for calculating CEIA_B is as follows:

[0014] CEIA_B = {1, if log2(ExpressionA+1)>log2(ExpressionB+1); else,0};

[0015] Here, ExpressionA and ExpressionB are the expression values ​​of two genes in the sample. If the expression value of ExpressionA is greater than or equal to that of ExpressionB, it is assigned a value of 1; otherwise, it is assigned a value of 0.

[0016] Furthermore, the hidden layer contains two consecutive layers: a first hidden layer and a second hidden layer;

[0017] The first hidden layer contains 16 neurons, uses the Rectified Linear Unit (ReLU) activation function, and sets the dropout rate to 0.3.

[0018] Furthermore, the second hidden layer contains 8 neurons and uses the ReLU activation function.

[0019] Furthermore, the output layer contains a single neuron that uses the Sigmoid activation function, with an output value between 0 and 1, representing the predicted probability of the patient achieving complete pathological remission.

[0020] Furthermore, in practice, Bulk RNA-seq, microarray transcriptome data, or single-cell RNA-seq data are converted into a vector of the same length as the number of gene pairs, input into a neural network model, and the prognosis of triple-negative breast cancer is predicted based on the output results; when the output value is less than or equal to 0.412, it is judged as non-pathological complete remission; when the output value is greater than 0.412, it is judged as pathological complete remission.

[0021] The beneficial effects of this invention are:

[0022] (1) This invention breaks down cross-platform data barriers, solves the problem that the model cannot be generalized to the single-cell level due to the inconsistent data preprocessing procedures of Bulk RNA-seq, microarray transcriptomics and single-cell RNA-seq, realizes cross-platform data integration and improves the amount of data.

[0023] (2) This invention utilizes gene pair co-characterization, which eliminates the batch effect of Bulk RNA-seq, microarray transcriptomics and single-cell RNA-seq data without the need for standardization and processing. It overcomes the shortcomings of traditional models based on differentially expressed genes (DEGs), such as noise interference from single gene expression abundance, influence of detection platform, significantly limited classification performance and low prediction model accuracy. It uses the GPCS model to break through the noise interference of single genes, accurately capture the interaction between genes, and creates a co-characterization system based on the relative expression of gene pairs, compressing feature redundancy and greatly improving classification accuracy.

[0024] (3) This invention successfully constructed a neural network classification model as a neoadjuvant response prediction model. The AUC of the training set was 0.914, the AUC of the validation set was 0.835, and the AUC of the test set was 0.859, with a sensitivity of up to 95% and a specificity of 93%. This invention's neoadjuvant response prediction model utilizes gene pair features to achieve accurate prediction of neoadjuvant chemotherapy (NAC) patient responses for the first time, providing scientific guidance for personalized clinical decision-making in neoadjuvant chemotherapy and possessing high clinical translational value. Attached Figure Description

[0025] Figure 1 This is a technical roadmap for the present invention.

[0026] Figure 2 This is a flowchart of the computational process for gene pair co-expression.

[0027] Figure 3 This is a distribution map of candidate gene pairs.

[0028] Figure 4 Frequency distribution of the 20 core gene pairs that were ultimately retained.

[0029] Figure 5 This is a diagram of a neural network architecture.

[0030] Figure 6 The AUC plot is used to evaluate the performance of the new auxiliary response prediction model. Detailed Implementation

[0031] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0032] To enable those skilled in the art to better understand the technical solution of this application, the technical solution of this application will be described in detail below with reference to specific embodiments.

[0033] The "gene pair co-expression" strategy proposed in this study has achieved breakthroughs and optimizations at multiple levels. First, the traditional DEG strategy usually relies on the difference in the mean expression of individual genes and the results of statistical tests (such as t-tests, limma analysis, etc.), which is sensitive to technical noise between samples, plateau batch effects, and intra-sample heterogeneity, and is prone to the generation of spurious differences or loss of biological information.

[0034] In contrast, gene-pair features, by constructing the relative expression relationship (rather than absolute expression values) between two genes within the same sample, can effectively eliminate cross-platform differences in expression levels, batch bias, and biases during numerical normalization; even in scenarios with inconsistent data distributions (such as Bulk and scRNA-seq), they maintain strong discriminative stability. Furthermore, we introduce a cross-dataset repeatability criterion, further improving the universality of the features and overcoming the long-standing bottleneck of "strong platform dependence and low reproducibility" in the DEG method.

[0035] Therefore, this invention revolves around the core technical route of "cross-platform gene expression data integration—gene pair feature construction—precise clinical decision-making". Figure 1 A high-precision and highly generalizable response prediction system for neoadjuvant therapy (NAC) in breast cancer was constructed.

[0036] The test materials used in the embodiments of this invention, unless otherwise specified, are all conventional test materials in the art and can be purchased through commercial channels.

[0037] Example 1: Construction of a cross-platform Gene-Pair Classification System (GPCS)

[0038] (1) Data download and processing

[0039] The Bulk RNA-seq and microarray transcriptome data used in this invention were obtained from the GEO (GeneExpression Omnibus) public database, which contains 11 independent datasets covering a total of 705 samples (Table 1). Of these, 251 samples belonged to the PCR category, and the remaining 454 samples were classified as non-PCR (non-pathological complete remission). These datasets cover two different types of RNA sequencing technologies: 583 microarray transcriptome data samples obtained through microarray technology, and 122 Bulk RNA-seq data samples obtained through high-throughput sequencing technology.

[0040] Table 1 Dataset and Sample Distribution

[0041]

[0042] For Bulk RNA-seq data, we downloaded the standardized expression matrix, specifically in TPM (Transcripts Per Million) or FPKM (Fragments Per Kilobase Million) format, to facilitate comparison and analysis between different samples.

[0043] For the single-cell data, the source is the public dataset: GSE246613, which contains 35 samples, including 24 response samples and 11 non-response samples.

[0044] All downloaded data underwent PCA / clustering quality control to ensure consistency and accuracy in subsequent analyses. The resulting gene expression matrix can be used for subsequent differential expression analysis, bioinformatics mining, and identification of potential biomarkers. Data was randomly assigned to groups, with 90% forming the training set and 10% forming the validation set.

[0045] (2) Calculation of gene pair co-expression

[0046] Input data: Gene expression matrix, including 11 public datasets with 705 samples, including 251 PCR samples and 454 non-PCR samples, using Bulk RNA-seq data (Bulk RNA-seq uses TPM / FPKM values ​​or counts values) and microarray transcriptome data; Single-cell data, totaling 35 samples, including 24 responding samples and 11 non-responding samples.

[0047] Selection of gene pairs:

[0048] CEIA_B represents the relative magnitude of expression of two genes within the same sample. During the calculation of CEIA_B, binary encoding (0 or 1) is used to avoid data normalization issues, improve model robustness, avoid cross-platform expression level differences, and reduce interference from differences in platform and experimental conditions.

[0049] The formula for calculating CEIA_B is as follows:

[0050] CEIA_B = {1, if log2(ExpressionA+1)>log2(ExpressionB+1); else,0} (Note: Add a pseudo-count +1 to eliminate zero value interference).

[0051] in:

[0052] • ExpressionA and ExpressionB are the expression values ​​of two genes in a sample.

[0053] • If the expression value of ExpressionA is greater than or equal to ExpressionB, assign a value of 1; otherwise, assign a value of 0.

[0054] Meanwhile, calculate the ratio of 1 / 0 in both the PCR group and the non-PCR group, and then use Fisher's test to evaluate whether the gene pair has significant classification ability. For each gene pair, we count the 1 / 0 counts in the PCR and non-PCR groups and construct the following 2×2 contingency table 2:

[0055] Table 2 1 / 0 counts in the PCR and non-PCR groups

[0056]

[0057] Where:

[0058] X_1, X_0 are the 1 / 0 counts in the PCR group; Y_1, Y_0 are the 1 / 0 counts in the non-PCR group.

[0059] To ensure directionality, we calculate the ratio of 1 to 0 in the PCR and non-PCR groups:

[0060] Calculation of the ratio of 1 to 0 in the PCR group:

[0061] PCR group ratio = X1 / X0

[0062] Calculation of the ratio of 1 to 0 in the non-PCR group:

[0063] Non-PCR group ratio = Y1 / Y0

[0064] Compare the above two ratios:

[0065] If the PCR group ratio > non-PCR group ratio, the p-value remains unchanged; if the PCR group ratio < non-PCR group ratio, the p-value is taken as negative.

[0066] Classification ability evaluation: Using Fisher's test, screen gene pairs with significant classification ability by counting the 1 / 0 ratios in the PCR and non-PCR groups, and select features with significant differences (|p| < 0.1) and stably appearing in multiple datasets. This can effectively ensure that the final classification model focuses on the gene pairs that are most discriminative for the PCR vs non-PCR groups and is not disturbed by data noise or directional errors. The calculation process of gene pair co-expression is shown in Figure 2 .

[0067] (3) Screening core gene pairs

[0068] Initial candidate pool

[0069] Standard settings

[0070] The initial candidate pool was selected based on the criteria of |p| < 0.1 and a frequency of ≥ 4 across the 11 datasets. The reasons for setting this criterion are as follows:

[0071] 1. Considering the priority given to feature “sensitivity” in the initial screening stage, we set the significance threshold to |p|<0.1 in order to retain marginal signals with potential classification value.

[0072] 2. However, relying solely on statistical significance is easily affected by the sample size or noise of a single dataset, leading to a lack of reproducibility in the screening results. Therefore, we further propose a frequency-based filtering strategy across datasets:

[0073] When the criterion is further set to "frequency ≥ 4 in 11 datasets", the number of gene pairs is in the dozens, facilitating subsequent screening. This strategy can effectively filter out noise features that are incidentally significant but unstable in other cohorts, ensuring that the retained features have a certain degree of biological universality and technical robustness.

[0074] Following the method in step (2), and based on the criteria of |p|<0.1 and a frequency ≥4 in 11 datasets, 67 candidate gene pairs were selected from 42,486 gene pairs. See the results below. Figure 3 .

[0075] Feature optimization

[0076] First, Lasso regression was used to narrow down the gene pairs to 46. Then, the Recursive Feature Elimination (RFE) method was used to iteratively remove redundant features.

[0077] Recursive feature elimination recursively eliminates the least important features based on the weights or importance of the machine learning model, thereby improving the model's generalization ability and performance. Its main steps are as follows:

[0078] 1. Train the initial model: Train an SVM supervised learning model using all features.

[0079] 2. Assess feature importance: Calculate the importance of each feature (such as weight coefficient or feature contribution) through the model.

[0080] 3. Recursively remove the least important features: In each round, remove the feature with the smallest weight or the lowest contribution, retain the remaining features, and retrain the model.

[0081] 4. Repeat steps 2-3 until the set number of features (20) is reached.

[0082] The reasons for choosing the number of features are as follows:

[0083] On the one hand, after screening by combining Lasso regression and recursive feature elimination (RFE), it was found that when the number of features was reduced to 20, the model performance tended to saturate, and further increasing the number of features could not significantly improve the classification effect. On the other hand, this number of features can effectively control the parameter scale of the neural network, avoid the risk of overfitting, and ensure the trainability and generalization ability of the model under medium sample size.

[0084] After feature optimization, 20 core gene pairs were ultimately retained. Figure 4 The quantitative relationships of the core gene pairs are as follows:

[0085] ARG2_FANCA——(0), BAMBI_EZH2——(0), CASP1_GPSM3——(1), CD36_CXCL13——(0), COL11A1_PLOD2—— (0), COL5A1_TWF1——(0), CPA3_IRF5——(0), CSF1R_CXCL13——(0), CSF1R_TYMS——(0), CXCL13_FAS—— (1), CXCL13_ND2——(1), EZH2_HK2——(1), EZH2_PARP4——(1), FADD_ITGA6——(0), FANCA_RORC——(1), FANCA_THBD——(1), FSTL3_KRAS——(0), IL1B_ITGA6——(0), IRF4_ROR2——(1), RPS6KB1_SFRP4——(1).

[0086] The GeneID numbers of the above genes in NCBI are as follows: ARG2: 384, FANCA: 2175, BAMBI: 25805, EZH2: 2146, CASP1: 834, GPSM3: 63940, CD36: 948, CXCL13: 10563, COL11A1: 1301, PLOD2: 5352, COL5A1: 1289, TWF1: 5756, CPA3: 1359, IRF5: 3663, CSF1R: 1436, CXCL13: 1 0563, CSF1R: 1436, TYMS: 7298, FAS: 355, ND2: 6775094, HK2: 3099, PARP4: 143, FADD: 8772, ITGA6: 3655, FANCA: 2175 , RORC: 6097, THBD: 7065, FSTL3: 10272, KRAS: 3845, IL1B: 3553, IRF4: 3662, ROR2: 4920, RPS6KB1: 6198, SFRP4: 6424.

[0087] Of the 20 core gene pairs mentioned above, there are a total of 9 positive gene pairs, i.e., a CEIA_B value of 1; and 11 negative gene pairs, i.e., a CEIA_B value of 0.

[0088] Output: A 20-dimensional gene pair feature matrix (n×20, where n is the number of samples) was generated, and the GPCS was successfully constructed. These 20 gene pair features are rank features constructed based on the relative expression relationships of gene pairs.

[0089] Example 2: Construction of a Neural Network Classification Model

[0090] (1) Model architecture design (see Figure 5 )

[0091] To accurately predict the responses of patients achieving pathological complete remission (pCR) and non-pathological complete remission (non-pCR), we designed a multilayer artificial neural network (ANN) model. The input layer used 20 gene pairs of features obtained from preliminary screening, resulting in an input dimension of 20. The hidden layer consists of two consecutive layers: the first hidden layer contains 16 neurons, employing the ReLU (Rectified Linear Unit) activation function to increase the network's non-linear fitting ability, and a Dropout rate of 0.3 is set to reduce the risk of overfitting. The second hidden layer contains 8 neurons, also using the ReLU activation function to further extract higher-order features from the input data. The output layer contains a single neuron, using the Sigmoid activation function, with an output value between 0 and 1, representing the predicted probability of the patient achieving pathological complete remission.

[0092] For the model input layer, the cross-platform gene-pair classification system (GPCS) constructed according to Example 1 is used. The model input is a vector of length 20, where each dimension is 0 or 1, representing the expression relationship of a gene pair.

[0093] In terms of judging the prediction results based on the output value, in order to convert the probability value output by the neural network model into the actual response prediction category (pCR or non-pCR), we use the maximum Youden index as the classification threshold optimization criterion based on the ROC curve of the training set.

[0094] Specifically, we use the `coords()` function in the `pROC` package to extract the point on the ROC curve of the training set that maximizes (sensitivity + specificity - 1) as the optimal classification threshold. This strategy aims to achieve an optimal balance between sensitivity and specificity, avoiding bias caused by a fixed threshold of 0.5. The obtained optimal classification threshold (Youden index) is 0.412. When the output value is less than or equal to 0.412, it is judged as a non-pathological complete remission; when the output value is greater than 0.412, it is judged as a pathological complete remission.

[0095] (2) Model Training and Optimization

[0096] During model training, we employed a weighted cross-entropy loss function to effectively address the imbalance between PCR and non-PCR sample sizes. Since the dataset contains significantly more non-PCR samples (454) than PCR samples (251), we specifically set the class weights to 1:5, increasing the weight of the PCR class to enhance the model's predictive sensitivity for minority class samples.

[0097] The model optimization strategy uses the Adam optimization algorithm, with hyperparameters set to a learning rate of 0.001, an exponential decay rate β1 for the first moment estimation of 0.9, and an exponential decay rate β2 for the second moment estimation of 0.999, in order to achieve efficient and stable training of the model.

[0098] A neural network classification model was successfully constructed as a new auxiliary response prediction model.

[0099] To improve the model's generalization performance on unseen data, the loss function and classification accuracy on both the training and validation sets were monitored during the training process to select the optimal model parameters, further enhancing the reliability and clinical applicability of the predictive model. The resulting neo-assisted response prediction model achieved an AUC of 0.914 on the training set and 0.835 on the validation set, exhibiting a sensitivity of 95%, a specificity of 93%, and stable predictive value.

[0100] Compared with traditional prediction models (AUC<0.75) based on differentially expressed genes (DEG) screening, our model showed a significant advantage in distinguishing between PCR and non-PCR samples, further confirming the reliability and clinical applicability of the model in predicting responses to neoadjuvant chemotherapy.

[0101] Example 3: Performance Evaluation of the Novel Auxiliary Response Prediction Model

[0102] To validate the generalization performance of the neoadjuvant response prediction model on independent data, an independent test dataset was used for evaluation. The dataset used for testing consisted of our own neoadjuvant chemotherapy cohort, totaling 20 samples, including 9 PCR samples and 11 non-PCR samples.

[0103] Before model evaluation, strict quality control was performed to filter out low-expression genes (genes with TPM < 1 were expressed in >90% of samples). During use, the sample data was converted into a vector of the same length as the number of gene pairs (20), and input into the neural network model constructed in Example 2. Based on the output results, the prognostic effect of triple-negative breast cancer was predicted. When the output value was less than or equal to 0.412, it was judged as a non-pathological complete remission; when the output value was greater than 0.412, it was judged as a pathological complete remission. The confusion matrix obtained from testing 20 samples is shown in Table 3.

[0104] Table 3 shows the confusion matrix obtained from testing 20 samples.

[0105]

[0106] The results showed that, using the novel auxiliary response prediction model of this invention, 10 PCRs (8 of which were correctly predicted) and 10 nonPCRs (9 of which were correctly predicted) were predicted, with a test set AUC of 0.859. Figure 6 ).

[0107] This approach achieves the following key innovations, significantly enhancing its clinical application value:

[0108] (1) Gene pair co-representation: The GPCS model breaks through the noise interference of single genes, accurately captures the interaction between genes, and pioneers a co-representation system based on the relative expression of gene pairs, compressing feature redundancy and greatly improving classification accuracy (training set AUC reaches 0.914).

[0109] (2) Cross-platform data integration: The processing workflow effectively eliminates batch effects of Bulk RNA-seq, microarray transcriptomics and single-cell RNA-seq data;

[0110] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A set of gene pairs for evaluating the prognostic effects of triple-negative breast cancer, characterized in that, The gene pair set includes 20 core gene pairs, as follows: ARG2_FANCA, BAMBI_EZH2, CASP1_GPSM3, CD36_CXCL13, COL11A1_PLOD2, COL5A1_TWF1, CPA3_IRF5, CSF1R_CXCL13, CSF1R_TYMS, CXCL13_ FAS, CXCL13_ND2, EZH2_HK2, EZH2_PARP4, FADD_ITGA6, FANCA_RORC, FANCA_THBD, FSTL3_KRAS, IL1B_ITGA6, IRF4_ROR2, RPS6KB1_SFRP4.

2. The gene pair set for evaluating the prognostic effect of triple-negative breast cancer according to claim 1, characterized in that, The GeneID numbers of the genes in the core gene pair in NCBI are as follows: ARG2: 384, FANCA: 2175, BAMBI: 25805, EZH2: 2146, CASP1: 834, GPSM3: 63940, CD36: 948, CXCL13: 10563, COL11A1: 1301, PLOD2: 5352, COL5A1: 1289, TWF1: 5756, CPA3: 1359, IRF5: 3663 , CSF1R: 1436, TYMS: 7298, FAS: 355, ND2: 6775094, HK2: 3099, PARP4: 143, FADD: 8772, ITGA6: 3655, RORC: 60 97, THBD: 7065, FSTL3: 10272, KRAS: 3845, IL1B: 3553, IRF4: 3662, ROR2: 4920, RPS6KB1: 6198, SFRP4: 6424.

3. The use of the gene pair set according to claim 1 or 2 in the following (i) or (ii): (i) evaluating the prognostic effect of triple-negative breast cancer; (ii) constructing a predictive model for adjuvant therapy of breast cancer.

4. A method for predicting the prognosis of triple-negative breast cancer based on a breast cancer adjuvant therapy prediction model, characterized in that, The breast cancer adjuvant therapy prediction model described herein uses the gene pair set described in claim 1 to construct a multi-layer artificial neural network to predict the prognosis of triple-negative breast cancer.

5. The method for predicting the prognosis of triple-negative breast cancer according to claim 4, characterized in that, The multilayer artificial neural network includes an input layer, a hidden layer, and an output layer; the input layer of the model is a vector of length 20, where each dimension is 0 or 1, representing the expression relationship of a gene pair.

6. The method for predicting the prognostic effect of triple-negative breast cancer according to claim 5, characterized in that, The vector is obtained as follows: the relative magnitude of the expression of two genes within the same gene pair in the same sample is calculated using CEIA_B, and represented by binary encoding 0 or 1; The formula for calculating CEIA_B is as follows: CEIA_B = {1, if log2(ExpressionA+1) > log2(ExpressionB+1); else,0}; Here, ExpressionA and ExpressionB are the expression values ​​of two genes in the sample. If the expression value of ExpressionA is greater than or equal to that of ExpressionB, it is assigned a value of 1; otherwise, it is assigned a value of 0.

7. The method for predicting the prognostic effect of triple-negative breast cancer according to claim 5, characterized in that, The hidden layer contains two consecutive layers: the first hidden layer and the second hidden layer; The first hidden layer contains 16 neurons, uses the ReLU activation function, and sets the dropout rate to 0.

3.

8. The method for predicting the prognosis of triple-negative breast cancer according to claim 5, characterized in that, The second hidden layer contains 8 neurons and uses the ReLU activation function.

9. The method for predicting the prognosis of triple-negative breast cancer according to claim 5, characterized in that, The output layer contains a single neuron that uses the Sigmoid activation function. The output value is between 0 and 1, representing the predicted probability that the patient will achieve complete pathological remission.

10. The method for predicting the prognostic effect of triple-negative breast cancer according to claim 4, characterized in that, When using it, Bulk RNA-seq, microarray transcriptome data, or single-cell RNA-seq data are converted into a vector with the same length as the number of gene pairs, input into the neural network model, and the prognosis of triple-negative breast cancer is predicted based on the output results. When the output value is less than or equal to 0.412, it is judged as a non-pathological complete remission; when the output value is greater than 0.412, it is judged as a pathological complete remission.

Citation Information

Patent Citations

  • Method and system for predicting pathological complete remission probability after breast cancer neoadjuvant chemotherapy

    CN112542247A

  • Three-negative breast cancer prognosis risk assessment system

    CN117373534A

  • Gene pair marker composition for prognostic risk prediction of complete pathologic reaction in assisted chemotherapy of sex hormone receptor positive breast cancer and application of gene pair marker composition

    CN118374599A