A prediction method and system for triple negative breast cancer subtype classification

By constructing a prediction model using graph convolutional neural networks and utilizing messenger RNA and long non-coding RNA feature data, the problems of limited sample size and unstable analysis results in the classification of triple-negative breast cancer subtypes were solved, achieving higher classification accuracy and stability.

CN119069000BActive Publication Date: 2026-03-24THE ACAD OF TIANJIN UNIV HEFEI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies for classifying triple-negative breast cancer subtypes suffer from limitations such as limited sample size, unstable analysis results, and unreliability due to differences in standards, making it difficult to achieve accurate subtype classification and personalized treatment.

Method used

A prediction model is constructed using a graph convolutional neural network (GCN). By building a graph dataset and an adjacency matrix, feature extraction and classification are performed using messenger RNA and long non-coding RNA feature data. Complex neural network structures such as multilayer perceptrons are combined to enhance feature expression and learning capabilities.

Benefits of technology

This study improves the accuracy and stability of triple-negative breast cancer subtype classification, provides a new method for determining the correlation between biological characteristics and disease, and guides future research directions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119069000B_ABST
    Figure CN119069000B_ABST
Patent Text Reader

Abstract

The application discloses a kind of triple-negative breast cancer subtype classification prediction method and system, method includes: constructing graph dataset;Graph convolutional neural network prediction classification model is constructed;Graph dataset is used as graph convolutional neural network prediction classification model and is trained until accuracy meets requirements, obtains best graph convolutional neural network prediction classification model;Best graph convolutional neural network prediction classification model is applied to predict triple-negative breast cancer subtype classification.Through the triple-negative breast cancer subtype classification prediction method and system disclosed in the application, the accuracy of classification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of breast cancer identification technology, specifically to a method and system for predicting triple-negative breast cancer subtypes. Background Technology

[0002] Triple-negative breast cancer (TNBC) is a specific type of breast cancer with a relatively poor prognosis, accounting for approximately one-fifth of all breast cancers. Nearly 20% of TNBC patients have BRCA1 / 2 gene mutations. TNBC is characterized by the lack of expression of estrogen receptor (ER), progesterone receptor (PR), and human epidermal growth factor receptor 2 (HER2) in breast tissue. Specifically, progesterone and estrogen receptor values ​​of 1% or less are considered negative, and HER2 results of (- / +) are also considered negative. Additionally, a result of (++) requires FISH testing; no amplification indicates a negative result, while amplification indicates a positive result. Due to the lack of these receptors (ER, PR, HER2), traditional hormone therapy and targeted therapy are usually ineffective against TNBC. Furthermore, due to its unique characteristics, TNBC grows very rapidly compared to other breast cancers, potentially escalating quickly in a short period. Therefore, accurate identification of the TNBC type and timely targeted treatment are crucial.

[0003] Multiple studies have shown that triple-negative breast cancer (TNBC) primarily occurs in premenopausal young women, with a higher chance of visceral metastasis than bone metastasis, and a high probability of brain metastasis. It is a highly diverse group of cancers and one of the most difficult to cure. Its high mortality rate is largely due to the complexity of the cancer and the significant variability in clinical outcomes, necessitating subtyping to better identify molecular-based therapies. Based on cancer genomic clustering analysis, six TNBC subtypes were identified, including two basal-like (BL1 and BL2), immunomodulatory (IM), mesenchymal (M), mesenchymal stem-like (MSL), and luminal androgen receptor (LAR) subtypes. Cell lines representing various subtypes exhibit different sensitivities to targeted therapies currently being explored in laboratory and clinical studies. Identifying various TNBC subtypes and molecular drivers in corresponding cell line models can provide profound insights into the heterogeneity of this disease, aligning TNBC patients with appropriate targeted therapies and providing a preclinical platform for developing effective treatments.

[0004] As we enter the information age, the integration of machine learning and medicine has become an irreversible trend. More and more machine learning methods are being used to detect cancer, such as decision trees and neural networks. With the rapid development of high-throughput sequencing technology, abundant gene expression profiling data provides highly accurate information for cancer patient decision-making and precise diagnosis. However, most of the aforementioned machine learning methods are supervised learning methods. Currently, in bioinformatics, acquiring data is very expensive, and readily available data is often insufficient. Therefore, the role of semi-supervised learning is particularly important. Experts have now developed a semi-supervised learning method for predicting cancer clinical outcomes based on Graph Convolutional Networks (GCNs). Among these, experts have proposed an innovative method applying GCNs to gene selection. For example, Ning Shiqi et al. (A semi-supervised learning method for predicting clinical outcomes of cancer based on graph convolutional networks [J]. Intelligent Computer and Applications, 2018, 8(06):44-48+53.) made some improvements to the network structure of GCN in order to distinguish between normal samples and cancer samples. As a result, more disease-related genes were selected. At the same time, some more promising related genes that had not been discovered before were also explored and discovered, such as CLEC3B. Before the improvement, it could only be used to solve classification problems. After the improvement, it was used for feature selection problems to select some genes related to the disease. Dong Hua et al. (Prediction Model of Triple-Negative Breast Cancer Based on Machine Learning [J]. Journal of Yunnan University (Natural Science Edition), 2017, 39(S1):111-115.) used machine learning algorithms and combined gene expression data of triple-negative breast cancer and non-triple-negative breast cancer in the TCGA database. They used decision tree algorithm and support vector machine feature elimination algorithm in machine learning to build a classification model to distinguish between triple-negative breast cancer and non-triple-negative breast cancer. The decision tree algorithm extracted 9 feature genes with an accuracy of 95.5%, while the support vector machine feature elimination algorithm extracted 6 feature genes with an accuracy of 97.8%. This broke through the unavoidable shortcomings of traditional diagnostic methods and can be further improved through continuous optimization.

[0005] In addition, many researchers abroad have made significant contributions to the study of triple-negative breast cancer. For example, in 2016, Brian et al. (Generation of an algorithm based on minimal gene sets toclinically subtype triple-negative breast cancer patients. BMC Cancer 16, 143 (2016)) constructed a new classification model using the same expression dataset as the original TNBC algorithm. They used gene set enrichment, followed by compact centroid analysis to reduce features, and then used elastic net regularized linear modeling to identify genes in the centroid model for classification across all subtypes. The small gene set model was used to generalize the TNBC subtypes identified by the original 2188 gene model, which has the ability to predict treatment response in the case of standard chemotherapy; in 2018, Santonja et al. (Triple negative breast cancer subtypes and pathologic complete response rate to neoadjuvant chemotherapy. Oncotarget. 2018; 9(41):26406-26416.) identified the Lehmann subtype by gene expression profiles in tumor biopsies from paraffin-embedded tumors of 125 TNBC patients who received neoadjuvant anthracyclines and / or taxanes + / - carboplatin. At the same time, the clinicopathological features of the Lehmann subtype and its relationship with pathological complete response (PCR) of different treatment methods were explored. In the study, only tumors with the LAR phenotype showed a non-basal-like intrinsic subtype (HER2-rich luminal subtype). Meanwhile, it was confirmed that TNBC tumors have high genetic diversity and, from a molecular perspective, they exhibit great heterogeneity. LAR is the least proliferating tumor subtype, while the BL1 subtype appears to be particularly sensitive to chemotherapy regimens, including platinum-based drugs.

[0006] Although many studies at home and abroad have used various methods to classify the molecular subtypes of triple-negative breast cancer, the low incidence and heterogeneity of triple-negative breast cancer and the limited sample size may lead to instability and unreliability of the analysis results. In addition, the subtype classification criteria may differ between different studies, resulting in incomparability of the results. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a method for predicting the subtype classification of triple-negative breast cancer.

[0008] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0009] A predictive method for triple-negative breast cancer subtype classification includes:

[0010] S10, Construct the graph dataset;

[0011] S20, Construct a graph convolutional neural network prediction and classification model;

[0012] S30, train the graph dataset as a graph convolutional neural network prediction and classification model until the accuracy meets the requirements, and obtain the best graph convolutional neural network prediction and classification model.

[0013] S40, applying the best graph convolutional neural network prediction classification model to predict the subtype classification of triple-negative breast cancer;

[0014] S20 involves constructing a graph convolutional neural network prediction and classification model using the following method:

[0015] S21, using the adjacency matrix in the graph dataset as the input node of the neural network, construct the weight matrix. In each layer of the neural network, the input of the node is linearly transformed through the weight matrix to obtain the output features of the node.

[0016] S22, normalize the adjacency matrix to obtain the normalized adjacency matrix;

[0017] S23. Construct a graph convolutional network, using the weight matrix and normalized adjacency matrix as inputs. The graph convolutional network uses the adjacency matrix to propagate the features of nodes, enabling each node to obtain the feature information of its neighboring nodes. By sharing the weight matrix, the feature propagation of each node is processed. Feature extraction is performed on the graph data, calculating the weighted sum of the features of each node and all its neighbors, outputting new features for each node without changing the adjacency structure of the graph.

[0018] In one embodiment of the present invention, constructing a graph dataset includes:

[0019] S11, acquire messenger RNA biological signature data and long non-coding RNA biological signature data; and the six TNBC molecular subtypes contained in these two biological signature data;

[0020] S12, from messenger RNA biological feature data and long non-coding RNA biological feature data, lnc.FAM83H.AS1, gene HSPA14 and gene MGARP were obtained and used as feature vectors;

[0021] S13, Discretize the data values ​​of each feature vector with equal width to obtain multi-interval feature vectors;

[0022] S14: Map the feature values ​​of the cases to the intervals corresponding to the feature vectors of the multiple intervals. Within each interval, connect the cases according to their feature values. Establish the connection between cases in the same interval to represent their similarity or correlation in the feature space. Traverse all intervals to establish the connection and obtain the adjacency matrix.

[0023] S15, the adjacency matrices corresponding to the three eigenvectors form a graph dataset.

[0024] In one embodiment of the present invention, the interval division of the multi-interval feature vector is achieved through the following method:

[0025] Bin i =[min+i×width, min+(i+1)×width];

[0026] In the formula, Bin i Let represent the range of the i-th interval, min represent the minimum value of the interval of the feature vector before partitioning, and width represent the width of each interval.

[0027] In one embodiment of the present invention, the output features of a node are obtained using the following formula:

[0028] Z = X·W + b;

[0029] In the formula, X is the input data, representing the node of the current layer; W is the weight matrix, representing the connection weights between nodes; Z is the output of the node; and b is the bias.

[0030] In one embodiment of the present invention, the output Z of a node, after passing through an activation function, is used as the input of the next layer node, thereby completing the forward propagation calculation; wherein, the activation function of the node's output Z should be obtained by the following formula:

[0031]

[0032] In the formula, σ(Z) is the activation function of the node's output Z, and exp is the natural exponent.

[0033] In one embodiment of the present invention, the normalized adjacency matrix is ​​obtained by the following formula:

[0034]

[0035] In the formula, Let A be the normalized adjacency matrix, D be the degree matrix of adjacency matrix A, and A be the original adjacency matrix.

[0036] In one embodiment of the present invention, a new feature for each node is output without changing the adjacency structure of the graph, obtained by the following formula:

[0037]

[0038] In the formula, H (l+1) H represents the output feature of the (l+1)th layer. (l) This represents the output feature of the l-th layer, where σ is a non-linear activation function. It is the normalized adjacency matrix, W (l) It is the weight matrix of the l-th layer. It is the normalized adjacency matrix The degree matrix.

[0039] This invention also provides a prediction system for triple-negative breast cancer subtype classification, which applies the above-described prediction method for triple-negative breast cancer subtype classification, including:

[0040] The data module is used to build graph datasets;

[0041] The prediction model building module is used to build graph convolutional neural network prediction classification models;

[0042] The training module is used to train the graph dataset as a graph convolutional neural network prediction and classification model until the accuracy meets the requirements and the best graph convolutional neural network prediction and classification model is obtained.

[0043] The prediction module is used to apply the best graph convolutional neural network prediction classification model to predict the subtype classification of triple-negative breast cancer;

[0044] The graph convolutional neural network prediction and classification model is constructed in the following way:

[0045] S21, using the adjacency matrix in the graph dataset as the input node of the neural network, construct the weight matrix. In each layer of the neural network, the input of the node is linearly transformed through the weight matrix to obtain the output features of the node.

[0046] S22, normalize the adjacency matrix to obtain the normalized adjacency matrix;

[0047] S23. Construct a graph convolutional network, using the weight matrix and normalized adjacency matrix as inputs. The graph convolutional network uses the adjacency matrix to propagate the features of nodes, enabling each node to obtain the feature information of its neighboring nodes. By sharing the weight matrix, the feature propagation of each node is processed. Feature extraction is performed on the graph data, calculating the weighted sum of the features of each node and all its neighbors, outputting new features for each node without changing the adjacency structure of the graph.

[0048] In one embodiment of the present invention, constructing a graph dataset includes:

[0049] Acquire biosignature data of messenger RNA and long non-coding RNA; and obtain six TNBC molecular subtypes contained in these two biosignature data.

[0050] We screened messenger RNA biological characteristic data and long non-coding RNA biological characteristic data to obtain lnc.FAM83H.AS1, gene HSPA14 and gene MGARP, and used them as feature vectors.

[0051] Discretize the data values ​​of each feature vector with equal width to obtain multi-interval feature vectors;

[0052] The feature values ​​of cases are mapped to the intervals corresponding to the feature vectors of multiple intervals. Within each interval, edges are connected according to the feature values ​​of the cases. Edges are established between cases in the same interval to represent their similarity or correlation in the feature space. All intervals are traversed to establish edge connections and obtain the adjacency matrix.

[0053] The adjacency matrices corresponding to the three eigenvectors form a graph dataset.

[0054] In one embodiment of the present invention, a new feature for each node is output without changing the adjacency structure of the graph, obtained by the following formula:

[0055]

[0056] In the formula, H (l+1) H represents the output feature of the (l+1)th layer. (l) This represents the output feature of the l-th layer, where σ is a non-linear activation function. It is the normalized adjacency matrix, W (l) It is the weight matrix of the l-th layer. It is the normalized adjacency matrix The degree matrix.

[0057] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention uses graph convolutional neural networks to determine the subtype classification of triple-negative breast cancer through molecular biological features, proposes a method for constructing an adjacency graph of cases, and classifies and verifies graph data constructed with different biological features, providing a new way to determine the correlation between biological features and diseases, and providing guidance for future research directions.

[0058] This invention evaluates the efficiency and stability of graph convolutional neural network models by training them with different numbers of cases, providing an effective solution for related research.

[0059] Graph Convolutional Neural Networks (GCNNs) introduce more complex neural network structures, such as multilayer perceptrons with activation functions, on top of graph convolutional networks to enhance feature expression and learning capabilities. They can learn effective feature representations from complex molecular data, including gene expression patterns and protein-protein interactions, which helps in the discovery of features associated with triple-negative breast cancer. Attached Figure Description

[0060] Figure 1 This is a flowchart of a method for predicting the classification of triple-negative breast cancer subtypes according to an embodiment of the present invention.

[0061] Figure 2 This is a schematic diagram of a graph convolutional neural network according to an embodiment of the present invention.

[0062] Figure 3 This is a schematic diagram of the ROC curve in an embodiment of the present invention.

[0063] Figure 4 This is a block diagram of a prediction system for classifying triple-negative breast cancer subtypes according to an embodiment of the present invention. Detailed Implementation

[0064] To facilitate understanding of the technical solution of the present invention by those skilled in the art, the technical solution of the present invention will now be further described in conjunction with the accompanying drawings.

[0065] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0066] Please see Figure 1 As shown, this invention discloses a method for predicting subtype classification of triple-negative breast cancer, comprising:

[0067] S10, Construct a graph dataset.

[0068] In one embodiment of the present invention, constructing a graph dataset includes:

[0069] S11, acquire messenger RNA biological signature data and long non-coding RNA biological signature data; and six TNBC molecular subtypes contained in these two biological signature data.

[0070] In this embodiment, the biological characteristic data comes from The Cancer Genome Atlas (TCGA), a project jointly initiated by the National Human Genome Research Institute (NHGRI) and the National Cancer Institute (NCI). The TCGA project provides extensive cancer data resources, including genomic data, epigenomic data, and transcriptome data from patient samples. This data is of significant reference value for researchers conducting research and discovering therapies in the field of cancer. The data used in this embodiment comes from confirmed cases in this database, with 305 cases. It includes two types of biological characteristics: messenger RNA (mRNA) and long non-coding RNA (lncRNA), and also includes six TNBC molecular subtypes. The dataset contains 1391 features.

[0071] S12, from messenger RNA biological characteristic data and long non-coding RNA biological characteristic data, lnc.FAM83H.AS1, gene HSPA14 and gene MGARP are screened and used as feature vectors.

[0072] In this embodiment, constructing a graph convolutional neural network prediction classification model requires building a case relationship graph. Considering the number of cases to be classified and the sensitivity, this embodiment uses three feature vectors, lnc.FAM83H.AS1, gene HSPA14, and gene MGARP, to construct an adjacency matrix.

[0073] AM83H.AS1 ​​is a long non-coding RNA (lncRNA) that interacts with other molecules such as miRNAs and proteins, participating in the regulation of cancer-related signaling pathways. The gene HSPA14 encodes a heat shock protein that possesses molecular chaperone and helicase activities within cells, participating in processes such as protein folding, assembly, and disassembly. The gene MGARP is a transcription factor involved in biological processes such as intracellular signal transduction and gene expression regulation.

[0074] S13, discretize the data values ​​of each feature vector with equal width to obtain multi-interval feature vectors.

[0075] The data values ​​of the three feature vectors extracted in step S12 are discretized with equal width, dividing them into 11 equal regions. Equal width discretization is a method of dividing data into several intervals based on their numerical range. Let the maximum value of the original data interval be `max`, the minimum value be `min`, and the data be divided into `k` intervals, with each interval having a width of `width`. The equal width discretization can be expressed as:

[0076] Bin i =[min+i×width, min+(i+1)×width];

[0077] Where, 0≤i <k,Bin i Let represent the range of the i-th interval, min represent the minimum value of the interval of the feature vector before partitioning, and width represent the width of each interval.

[0078] S14: Map the feature values ​​of cases to the intervals corresponding to the multi-interval feature vectors. Within each interval, establish edges based on the feature values ​​of the cases, connecting cases within the same interval to represent their similarity or association in the feature space. Iterate through all intervals to establish edge connections and obtain the adjacency matrix. Each adjacency matrix represents a case association network within a region.

[0079] In this embodiment, edge connections are established between cases within the same interval. For example, if the feature values ​​of case A and case B are similar or the distance is less than a certain threshold, an edge is established between them to represent their similarity.

[0080] S15, the adjacency matrices corresponding to the three eigenvectors form a graph dataset.

[0081] S20, Construct a graph convolutional neural network prediction and classification model.

[0082] Please see Figure 1 and Figure 2 As shown, in one embodiment of the present invention, a graph convolutional neural network prediction and classification model is constructed in the following manner:

[0083] S21, using the adjacency matrix in the graph dataset as the input node of the neural network, construct the weight matrix. In each layer of the neural network, the input of the node is linearly transformed through the weight matrix to obtain the output features of the node.

[0084] In this embodiment, the output features of the nodes are obtained using the following formula:

[0085] Z = X·W + b;

[0086] In the formula, X is the input data, representing the node of the current layer; W is the weight matrix, representing the connection weights between nodes; Z is the node output, obtained by linear transformation of the weight matrix and the input data; and b is the bias, used to adjust the model output.

[0087] During the propagation process of the neural network, the input data X of each node is multiplied by the weight matrix W, and a bias term b is added to obtain the linear transformation output Z of the node. The output Z is then activated by the sigmoid function and used as the input to the next layer node, thus completing the forward propagation calculation. The formula is as follows:

[0088]

[0089] In the formula, σ(Z) is the activation function of the node's output Z, and exp is the natural exponent. This process maps the input value Z to a continuous range between 0 and 1.

[0090] S22, normalize the adjacency matrix to obtain the normalized adjacency matrix.

[0091] In this embodiment, the adjacency matrix A is used to represent the relationship between nodes. It undergoes symmetric normalization to reduce the influence between different node degrees, making the network structure more comparable. The formula is as follows:

[0092]

[0093] In the formula, Let A be the normalized adjacency matrix, D be the degree matrix of adjacency matrix A, and A be the original adjacency matrix. This is achieved by dividing the adjacency matrix A by the square root of the degree of each node.

[0094] In this embodiment of the invention, each graph data consists of a weight matrix W and an adjacency matrix. This is used to represent the input data of the model.

[0095] S23. Construct a graph convolutional network, using the weight matrix and normalized adjacency matrix as inputs. The graph convolutional network uses the adjacency matrix to propagate the features of nodes, enabling each node to obtain the feature information of its neighboring nodes. By sharing the weight matrix, the feature propagation of each node is processed. Feature extraction is performed on the graph data, calculating the weighted sum of the features of each node and all its neighbors, outputting new features for each node without changing the adjacency structure of the graph.

[0096] In one embodiment of the present invention, a local first-order approximation of the Chebyshev polynomial is used to connect the weight matrix and the link matrix, simplifying the calculation process of the graph convolution kernel and reducing the computational complexity of calculating the Laplacian matrix in graph convolution. The formula for outputting new features for each node without changing the adjacency structure of the graph is as follows:

[0097]

[0098] In the formula, H (l+1) H represents the output feature of the (l+1)th layer. (l) This represents the output feature of the l-th layer, where σ is a non-linear activation function. It is the normalized adjacency matrix, W (l) It is the weight matrix of the l-th layer. It is the normalized adjacency matrix The degree matrix.

[0099] S30: Train the graph dataset as a graph convolutional neural network prediction and classification model until the accuracy meets the requirements, and obtain the best graph convolutional neural network prediction and classification model.

[0100] In one embodiment of the invention, each layer of the model shares a parameter transformation matrix (weight matrix W and adjacency matrix A), the dimension of which is independent of the number of vertices in the input graph. Therefore, this parameter transformation matrix is ​​shared by all vertices within the same layer, which can be understood as a shared parameter transformation for the features of different vertices. Furthermore, the model employs a graph convolutional network with two hidden layers, each with embedding dimensions of 32 and 16. The model transforms and updates the input features through these two graph convolutional layers, mapping the features to a 32-dimensional space and then to a 16-dimensional space. This two-layer architecture helps the model better learn the representation and features of the data, improving the nonlinear modeling ability for classification tasks while ensuring learning efficiency and accuracy.

[0101] S40, using the best graph convolutional neural network prediction classification model to predict the subtype classification of triple-negative breast cancer.

[0102] Please see Figures 1 to 3 As shown, in one embodiment of the present invention, the model performance is evaluated by the following three aspects of performance results (a), (b), and (c).

[0103] (a) The impact of different data splits on accuracy

[0104] The graph dataset was divided into training and testing sets, and further partitioned in five ratios: 1:2, 1:3, 2:1, 3:1, and 1:1. After 200 iterations (epochs), the efficiency and accuracy of the model were validated, and the results are shown in Table 1.

[0105] Table 1. Impact of segmentation on accuracy for different datasets

[0106] Segmentation ratio Accuracy after 200 epochs 1∶1 75.1% 1∶2 71.4% 1∶3 64.8% 2∶1 89.2% 3∶1 89.3%

[0107] As can be seen from Table 1, the accuracy of the model gradually increases with the number of training cases. After reaching 200 cases, the accuracy of the model is basically stable. This shows that the graph convolutional neural network prediction and classification model used in this invention can converge quickly and has high efficiency.

[0108] (b) The impact of different training iterations on accuracy

[0109] As shown in result (a), using a 3:1 split ratio for training and testing the data resulted in an accuracy improvement of less than 0.01, and the impact of different numbers of iterations on training accuracy is shown in Table 2.

[0110] Table 2. Effect of different iteration numbers on accuracy

[0111] epoch Training accuracy 1 40.7% 10 72.0% 20 85.3% 50 92.2% 100 93.3% 200 98.6%

[0112] As can be seen from Table 2, the graph convolutional neural network prediction classification model can quickly optimize parameters and approach a local optimum in the early stage of training. As the iteration progresses, after 100 iterations, the model convergence speed becomes relatively flat and the optimization speed is relatively slow. Therefore, it is a relatively reasonable choice to use different cases to train multiple times in this invention to obtain a better model.

[0113] (c) The impact of different features on accuracy

[0114] The experiment used three features—lnc.FAM83H.AS1, gene HSPA14, and gene MGARP—to construct an adjacency matrix to examine the magnitude of different features' influence on disease factors. Table 3 lists the impact of different feature vectors on training accuracy:

[0115] Table 3. The impact of different feature vectors on accuracy

[0116] feature Training accuracy lnc.FAM83H.AS1 95.8% HSPA14 gene 92.1% MGARP gene 92.3%

[0117] As can be seen from Table 3, different feature vector methods have a significant impact on classification accuracy. The lnc.FAM83H.AS1 ​​feature shows a greater correlation with the disease, which suggests that this feature vector should be given more attention than other feature vectors.

[0118] Determine the evaluation metric for the model: The ROC curve is used as the evaluation metric, mainly to reflect the comprehensive index of continuous variables such as sensitivity and specificity. The ROC curve uses the false positive rate (FPR) as the x-axis and the true positive rate (TPR, also known as recall) as the y-axis to show the classification performance of the model at different thresholds.

[0119] In the ROC curve, the horizontal axis is FPR, and the calculation formula is:

[0120]

[0121] Where FP represents a false positive and TN represents a true negative; the vertical axis represents TPR, calculated using the following formula:

[0122]

[0123] TP represents a true positive and FN represents a false negative.

[0124] The area under the ROC curve (AUC area under the curve) is called the AUC value. A value closer to 1 indicates better model performance and a better ability to distinguish between positive and negative examples. An AUC value of 0.5 indicates that the graph convolutional neural network's predictive classification model performs similarly to random guessing. This invention uses the ratio of accurate cases to (accurate cases + missed cases) to represent the true positive rate, and the ratio of misclassified cases to (misclassified cases + missed cases) to approximate the false positive rate. The ROC curve is attached. Figure 3 As shown in the figure, the AUC of this model reaches between 0.8 and 0.9, indicating that the graph convolutional neural network prediction classification model has high classification accuracy.

[0125] Please see Figures 1 to 4 As shown, the present invention also provides a prediction system for triple-negative breast cancer subtype classification, which applies the above-described prediction method for triple-negative breast cancer subtype classification, including:

[0126] The data module is used to build graph datasets;

[0127] The prediction model building module is used to build graph convolutional neural network prediction classification models;

[0128] The training module is used to train the graph dataset as a graph convolutional neural network prediction and classification model until the accuracy meets the requirements and the best graph convolutional neural network prediction and classification model is obtained.

[0129] The prediction module is used to apply the best graph convolutional neural network prediction classification model to predict the subtype classification of triple-negative breast cancer.

[0130] The graph convolutional neural network prediction and classification model is constructed in the following way:

[0131] S21, using the adjacency matrix in the graph dataset as the input node of the neural network, construct the weight matrix. In each layer of the neural network, the input of the node is linearly transformed through the weight matrix to obtain the output features of the node.

[0132] S22, normalize the adjacency matrix to obtain the normalized adjacency matrix;

[0133] S23. Construct a graph convolutional network, using the weight matrix and normalized adjacency matrix as inputs. The graph convolutional network uses the adjacency matrix to propagate the features of nodes, enabling each node to obtain the feature information of its neighboring nodes. By sharing the weight matrix, the feature propagation of each node is processed. Feature extraction is performed on the graph data, calculating the weighted sum of the features of each node and all its neighbors, outputting new features for each node without changing the adjacency structure of the graph.

[0134] In this embodiment, constructing a graph dataset includes:

[0135] Acquire biosignature data of messenger RNA and long non-coding RNA; and obtain six TNBC molecular subtypes contained in these two biosignature data.

[0136] We screened messenger RNA biological characteristic data and long non-coding RNA biological characteristic data to obtain lnc.FAM83H.AS1, gene HSPA14 and gene MGARP, and used them as feature vectors.

[0137] Discretize the data values ​​of each feature vector with equal width to obtain multi-interval feature vectors;

[0138] The feature values ​​of cases are mapped to the intervals corresponding to the feature vectors of multiple intervals. Within each interval, edges are connected according to the feature values ​​of the cases. Edges are established between cases in the same interval to represent their similarity or correlation in the feature space. All intervals are traversed to establish edge connections and obtain the adjacency matrix.

[0139] The adjacency matrices corresponding to the three eigenvectors form a graph dataset.

[0140] In this embodiment, the new features of each node are output without changing the adjacency structure of the graph, and are obtained using the following formula:

[0141]

[0142] In the formula, H (l+1) H represents the output feature of the (l+1)th layer. (l) This represents the output feature of the l-th layer, where σ is a non-linear activation function. It is the normalized adjacency matrix, W (l) It is the weight matrix of the l-th layer. It is the normalized adjacency matrix The degree matrix.

[0143] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.

[0144] The above embodiments are merely examples of implementation methods of the invention. The scope of protection of the present invention is not limited to the above embodiments. For those skilled in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all fall within the scope of protection of the present invention.

Claims

1. A predictive method for subtype classification of triple-negative breast cancer, characterized in that, This provides a method for determining the correlation between biological characteristics and diseases, including: S10, Construct the graph dataset, including: S11, acquire messenger RNA biological signature data and long non-coding RNA biological signature data; and the six TNBC molecular subtypes contained in these two biological signature data; S12, from messenger RNA biological feature data and long non-coding RNA biological feature data, lnc.FAM83H.AS1, gene HSPA14 and gene MGARP were obtained and used as feature vectors; S13, Discretize the data values ​​of each feature vector with equal width to obtain multi-interval feature vectors; S14: Map the feature values ​​of cases to the intervals corresponding to the feature vectors of multiple intervals. Within each interval, establish edge connections based on the feature values ​​of cases. Establish edge connections between cases within the same interval to represent their similarity or correlation in the feature space. Traverse all intervals to establish edge connections and obtain the adjacency matrix. Each adjacency matrix represents the case association network within an interval. S15, the adjacency matrices corresponding to the three feature vectors form a graph dataset; S20, Construct a graph convolutional neural network prediction and classification model; S30: Use the graph dataset as the training set for the graph convolutional neural network prediction and classification model until the accuracy meets the requirements, and obtain the best graph convolutional neural network prediction and classification model. S40, applying the best graph convolutional neural network prediction classification model to predict the subtype classification of triple-negative breast cancer; S20 involves constructing a graph convolutional neural network prediction and classification model using the following method: S21, using the adjacency matrix in the graph dataset as the input node of the neural network, construct the weight matrix. In each layer of the neural network, the input of the node is linearly transformed through the weight matrix to obtain the output features of the node. S22, normalize the adjacency matrix to obtain the normalized adjacency matrix; S23. Construct a graph convolutional network, using the weight matrix and normalized adjacency matrix as inputs. The graph convolutional network uses the adjacency matrix to propagate the features of nodes, enabling each node to obtain the feature information of its neighboring nodes. By sharing the weight matrix, the feature propagation of each node is processed. Feature extraction is performed on the graph data, calculating the weighted sum of the features of each node and all its neighbors, outputting new features for each node without changing the adjacency structure of the graph.

2. The method for predicting triple-negative breast cancer subtypes according to claim 1, characterized in that, The interval division of multi-interval feature vectors is achieved through the following methods: ; In the formula, Represented as the first The range of each interval This represents the minimum value of the interval of the eigenvectors before partitioning. This represents the width of each interval.

3. The method for predicting triple-negative breast cancer subtypes according to claim 1, characterized in that, The output features of a node are obtained using the following formula: ; In the formula, It is the input data, representing the nodes of the current layer; It is a weight matrix, representing the connection weights between nodes; It is the output of the node. It is a bias.

4. The method for predicting triple-negative breast cancer subtypes according to claim 3, characterized in that, Node output After passing through the activation function, the result serves as the input to the next layer of nodes, thus completing the forward propagation computation; where the node's output... The activation function should be obtained using the following formula: ; In the formula, For the output of the node Activation function, It is the natural index.

5. The method for predicting triple-negative breast cancer subtypes according to claim 1, characterized in that, The normalized adjacency matrix is ​​obtained using the following formula: ; In the formula, This is the normalized adjacency matrix. Adjacency matrix The degree matrix, This is the original adjacency matrix.

6. The method for predicting triple-negative breast cancer subtypes according to claim 1, characterized in that, The new features of each node are output without changing the adjacency structure of the graph, obtained using the following formula: ; In the formula, Indicates the first The output features of the layer Indicates the first The output features of the layer It is a non-linear activation function. It is the normalized adjacency matrix. It is the first The weight matrix of the layer, It is the normalized adjacency matrix The degree matrix.

7. A predictive system for triple-negative breast cancer subtype classification, employing the predictive method for triple-negative breast cancer subtype classification according to any one of claims 1-6, characterized in that, include: The data module is used to build graph datasets; The prediction model building module is used to build graph convolutional neural network prediction classification models; The training module is used to train the graph dataset as the training set for the graph convolutional neural network prediction and classification model until the accuracy meets the requirements and the best graph convolutional neural network prediction and classification model is obtained. The prediction module is used to apply the best graph convolutional neural network prediction classification model to predict the subtype classification of triple-negative breast cancer; The graph convolutional neural network prediction and classification model is constructed in the following way: S21, using the adjacency matrix in the graph dataset as the input node of the neural network, construct the weight matrix. In each layer of the neural network, the input of the node is linearly transformed through the weight matrix to obtain the output features of the node. S22, normalize the adjacency matrix to obtain the normalized adjacency matrix; S23. Construct a graph convolutional network, using the weight matrix and normalized adjacency matrix as inputs. The graph convolutional network uses the adjacency matrix to propagate the features of nodes, enabling each node to obtain the feature information of its neighboring nodes. By sharing the weight matrix, the feature propagation of each node is processed. Feature extraction is performed on the graph data, calculating the weighted sum of the features of each node and all its neighbors, outputting new features for each node without changing the adjacency structure of the graph.

8. The prediction system for triple-negative breast cancer subtype classification according to claim 7, characterized in that, The new features of each node are output without changing the adjacency structure of the graph, obtained using the following formula: ; In the formula, Indicates the first The output features of the layer Indicates the first The output features of the layer It is a non-linear activation function. It is the normalized adjacency matrix. It is the first The weight matrix of the layer, It is the normalized adjacency matrix The degree matrix.

Citation Information

Patent Citations

  • Breast cancer subtype classification method based on convolutional neural network

    CN113283477A

  • Breast cancer subtype classification method and system based on graph convolutional neural network

    CN116959572A