Cancer clinical staging prediction method and system based on VAE-transformer
By combining the variational autoencoder and Transformer network methods, the problem of poor feature extraction and model generalization ability of high-dimensional gene expression data is solved, and efficient and accurate diagnosis of cancer clinical staging is achieved, and personalized treatment is supported.
Patent Information
- Application Number
- CN202510555783.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-08
AI Technical Summary
The existing clinical staging methods of cancer have problems such as insufficient feature extraction, poor generalization ability of model, low diagnostic efficiency and low accuracy when processing high-dimensional gene expression data, especially in high-dimensional and low-sample-sized datasets.
Using a combination of variational autoencoder (VAE) and Transformer network, efficient processing of gene expression data in cancer patients is achieved through feature extraction and global dependency modeling. The specific steps include: obtaining gene expression data, pre-processing, using VAE for low-dimensional feature extraction, using Transformer for feature conversion, and finally inputting a classification model for clinical staging prediction of cancer.
It significantly improves the level and accuracy of automated diagnosis of clinical staging of cancer, can effectively process high-dimensional and low sample size gene expression data, avoids the problems of insufficient feature extraction and poor generalization ability in traditional methods, and provides support for the accurate diagnosis and personalized treatment of cancer.
Smart Images

Figure CN120452788A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cancer analysis, and in particular to a method and system for predicting cancer clinical stages based on a VAE-transformer. Background Art
[0002] Clinical staging of cancer is a key step in developing treatment plans and assessing prognosis. Traditional cancer staging methods rely primarily on pathological examinations and imaging analysis, which rely heavily on the physician's experience and are subject to subjectivity and time-consuming. With the continuous advancement of medical technology, the application of molecular biology-based gene expression data in cancer diagnosis and staging is gaining increasing attention. However, cancer staging methods based on gene expression data still face many challenges:
[0003] High-dimensional data: Gene expression data is typically high-dimensional and has low sample size. Traditional machine learning methods are prone to the "curse of dimensionality" when processing such high-dimensional data, leading to problems such as high computational complexity and poor model generalization.
[0004] Inadequate feature extraction: Traditional methods often rely on manually designed features for feature extraction, which struggles to fully capture the complex nonlinear relationships in gene expression data. Gene expression data contains rich biological information, requiring effective methods to automatically extract these key features.
[0005] Poor model generalization: Existing machine learning-based cancer staging models have limited generalization capabilities and are difficult to adapt to the needs of different datasets and clinical scenarios. Due to the heterogeneity of cancer and individual differences in patients, improving the model's generalization ability is crucial for clinical application.
[0006] Diagnostic efficiency and accuracy: Traditional methods suffer from low diagnostic efficiency and accuracy in clinical applications. With the advent of the era of medical big data, more efficient and accurate automated diagnostic tools are needed to assist doctors in decision-making.
[0007] It can be seen that the existing cancer clinical staging prediction methods have major defects. Summary of the Invention
[0008] The technical problem to be solved by the present invention is to provide a cancer clinical staging prediction method and system based on VAE-transformer, which can combine the feature extraction capability of VAE and the global dependency modeling capability of Transformer, and has the advantages of being able to solve the limitations of traditional methods in processing high-dimensional gene expression data and improve the automation level, accuracy and efficiency of cancer clinical staging.
[0009] In order to solve the above technical problems, the first aspect of the present invention discloses a method for predicting cancer clinical stages based on VAE-transformer, the method comprising:
[0010] Obtaining gene expression data of a target patient, and preprocessing the gene expression data to obtain initial data;
[0011] Extracting features from the initial data using a variational autoencoder to obtain a low-dimensional latent feature representation;
[0012] Performing feature transformation on the low-dimensional potential feature representation through a Transformer network to obtain a high-dimensional feature representation;
[0013] The high-dimensional feature representation is input into a classification model to obtain a cancer clinical stage prediction result.
[0014] As an optional embodiment, in the first aspect of the present invention, obtaining gene expression data of a target patient and preprocessing the gene expression data to obtain initial data comprises the following steps:
[0015] Missing values are deleted from the gene expression data, and the gene expression data after the missing values are deleted are normalized to obtain initial data.
[0016] As an optional embodiment, in the first aspect of the present invention, extracting features from the initial data using a variational autoencoder to obtain a low-dimensional potential feature representation comprises the following steps:
[0017] Mapping the initial data to a latent space through an encoder network in the variational autoencoder to obtain latent variable mean features and logarithmic variance features;
[0018] According to the latent variable mean feature and the logarithmic variance feature, a low-dimensional latent feature representation is generated through a reparameterization technique.
[0019] As an optional implementation, in the first aspect of the present invention, performing feature conversion on the low-dimensional latent feature representation through a Transformer network to obtain a high-dimensional feature representation includes the following steps:
[0020] Inputting the low-dimensional latent feature representation into the multi-head self-attention mechanism of the Transformer network, and calculating the global dependency weights between each feature element in the low-dimensional latent feature representation;
[0021] performing weighted summation on the feature elements in the low-dimensional latent feature representation according to the global dependency weight to obtain a weighted feature representation;
[0022] The weighted feature representation is input into the feedforward neural network in the Transformer network, and a high-dimensional feature representation is obtained through nonlinear transformation.
[0023] As an optional embodiment, in the first aspect of the present invention, inputting the low-dimensional latent feature representation into the multi-head self-attention mechanism of the Transformer network and calculating the global dependency weights between the feature elements in the low-dimensional latent feature representation includes:
[0024] Input the low-dimensional latent feature representation into the multi-head self-attention mechanism of the Transformer network to generate a query vector, a key vector, and a value vector for each feature;
[0025] The dot product between the key vector and the value vector is calculated and scaled by a scaling factor to get the global dependency weight.
[0026] As an optional implementation, in the first aspect of the present invention, performing weighted summation of feature elements in the low-dimensional potential feature representation according to the global dependency weight to obtain a weighted feature representation includes:
[0027] Normalizing the global dependency weight to obtain a correlation weight;
[0028] The correlation weight and the value vector are weighted and summed to obtain a weighted feature representation.
[0029] As an optional embodiment, in the first aspect of the present invention, inputting the high-dimensional feature representation into a classification model to obtain a cancer clinical stage prediction result includes:
[0030] The classification model is a multi-layer perceptron. The high-dimensional feature representation is input into the multi-layer perceptron and processed by a fully connected layer and an activation function to obtain a cancer clinical stage prediction result.
[0031] The second aspect of the present invention discloses a cancer clinical staging prediction system based on VAE-transformer, the system comprising:
[0032] an acquisition module, the acquisition module being used to acquire gene expression data of a target patient and preprocess the gene expression data to obtain initial data;
[0033] An extraction module, configured to extract features from the initial data using a variational autoencoder to obtain a low-dimensional latent feature representation;
[0034] A conversion module, configured to perform feature conversion on the low-dimensional potential feature representation through a Transformer network to obtain a high-dimensional feature representation;
[0035] A prediction module is used to input the high-dimensional feature representation into a classification model to obtain a cancer clinical staging prediction result.
[0036] As an optional embodiment, in the second aspect of the present invention, the acquisition module acquires the gene expression data of the target patient and preprocesses the gene expression data to obtain initial data, including the following steps:
[0037] Missing values are deleted from the gene expression data, and the gene expression data after the missing values are deleted are normalized to obtain initial data.
[0038] As an optional embodiment, in the second aspect of the present invention, the extraction module performs feature extraction on the initial data through a variational autoencoder to obtain a low-dimensional potential feature representation, comprising the following steps:
[0039] Mapping the initial data to a latent space through an encoder network in the variational autoencoder to obtain latent variable mean features and logarithmic variance features;
[0040] According to the latent variable mean feature and the logarithmic variance feature, a low-dimensional latent feature representation is generated through a reparameterization technique.
[0041] As an optional implementation, in the second aspect of the present invention, the conversion module performs feature conversion on the low-dimensional latent feature representation through a Transformer network to obtain a high-dimensional feature representation, comprising the following steps:
[0042] Inputting the low-dimensional latent feature representation into the multi-head self-attention mechanism of the Transformer network, and calculating the global dependency weights between each feature element in the low-dimensional latent feature representation;
[0043] performing weighted summation on the feature elements in the low-dimensional latent feature representation according to the global dependency weight to obtain a weighted feature representation;
[0044] The weighted feature representation is input into the feedforward neural network in the Transformer network, and a high-dimensional feature representation is obtained through nonlinear transformation.
[0045] As an optional embodiment, in the second aspect of the present invention, the conversion module inputs the low-dimensional latent feature representation into the multi-head self-attention mechanism of the Transformer network, and calculates the global dependency weights between each feature element in the low-dimensional latent feature representation, including:
[0046] Input the low-dimensional latent feature representation into the multi-head self-attention mechanism of the Transformer network to generate a query vector, a key vector, and a value vector for each feature;
[0047] The dot product between the key vector and the value vector is calculated and scaled by a scaling factor to get the global dependency weight.
[0048] As an optional implementation, in the second aspect of the present invention, the conversion module performs weighted summation on the feature elements in the low-dimensional latent feature representation according to the global dependency weight to obtain a weighted feature representation, including:
[0049] Normalizing the global dependency weight to obtain a correlation weight;
[0050] The correlation weight and the value vector are weighted and summed to obtain a weighted feature representation.
[0051] As an optional embodiment, in the second aspect of the present invention, the prediction module inputs the high-dimensional feature representation into a classification model to obtain a cancer clinical stage prediction result, including:
[0052] The classification model is a multi-layer perceptron. The high-dimensional feature representation is input into the multi-layer perceptron and processed by a fully connected layer and an activation function to obtain a cancer clinical stage prediction result.
[0053] The third aspect of the present invention discloses another cancer clinical staging prediction device based on VAE-transformer, the device comprising:
[0054] a memory storing executable program code;
[0055] a processor coupled to the memory;
[0056] The processor calls the executable program code stored in the memory to execute some or all of the steps in the VAE-transformer-based cancer clinical staging prediction method disclosed in the first aspect of the embodiment of the present invention.
[0057] A fourth aspect of an embodiment of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute some or all of the steps in the VAE-transformer-based cancer clinical staging prediction method disclosed in the first aspect of the embodiment of the present invention.
[0058] Compared with existing technologies, the embodiments of the present invention have the following advantages: By combining the advantages of variational autoencoders (VAEs) and Transformer networks, they achieve efficient feature extraction and global dependency modeling of cancer patient gene expression data, significantly improving the automated diagnosis, accuracy, and efficiency of cancer clinical staging. This prediction method not only handles high-dimensional, low-sample-volume gene expression data but also effectively avoids the problems of insufficient feature extraction and poor model generalization in traditional methods, providing strong support for accurate cancer diagnosis and personalized treatment. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0060] Figure 1 This is a flowchart of a method for predicting cancer clinical stages based on a VAE-transformer, as disclosed in an embodiment of the present invention;
[0061] Figure 2 This is an ROC curve diagram drawn based on the experimental results of different cancer data sets using a VAE-transformer-based cancer clinical stage prediction method disclosed in an embodiment of the present invention;
[0062] Figure 3 1 is a schematic diagram of the structure of a cancer clinical staging prediction system based on VAE-transformer disclosed in an embodiment of the present invention;
[0063] Figure 4 This is a schematic structural diagram of another VAE-transformer-based cancer clinical staging prediction device disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0064] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0065] The terms "first," "second," and so on, in the description and claims of the present invention and the accompanying drawings are used to distinguish between different items, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or end comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed therein, or may optionally include other steps or elements inherent to such process, method, product, or end.
[0066] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0067] The present invention discloses a cancer clinical staging prediction method and system based on VAE-transformer. By combining the advantages of variational autoencoders (VAE) and Transformer networks, efficient feature extraction and global dependency modeling of cancer patient gene expression data are achieved, significantly improving the automated diagnosis level, accuracy and efficiency of cancer clinical staging. This prediction method can not only process high-dimensional, low-sample-volume gene expression data, but also effectively avoid problems such as insufficient feature extraction and poor model generalization ability in traditional methods, providing strong support for accurate diagnosis and personalized treatment of cancer. The following are detailed descriptions.
[0068] Example 1
[0069] See also Figure 1 , Figure 1 This is a flowchart of a method for predicting cancer clinical stages based on VAE-transformer disclosed in an embodiment of the present invention. Figure 1The described method is applied to a cancer clinical staging prediction device based on VAE-transformer, which can be a corresponding prediction terminal, prediction device or server, and the server can be a local server or a cloud server, which is not limited in the embodiment of the present invention. Figure 1 As shown, the VAE-transformer-based cancer clinical stage prediction method can include the following operations:
[0070] 101. Obtain gene expression data of the target patient and preprocess the gene expression data to obtain initial data.
[0071] 102. The initial data is extracted through variational autoencoder to obtain low-dimensional latent feature representation.
[0072] 103. The low-dimensional latent feature representation is transformed through the Transformer network to obtain a high-dimensional feature representation.
[0073] 104. Input the high-dimensional feature representation into the classification model to obtain the cancer clinical stage prediction results.
[0074] As can be seen, the method described in the embodiment of the present invention, by combining the advantages of variational autoencoders (VAEs) and Transformer networks, can achieve efficient feature extraction and global dependency modeling of cancer patient gene expression data, significantly improving the level, accuracy, and efficiency of automated diagnosis of cancer clinical staging. This prediction method not only can handle high-dimensional, low-sample-volume gene expression data, but also effectively avoids the problems of insufficient feature extraction and poor model generalization in traditional methods, providing strong support for accurate cancer diagnosis and personalized treatment.
[0075] In an optional embodiment, the step 101 of obtaining gene expression data of a target patient and preprocessing the gene expression data to obtain initial data includes the following steps:
[0076] Missing values were deleted from the gene expression data, and the gene expression data after deletion of missing values were normalized to obtain initial data.
[0077] Specifically, miRNA gene features in gene expression data were selected, missing values were deleted, and StandardScaler was used to standardize the features to ensure that their mean was 0 and variance was 1.
[0078] In an optional embodiment, step 102 of extracting features from the initial data using a variational autoencoder to obtain a low-dimensional latent feature representation includes the following steps:
[0079] The initial data is mapped to the latent space through the encoder network in the variational autoencoder to obtain the latent variable mean feature and logarithmic variance feature;
[0080] According to the mean and logarithmic variance characteristics of the latent variables, a low-dimensional latent feature representation is generated through reparameterization techniques.
[0081] Specifically, the variational autoencoder (VAE) used in this step is a generative model for learning the distribution of latent variables from input data. The input data is mapped to the latent space through a simple neural network encoder, and then the latent variables are sampled through the reparameterization technique. In an embodiment of the present invention, the encoder in the variational autoencoder consists of two fully connected layers and a ReLU activation function, outputs the latent variable mean feature mu and logarithmic variance feature logvar of the latent space, and samples the low-dimensional latent feature representation through the reparameterization technique, wherein the two fully connected layers have 128 and 68 hidden units respectively.
[0082] In an optional embodiment, in step 103, feature conversion is performed on the low-dimensional latent feature representation by using a Transformer network to obtain a high-dimensional feature representation, including the following steps:
[0083] The low-dimensional latent feature representation is input into the multi-head self-attention mechanism of the Transformer network to calculate the global dependency weights between each feature element in the low-dimensional latent feature representation;
[0084] According to the global dependency weight, the feature elements in the low-dimensional latent feature representation are weighted and summed to obtain the weighted feature representation;
[0085] The weighted feature representation is input into the feedforward neural network in the Transformer network, and a high-dimensional feature representation is obtained through nonlinear transformation.
[0086] In an optional embodiment, the step of inputting the low-dimensional latent feature representation into the multi-head self-attention mechanism of the Transformer network and calculating the global dependency weights between the feature elements in the low-dimensional latent feature representation includes:
[0087] The low-dimensional latent feature representation is fed into the multi-head self-attention mechanism of the Transformer network to generate a query vector, a key vector, and a value vector for each feature.
[0088] The dot product between the key vector and the value vector is calculated and scaled by a scaling factor to get the global dependency weight.
[0089] In an optional embodiment, in the above step, the feature elements in the low-dimensional latent feature representation are weighted and summed according to the global dependency weight to obtain the weighted feature representation, including:
[0090] Normalize the global dependency weights to get the correlation weights;
[0091] The relevance weight is weighted and summed with the value vector to obtain the weighted feature representation.
[0092] In this embodiment of the present invention, the Transformer network model used in this step is a powerful sequence modeling tool based on the self-attention mechanism. Its core formula includes the self-attention mechanism, multi-head attention, a feedforward neural network, residual connections, and layer normalization. The self-attention mechanism captures global dependencies by calculating the correlation between each element in the input sequence and other elements.
[0093] Specifically, the low-dimensional latent feature representation is linearly transformed to generate the query vector Q, key vector K and value vector V:
[0094] Q=XW Q ,K=XW K ,V=XW V
[0095] Where W Q 、W K 、W V is a learnable weight matrix.
[0096] Next, the attention score is calculated by calculating the correlation between the query vector Q and the key vector K through the dot product, and then scaled and Softmax normalized:
[0097]
[0098] Among them, d k is the dimension of Key, is a scaling factor used to prevent the dot product result from being too large.
[0099] Multi-head attention divides the input into multiple heads, calculates attention separately, and finally concatenates the results:
[0100] MultiHead(Q,K,V)=Concat(head1,…,head h )W O
[0101] in, W O is the output linear transformation matrix.
[0102] The feedforward neural network performs a nonlinear transformation on the output of the self-attention mechanism:
[0103] FFN(X)=max(0,xW1+b1)W2+b2
[0104] Residual connections and layer normalization are added after each sublayer (self-attention and feedforward network):
[0105] LayerNorm(x+Sublayer(x))
[0106] In an optional embodiment, in step 104, the high-dimensional feature representation is input into a classification model to obtain a cancer clinical stage prediction result, including:
[0107] The high-dimensional feature representation is input into the classification model to obtain the cancer clinical stage prediction results.
[0108] Specifically, the classification model is a multi-layer perceptron (MLP), which has two fully connected layers, an input dimension of 32, a hidden layer dimension of 64, and an output dimension of 4, corresponding to the four clinical stage categories. The Softmax function is then used to convert the output into a probability distribution to obtain the cancer clinical stage prediction results.
[0109] As can be seen, the method described in the embodiment of the present invention, by combining the advantages of variational autoencoders (VAEs) and Transformer networks, can achieve efficient feature extraction and global dependency modeling of cancer patient gene expression data, significantly improving the level, accuracy, and efficiency of automated diagnosis of cancer clinical staging. This prediction method not only can handle high-dimensional, low-sample-volume gene expression data, but also effectively avoids the problems of insufficient feature extraction and poor model generalization in traditional methods, providing strong support for accurate cancer diagnosis and personalized treatment.
[0110] To verify the superiority of a cancer clinical staging diagnosis method based on a combination of a variational autoencoder (VAE) and a Transformer disclosed in an embodiment of the present invention over traditional machine learning, in this embodiment, we used traditional machine learning methods such as support vector machine (SVM), logistic regression, random forest, decision tree, XGBoost, and ordinary neural network. The experimental results are shown in the following table.
[0111] Table 1 Comparison of results of different machine learning methods
[0112]
[0113] As an optional implementation, referring to Table 1, the AUC values of the method of the present invention on different data sets are higher than those of traditional machine learning methods. In terms of feature extraction capability, the variational autoencoder (VAE) and Transformer modules can automatically extract nonlinear features in high-dimensional gene expression data, while traditional methods rely on manually designed features and simple linear transformations; in terms of global dependency modeling, the Transformer module is used to model the global dependency relationships of latent space features, which can capture long-distance dependencies in the data, while traditional methods (such as SVM and logistic regression) can only process local features; in terms of generalization capability, through stratified K-fold cross-validation and early stopping strategy, the system of the present invention exhibits stronger generalization capability on multiple data sets, while traditional methods are prone to overfitting and underfitting. In terms of performance indicators such as AUC, accuracy and F1 score, the system of the present invention is significantly superior to traditional machine learning methods, especially in high-dimensional, low-sample-size gene expression data.
[0114] The following briefly describes the training process of this prediction model:
[0115] Gene expression data from 1,000 cancer patients were obtained from the public cancer gene expression database TCGA. MiRNA gene signatures were selected from the samples, and missing values were removed. The clinical stage label, stage, was encoded as an integer using LabelEncoder. Labels were converted from the original stage categories, stage1, stage2, stage3, and stage4, to integer values 0, 1, 2, and 3 to accommodate the requirements of the classification model. The feature data X was normalized using StandardScaler to ensure that each feature had a mean of 0 and a variance of 1. This provided the initial training data, thus preventing model training from being affected by varying feature scales.
[0116] During the training process, the model is trained using cross entropy, KL divergence, and reconstruction loss function, and the model parameters are optimized through back propagation. Specifically, the function of VAE is realized through the KL divergence and reconstruction loss function of the deep neural network system. The reconstruction error is a measure of the difference between the reconstructed data x' and the original data x, and the formula is:
[0117]
[0118] The reconstruction loss is the mean square error, where x i is the i-th sample of the input data, x′ i is the i-th sample reconstructed by the decoder, and N is the total number of samples.
[0119] KL divergence is used to measure the difference between two probability distributions. The formula is:
[0120]
[0121] where d is the dimension of the latent variable, and are the mean and variance of the underlying distribution, is the log variance.
[0122] The function of the classifier is realized through the cross entropy loss function, which is used to measure the difference between the category probability distribution predicted by the model and the true label. Its formula is:
[0123]
[0124] Where N is the number of samples, C is the number of categories, and y ij is the one-hot encoding of the true label (if sample i belongs to category j, then y ij =1, otherwise y ij =0), p ij is the probability that sample i belongs to category j predicted by the model.
[0125] That is, the total network structure loss function of this VAE-transformer-based cancer clinical staging prediction model can be expressed as:
[0126] MSE+KL(q(z|x)||p(z))+CrossEntropyLoss
[0127] Reference Figure 2 During the training process, stratified K-fold cross-validation was used to evaluate the performance of the model, calculate the AUC value of the model, and draw the ROC curve. Specifically, the dataset was divided into a training set and a validation set to ensure that the proportion of samples of each category in each fold was consistent with the original dataset; in each fold, the training set was used to train the model, and the AUC, accuracy, and F1 score of the model were evaluated on the validation set; the average AUC value of all folds was calculated, and the ROC curve was drawn as the final performance indicator of the model.
[0128] During the model training process, an early stopping strategy is used to prevent overfitting: after each training cycle (epoch), the AUC value on the validation set is calculated; if the AUC value on the validation set does not improve over multiple consecutive training cycles, the training is terminated early.
[0129] During the model training process, batch normalization technology is used to accelerate model training and improve model stability.
[0130] During the training process of the model, the Dropout technology is used to randomly discard some neurons to prevent overfitting.
[0131] During the model training process, transfer learning technology is used to initialize model parameters using pre-trained variational autoencoders and Transformer modules to accelerate model training and improve model performance.
[0132] During the model training process, data enhancement technology is used to generate new training samples by randomly perturbing the gene expression data to improve the generalization ability of the model.
[0133] During the model training process, model visualization technology is used to draw the AUC curve to intuitively display the model training process and performance.
[0134] Example 2
[0135] See also Figure 3 , Figure 3 This is a schematic diagram of the structure of a cancer clinical staging prediction system based on VAE-transformer disclosed in an embodiment of the present invention. Figure 3 The described system can be applied to corresponding prediction terminals, prediction devices or servers, and the server can be a local server or a cloud server, which is not limited in the embodiment of the present invention. Figure 3 As shown, the system may include:
[0136] The acquisition module 100 is used to acquire gene expression data of a target patient and pre-process the gene expression data to obtain initial data.
[0137] The extraction module 200 is used to extract features from the initial data through a variational autoencoder to obtain a low-dimensional potential feature representation.
[0138] The conversion module 300 is used to perform feature conversion on the low-dimensional potential feature representation through the Transformer network to obtain a high-dimensional feature representation.
[0139] The prediction module 400 is used to input the high-dimensional feature representation into the classification model to obtain the cancer clinical stage prediction result.
[0140] As can be seen, the system described in the embodiments of the present invention, by combining the advantages of variational autoencoders (VAEs) and Transformer networks, achieves efficient feature extraction and global dependency modeling of cancer patient gene expression data, significantly improving the automated diagnosis, accuracy, and efficiency of cancer clinical staging. This prediction method not only handles high-dimensional, low-sample-volume gene expression data but also effectively avoids the problems of insufficient feature extraction and poor model generalization in traditional methods, providing strong support for accurate cancer diagnosis and personalized treatment.
[0141] In an optional embodiment, the acquisition module 100 acquires gene expression data of a target patient and preprocesses the gene expression data to obtain initial data, including the following steps:
[0142] Missing values were deleted from the gene expression data, and the gene expression data after deletion of missing values were normalized to obtain initial data.
[0143] Specifically, miRNA gene features in gene expression data were selected, missing values were deleted, and StandardScaler was used to standardize the features to ensure that their mean was 0 and variance was 1.
[0144] In an optional embodiment, the extraction module 200 extracts features from the initial data using a variational autoencoder to obtain a low-dimensional latent feature representation, including the following steps:
[0145] The initial data is mapped to the latent space through the encoder network in the variational autoencoder to obtain the latent variable mean feature and logarithmic variance feature;
[0146] According to the mean and logarithmic variance characteristics of the latent variables, a low-dimensional latent feature representation is generated through reparameterization techniques.
[0147] Specifically, the variational autoencoder (VAE) used in this step is a generative model for learning the distribution of latent variables from input data. The input data is mapped to the latent space through a simple neural network encoder, and then the latent variables are sampled through the reparameterization technique. In an embodiment of the present invention, the encoder in the variational autoencoder consists of two fully connected layers and a ReLU activation function, outputs the latent variable mean feature mu and logarithmic variance feature logvar of the latent space, and samples the low-dimensional latent feature representation through the reparameterization technique, wherein the two fully connected layers have 128 and 68 hidden units respectively.
[0148] In an optional embodiment, the conversion module 300 performs feature conversion on the low-dimensional latent feature representation through a Transformer network to obtain a high-dimensional feature representation, including the following steps:
[0149] The low-dimensional latent feature representation is input into the multi-head self-attention mechanism of the Transformer network to calculate the global dependency weights between each feature element in the low-dimensional latent feature representation;
[0150] According to the global dependency weight, the feature elements in the low-dimensional latent feature representation are weighted and summed to obtain the weighted feature representation;
[0151] The weighted feature representation is input into the feedforward neural network in the Transformer network, and a high-dimensional feature representation is obtained through nonlinear transformation.
[0152] In an optional embodiment, the conversion module 300 inputs the low-dimensional latent feature representation into the multi-head self-attention mechanism of the Transformer network and calculates the global dependency weights between the feature elements in the low-dimensional latent feature representation, including:
[0153] The low-dimensional latent feature representation is fed into the multi-head self-attention mechanism of the Transformer network to generate a query vector, a key vector, and a value vector for each feature.
[0154] The dot product between the key vector and the value vector is calculated and scaled by a scaling factor to get the global dependency weight.
[0155] In an optional embodiment, the conversion module 300 performs weighted summation on the feature elements in the low-dimensional latent feature representation according to the global dependency weight to obtain a weighted feature representation, including:
[0156] Normalize the global dependency weights to obtain the correlation weights;
[0157] The relevance weight is weighted and summed with the value vector to obtain the weighted feature representation.
[0158] In this embodiment of the present invention, the Transformer network model used in this step is a powerful sequence modeling tool based on the self-attention mechanism. Its core formula includes the self-attention mechanism, multi-head attention, a feedforward neural network, residual connections, and layer normalization. The self-attention mechanism captures global dependencies by calculating the correlation between each element in the input sequence and other elements.
[0159] Specifically, the low-dimensional latent feature representation is linearly transformed to generate the query vector Q, key vector K and value vector V:
[0160] Q=XW Q ,K=XW K ,V=XW V
[0161] Where W Q 、W K 、W V is a learnable weight matrix.
[0162] Next, the attention score is calculated by calculating the correlation between the query vector Q and the key vector K through the dot product, and then scaled and Softmax normalized:
[0163]
[0164] Among them, d k is the dimension of Key, is a scaling factor used to prevent the dot product result from being too large.
[0165] Multi-head attention divides the input into multiple heads, calculates attention separately, and finally concatenates the results:
[0166] MultiHead(Q,K,V)=Concat(head1,…,head h )W O
[0167] in, W O is the output linear transformation matrix.
[0168] The feedforward neural network performs a nonlinear transformation on the output of the self-attention mechanism:
[0169] FFN(X)=max(0,xW1+b1)W2+b2
[0170] Residual connections and layer normalization are added after each sublayer (self-attention and feedforward network):
[0171] LayerNorm(x+Sublayer(x))
[0172] In an optional embodiment, the prediction module 400 inputs the high-dimensional feature representation into the classification model to obtain a cancer clinical stage prediction result, including:
[0173] The high-dimensional feature representation is input into the classification model to obtain the cancer clinical stage prediction results.
[0174] Specifically, the classification model is a multi-layer perceptron (MLP), which has two fully connected layers, an input dimension of 32, a hidden layer dimension of 64, and an output dimension of 4, corresponding to the four clinical stage categories. The Softmax function is then used to convert the output into a probability distribution to obtain the cancer clinical stage prediction results.
[0175] As can be seen, the system described in the embodiments of the present invention, by combining the advantages of variational autoencoders (VAEs) and Transformer networks, achieves efficient feature extraction and global dependency modeling of cancer patient gene expression data, significantly improving the automated diagnosis, accuracy, and efficiency of cancer clinical staging. This prediction method not only handles high-dimensional, low-sample-volume gene expression data but also effectively avoids the problems of insufficient feature extraction and poor model generalization in traditional methods, providing strong support for accurate cancer diagnosis and personalized treatment.
[0176] Example 3
[0177] See also Figure 4 , Figure 4 This is a schematic diagram of the structure of another cancer clinical staging prediction device based on VAE-transformer disclosed in an embodiment of the present invention. Figure 4 As shown, the device may include:
[0178] A memory 301 storing executable program code;
[0179] a processor 302 coupled to the memory 301;
[0180] The processor 302 calls the executable program code stored in the memory 301 to execute part or all of the steps in the VAE-transformer-based cancer clinical staging prediction method disclosed in the first embodiment of the present invention.
[0181] Example 4
[0182] An embodiment of the present invention discloses a computer storage medium storing computer instructions. When the computer instructions are called, they are used to execute some or all of the steps in the VAE-transformer-based cancer clinical staging prediction method disclosed in Example 1 of the present invention.
[0183] The device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Those skilled in the art can understand and implement the present invention without inventive effort.
[0184] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus the necessary general hardware platform, or of course, by means of hardware. Based on this understanding, the above technical solution, in essence, or the portion that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
[0185] Finally, it should be noted that the VAE-transformer-based cancer clinical staging prediction method and system disclosed in the embodiments of the present invention are only preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features therein may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A cancer clinical stage prediction method based on VAE-transformer, characterized by: The method comprises: Obtaining gene expression data of a target patient, and preprocessing the gene expression data to obtain initial data; Extracting features from the initial data using a variational autoencoder to obtain a low-dimensional latent feature representation; Performing feature transformation on the low-dimensional potential feature representation through a Transformer network to obtain a high-dimensional feature representation; The high-dimensional feature representation is input into a classification model to obtain a cancer clinical stage prediction result.
2. The cancer clinical stage prediction method based on VAE-transformer according to claim 1, characterized in that: The step of obtaining gene expression data of a target patient and preprocessing the gene expression data to obtain initial data comprises the following steps: Missing values are deleted from the gene expression data, and the gene expression data after the missing values are deleted are normalized to obtain initial data.
3. The cancer clinical stage prediction method based on VAE-transformer according to claim 1, characterized in that: The process of extracting features from the initial data using a variational autoencoder to obtain a low-dimensional potential feature representation includes the following steps: Mapping the initial data to a latent space through an encoder network in the variational autoencoder to obtain latent variable mean features and logarithmic variance features; According to the latent variable mean feature and the logarithmic variance feature, a low-dimensional latent feature representation is generated through a reparameterization technique.
4. The cancer clinical stage prediction method based on VAE-transformer according to claim 1, characterized in that: The step of performing feature conversion on the low-dimensional potential feature representation through a Transformer network to obtain a high-dimensional feature representation includes the following steps: Inputting the low-dimensional latent feature representation into the multi-head self-attention mechanism of the Transformer network, and calculating the global dependency weights between each feature element in the low-dimensional latent feature representation; performing weighted summation on the feature elements in the low-dimensional latent feature representation according to the global dependency weight to obtain a weighted feature representation; The weighted feature representation is input into the feedforward neural network in the Transformer network, and a high-dimensional feature representation is obtained through nonlinear transformation.
5. The cancer clinical stage prediction method based on VAE-transformer according to claim 4, characterized in that: Inputting the low-dimensional latent feature representation into the multi-head self-attention mechanism of the Transformer network and calculating the global dependency weights between the feature elements in the low-dimensional latent feature representation include: Input the low-dimensional latent feature representation into the multi-head self-attention mechanism of the Transformer network to generate a query vector, a key vector, and a value vector for each feature; The dot product between the key vector and the value vector is calculated and scaled by a scaling factor to get the global dependency weight.
6. The cancer clinical stage prediction method based on VAE-transformer according to claim 5, characterized in that: The step of performing weighted summation on the feature elements in the low-dimensional potential feature representation according to the global dependency weight to obtain a weighted feature representation includes: Normalizing the global dependency weight to obtain a correlation weight; The correlation weight and the value vector are weighted and summed to obtain a weighted feature representation.
7. The cancer clinical stage prediction method based on VAE-transformer according to claim 1, characterized in that: Inputting the high-dimensional feature representation into a classification model to obtain a cancer clinical stage prediction result includes: The classification model is a multi-layer perceptron. The high-dimensional feature representation is input into the multi-layer perceptron and processed by a fully connected layer and an activation function to obtain a cancer clinical staging prediction result.
8. A cancer clinical staging prediction system based on VAE-transformer, characterized by: The system comprises: an acquisition module, the acquisition module being used to acquire gene expression data of a target patient and preprocess the gene expression data to obtain initial data; An extraction module, configured to extract features from the initial data using a variational autoencoder to obtain a low-dimensional latent feature representation; A conversion module, configured to perform feature conversion on the low-dimensional potential feature representation through a Transformer network to obtain a high-dimensional feature representation; A prediction module is used to input the high-dimensional feature representation into a classification model to obtain a cancer clinical staging prediction result.
9. A cancer clinical staging prediction device based on VAE-transformer, characterized in that: The device comprises: a memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the VAE-transformer-based cancer clinical staging prediction method according to any one of claims 1 to 7.
10. A computer storage medium, characterized in that The computer storage medium stores computer instructions, which, when called, are used to execute the VAE-transformer-based cancer clinical staging prediction method according to any one of claims 1 to 7.