Chest tumor prognosis prediction method and system based on data analysis
Through multimodal data acquisition and preprocessing, combined with graph convolution network and multi-layer perceptron algorithm, the degree of infiltration of tumor immune cells and tumor prognosis of tumor immune cells is predicted, and the problem of relying on a single data source and lack of flexibility in the existing technology is solved, personalized and dynamic tumor prognosis analysis is achieved, and treatment effect and survival rate are improved.
Patent Information
- Application Number
- CN202510113638.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing tumor prognosis prediction methods rely mostly on a single data source, lack the ability to comprehensively consider multiple factors, fail to fully utilize the impact of immune cell infiltration on tumor prognosis, and the model lacks flexibility, and cannot combine individual differences in patients and dynamic changes in the immune microenvironment.
By collecting and preprocessing multimodal data, including medical imaging data, genomic data, clinical record data and liquid biopsy data, a graph convolutional network model is constructed to predict the degree of infiltration of tumor immune cells, and comprehensive analysis is combined with multi-layer perceptron algorithms to generate prognosis prediction results for tumor patients.
It realizes accurate prediction of the tumor immune microenvironment, provides personalized and dynamic prognostic analysis, helps clinical decision-making to provide more accurate support, and improves the effectiveness of tumor treatment and patient survival rate.
Smart Images

Figure CN120048514A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tumor prognosis prediction, and particularly to a method and system for predicting the prognosis of chest tumors based on data analysis. Background Art
[0002] Chest tumors, especially lung cancer, are one of the cancers with a relatively high fatality rate globally. With the continuous progress of medical imaging, genomics, and liquid biopsy technologies, the early diagnosis and prognosis prediction of tumors have become an important part of tumor treatment. In recent years, the role of immune cell infiltration in the tumor immune microenvironment has been gradually recognized. Research shows that the degree of immune cell infiltration is closely related to tumor metastasis, recurrence, and the effect of immunotherapy. In order to improve the accuracy and individualization of tumor treatment, integrating multi-modal data (such as imaging data, genomic data, clinical data, etc.) and combining it with immune cell infiltration prediction has become a hot research direction in current tumor research.
[0003] Although the current tumor prognosis prediction methods have made certain progress, there are still some deficiencies. Traditional tumor prognosis prediction methods mostly rely on a single data source, such as imaging data or genomic data, lacking the ability to comprehensively consider multiple factors. In addition, most existing methods fail to fully utilize the impact of immune cell infiltration on tumor prognosis and cannot provide accurate prognosis assessment. Moreover, the existing prognosis prediction models often lack flexibility and fail to combine the individual differences of patients and the dynamic changes of the immune microenvironment, thus limiting their application in clinical practice.
[0004] The purpose of the present invention is to provide a method and system for predicting the prognosis of chest tumors based on data analysis, to achieve accurate prediction of the tumor immune microenvironment, provide personalized and dynamic prognosis analysis, and provide more accurate support for clinical decision-making. Summary of the Invention
[0005] The present invention provides a method and system for predicting the prognosis of chest tumors based on data analysis.
[0006] A method for predicting the prognosis of chest tumors based on data analysis includes the following steps:
[0007] S1, data collection: Collect multi-modal data of chest tumor patients, including medical imaging data, genomic data, clinical medical record data, and liquid biopsy data;
[0008] S2, data preprocessing: Preprocess the collected multi-modal data, including denoising, normalization, and standardization processing;
[0009] S3, prediction of tumor immune cell infiltration: Based on the preprocessed multi-modal data, predict the infiltration degree of tumor immune cells and evaluate its impact on the patient's prognosis, specifically including:
[0010] S31, Infiltration degree prediction: Based on the preprocessed multi-modal data, construct a tumor immune microenvironment model to predict the infiltration degree of tumor immune cells;
[0011] S32, Prognosis impact analysis: Based on the predicted infiltration degree of immune cells, evaluate the impact of the infiltration degree on the patient's prognosis;
[0012] S4, Prognosis analysis and prediction: Combine the results of the predicted infiltration of tumor immune cells, genomic data, and clinical medical record data of the patient, and conduct comprehensive analysis through the multi-layer perceptron (MLP) algorithm to generate the prognosis prediction results of tumor patients, including the metastasis risk, recurrence probability, and immunotherapy response of the tumor.
[0013] Optionally, the data collection in S1 includes:
[0014] S11, Medical image data collection: Collect the chest tumor image data of the patient through CT scan and MRI scan, including the size, shape, location, and boundary characteristics of the tumor;
[0015] S12, Genomic data collection: Collect the genomic data of the patient through gene sequencing technologies (whole-genome sequencing, targeted genome sequencing), including the gene expression profiles, mutation data, and copy number variations (CNVs) of tumor tissues and normal tissues;
[0016] S13, Clinical medical record data collection: Collect the clinical medical record data of the patient, including the patient's age, gender, family history, past medical history, treatment history, and clinical symptoms;
[0017] S14, Liquid biopsy data collection: Collect the liquid biopsy data of the patient through liquid biopsy technologies (blood, urine, saliva), including circulating tumor DNA (ctDNA), circulating tumor cells (CTC), and exosomes.
[0018] Optionally, the data preprocessing in S2 includes:
[0019] S21, Data denoising: Use the Gaussian filtering algorithm to perform denoising processing on medical image data and genomic data to remove noise and interference signals;
[0020] S22, Data normalization: Use the min-max normalization method to perform normalization processing on the collected multi-modal data;
[0021] S23, Data standardization: Use the z-score standardization method to perform standardization processing on the collected multi-modal data.
[0022] Optionally, the tumor immune microenvironment model in S31 adopts a graph convolutional network (GCN) model, and the graph convolutional network (GCN) model includes:
[0023] S311, data conversion to graph structure: Convert multi-modal data into a graph structure, where each node represents a tumor immune microenvironment feature, and the node features include medical imaging data, genomic data, clinical record data, and liquid biopsy data. The edges in the graph represent the relationships between nodes;
[0024] S312, graph convolution operation: Introduce dynamic weight adjustment and immune cell type-specific weighting to capture complex relationships in the tumor immune microenvironment;
[0025] S313, dynamic weight adjustment: In the graph convolution operation, dynamically adjust the information propagation between different nodes, use the feature differences between nodes to update the weights of the adjacency matrix, and obtain dynamic edge weights by calculating the similarity between node features;
[0026] S314, immune cell-specific weighting: By introducing specific weighting coefficients for immune cells, assign different weights to different types of immune cells to optimize the prediction of infiltration degree;
[0027] S315, infiltration degree prediction: Through the Sigmoid activation function, process the features of each node (i.e., immune cell) to obtain the final feature vector of each node, which is used to predict the infiltration degree of immune cells.
[0028] Optionally, the prognostic impact analysis in S32 includes:
[0029] S321, evaluation of the impact of immune cell infiltration on prognosis: Based on the results of infiltration degree prediction, calculate the prognostic risk score Risk Score of the patient;
[0030] S322, impact evaluation classification: Based on the calculated prognostic risk score Risk Score, compare it with the preset risk threshold T th to classify the impact of immune cell infiltration on the prognosis of the patient. When Risk Score > T th it is considered that immune cell infiltration has an impact on prognosis, and when Risk Score ≤ T th it is considered that immune cell infiltration has no impact on prognosis.
[0031] Optionally, the prognostic analysis and prediction in S4 include:
[0032] S41, data integration: According to the results of tumor immune cell infiltration prediction, integrate the results of infiltration degree prediction, genomic data, and clinical record data to generate comprehensive data;
[0033] S42, Prognosis result generation: Based on the generated comprehensive data, use the multi-layer perceptron (MLP) algorithm to predict the comprehensive data and generate the patient's prognosis results, including the metastasis risk of the tumor, the recurrence probability, and the immunotherapy response;
[0034] S43, Prognosis result evaluation: Evaluate the patient's prognosis results and classify the patient's risk levels, including low risk, medium risk, and high risk.
[0035] Optionally, the data integration in S41 includes:
[0036] S411, Determine whether the immune cell infiltration prediction value has an impact: By evaluating and classifying the degree of immune cell infiltration in the patient, determine whether to include the infiltration prediction value in the data integration. If it is determined that there is an impact, fuse the immune cell infiltration prediction value with the genomic data and clinical medical record data. If it is determined that there is no impact, only fuse the genomic data and clinical medical record data;
[0037] S412, Data fusion: Use the weighted average method to perform data fusion according to the result of determining whether the immune cell infiltration prediction value has an impact, and generate the comprehensive data D 综合 。
[0038] Optionally, the prognosis result generation in S42 includes:
[0039] S421, Input layer: Input the comprehensive data D 综合 as the input data into the input layer of the multi-layer perceptron algorithm;
[0040] S422, Hidden layer: Perform a non-linear transformation on the input data of the input layer through the hidden layer, and gradually extract the high-order features in the input data using the ReLU activation function;
[0041] S423, Output layer: Generate the final prognosis results through the output layer, including the metastasis risk of the tumor, the recurrence probability, and the probability of immunotherapy response;
[0042] S424, Loss function: Use the cross-entropy loss function to measure the difference between the prediction result of the multi-layer perceptron algorithm and the true label.
[0043] Optionally, the prognosis result evaluation in S43 includes:
[0044] S431, Risk level classification: According to the generated patient's prognosis results, use the weighted scoring method to perform weighted scoring on each prognosis factor (tumor metastasis risk, recurrence probability, and immunotherapy response), and calculate the comprehensive risk score R score ;
[0045] S432. Risk level classification: The comprehensive risk score R score is compared with the upper limit T high of the risk score threshold and the lower limit T low of the risk score threshold to classify the patient's risk level. When R score ≤T low , it indicates that the patient's risk level is low risk. When T lo < R score ≤T high , it indicates that the patient's risk level is medium risk. When R score >T high , it indicates that the patient's risk level is high risk;
[0046] S433. Result output: According to the classified risk levels, personalized treatment suggestions and management plans are provided for each patient. For high-risk patients, it is recommended to increase follow-up and monitoring. For low-risk patients, it is recommended to reduce the monitoring frequency.
[0047] A chest tumor prognosis prediction system based on data analysis, which is used to implement the above-mentioned chest tumor prognosis prediction method based on data analysis, includes the following modules:
[0048] Data acquisition module: Collect multi-modal data of chest tumor patients, including medical image data, genomic data, clinical record data, and liquid biopsy data;
[0049] Data preprocessing module: Preprocess the collected multi-modal data, including denoising, normalization, and standardization processing;
[0050] Tumor immune cell infiltration prediction module: Based on the preprocessed multi-modal data, predict the infiltration degree of tumor immune cells and evaluate its impact on the patient's prognosis;
[0051] Prognosis analysis and prediction module: Combine the results of the patient's tumor immune cell infiltration prediction, genomic data, and clinical record data, and conduct comprehensive analysis through the multi-layer perceptron algorithm to generate the prognosis prediction results of tumor patients, including the metastasis risk, recurrence probability, and immunotherapy response of the tumor.
[0052] Advantages of the present invention:
[0053] In the present invention, through the comprehensive collection of multi-modal data, including medical image data, genomic data, clinical record data, and liquid biopsy data, various aspects of the tumor can be accurately captured, such as the morphology of the tumor, gene mutations, the patient's clinical background, and the dynamic changes of the tumor. Through the denoising, normalization, and standardization processing of the collected data, the consistency between different data sources is ensured, providing high-quality and standardized input data for subsequent analysis.
[0054] In the present invention, by constructing a graph convolutional network model, the infiltration degree of tumor immune cells is accurately predicted, and based on this, its impact on the prognosis of patients is evaluated. Through the dynamic prediction of the infiltration degree of immune cells and the specific weighting of immune cells, the complexity of the tumor immune microenvironment can be better captured, thereby providing a more accurate prognosis assessment to help clinicians understand the potential impact of immune responses on tumor metastasis, recurrence, and immune therapy response.
[0055] In the present invention, through comprehensive analysis using the multi-layer perceptron algorithm, accurate prognosis prediction is provided for tumor patients, and the risk of tumor metastasis, recurrence probability, and immune therapy response can be calculated, thereby helping doctors formulate personalized treatment plans according to the specific conditions of each patient. Through the hierarchical assessment of risk scores, targeted treatment recommendations can be provided for patients with different risk levels, greatly improving the effect of tumor treatment and the survival rate of patients. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0057] Figure 1 It is a schematic flow chart of the prediction method according to an embodiment of the present invention;
[0058] Figure 2 It is a schematic diagram of the system function modules according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] The present invention will be described in detail below in conjunction with the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.
[0060] It should be pointed out that in the specification, it is mentioned that "an embodiment", "embodiments", "exemplary embodiments", "some embodiments", etc. indicate that the described embodiments may include specific features, structures, or characteristics, but not necessarily every embodiment includes such specific features, structures, or characteristics. Additionally, when combining embodiments to describe specific features, structures, or characteristics, implementing such features, structures, or characteristics in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.
[0061] Generally, terms can be understood, at least in part, from their use in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or property in a singular sense, or can be used to describe a combination of features, structures, or properties in a plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but rather can alternatively, depending at least in part on the context, allow for the existence of other factors that are not necessarily explicitly described.
[0062] As Figure 1 shown, a method for predicting the prognosis of chest tumors based on data analysis includes the following steps:
[0063] S1, Data collection: Collect multi-modal data of chest tumor patients, including medical imaging data, genomic data, clinical record data, and liquid biopsy data;
[0064] S2, Data preprocessing: Preprocess the collected multi-modal data, including denoising, normalization, and standardization;
[0065] S3, Prediction of tumor immune cell infiltration: Based on the preprocessed multi-modal data, predict the infiltration degree of tumor immune cells and evaluate its impact on the patient's prognosis, specifically including:
[0066] S31, Prediction of infiltration degree: Based on the preprocessed multi-modal data, construct a tumor immune microenvironment model to predict the infiltration degree of tumor immune cells;
[0067] S32, Analysis of prognosis impact: Based on the predicted infiltration degree of immune cells, evaluate the impact of the infiltration degree on the patient's prognosis;
[0068] S4, Prognosis analysis and prediction: Combine the results of the prediction of tumor immune cell infiltration of the patient, genomic data, and clinical record data, and conduct comprehensive analysis through the multi-layer perceptron (MLP) algorithm to generate the prognosis prediction results of tumor patients, including the metastasis risk, recurrence probability, and immune therapy response of the tumor;
[0069] Through the above content, it is possible to effectively predict the prognosis results of tumor patients, introduce the prediction of the tumor immune microenvironment, comprehensively consider the impact of immune cell infiltration on the patient's metastasis risk, recurrence probability, and immune therapy response, provide more accurate personalized prognosis analysis, thereby providing strong support for clinical decision-making, and having high clinical application value and promotion potential.
[0070] The data collection in S1 includes:
[0071] S11, Collection of medical imaging data: Collect chest tumor imaging data of the patient through CT scan and MRI scan, including the size, shape, location, and boundary characteristics of the tumor;
[0072] S12, Genome data collection: Through gene sequencing technologies (whole-genome sequencing, targeted genome sequencing), collect the patient's genome data, including gene expression profiles, mutation data, and copy number variations (CNVs) of tumor tissues and normal tissues;
[0073] S13, Clinical medical record data collection: Collect the patient's clinical medical record data, including the patient's age, gender, family history, past medical history, treatment history, and clinical symptoms;
[0074] S14, Liquid biopsy data collection: Through liquid biopsy technologies (blood, urine, saliva), collect the patient's liquid biopsy data, including circulating tumor DNA (ctDNA), circulating tumor cells (CTC), and exosomes;
[0075] Through the above, not only can the morphology, genetic characteristics, pathological background, and dynamic changes of tumors be accurately captured, but also non-invasive and real-time monitoring data can be provided through liquid biopsy. This comprehensive data collection method helps to comprehensively evaluate the characteristics of tumors and the individual differences of patients, thus providing a more accurate and personalized analysis basis for subsequent tumor prognosis prediction.
[0076] The data preprocessing in S2 includes:
[0077] S21, Data denoising: Use the Gaussian filtering algorithm to denoise medical image data and genome data, removing noise and interference signals, expressed as:
[0078]
[0079] where I(x+i,y+j) is the pixel value at position (x+i,y+j), w(i,j) is the weight of the Gaussian filter, k is the size of the filter, and I denoise (x,y) is the denoised medical image data;
[0080] S22, Data normalization: Use the min-max normalization method to normalize the collected multi-modal data to ensure that the features between different data sources have the same dimension, expressed as:
[0081]
[0082] where X is the original multi-modal data, X min and X max are the minimum and maximum values of the multi-modal data set respectively, and X ′ is the normalized multi-modal data, ensuring that the range of all data is scaled to the [0,1] range;
[0083] S23, Data Standardization: The z-score standardization method is used to standardize the collected multi-modal data, which is expressed as:
[0084]
[0085] where X is the original multi-modal data, μ is the mean of the multi-modal data, σ is the standard deviation of the multi-modal data, and X standardized is the standardized multi-modal data;
[0086] Through the above content, the consistency and quality of multi-modal data can be effectively improved, the scale differences and noise interferences between different data sources can be eliminated, and the unity of medical images, genomes, clinical records, and liquid biopsy data is ensured, thus avoiding the impact of data inconsistency on subsequent analysis.
[0087] The tumor immune microenvironment model in S31 adopts a graph convolutional network (GCN) model, and the graph convolutional network (GCN) model includes:
[0088] S311, Data Conversion to Graph Structure: Convert the multi-modal data into a graph structure, where each node represents a tumor immune microenvironment feature, and the node features include medical image data, genomic data, clinical record data, and liquid biopsy data. The edges in the graph represent the relationships between nodes, which is expressed as:
[0089]
[0090] where is the initial feature vector of node i, M is the number of modalities, and X i,m is the feature of the i-th node in the m-th modality, is the weighted coefficient of the m-th modality feature;
[0091]
[0092] where A ij is the edge weight between node i and node j, are the initial feature vectors of node i and node j respectively, and f is a function (cosine similarity) used to calculate the similarity between nodes;
[0093]
[0094] where h i and h j are the feature vectors of node i and node j at the current layer respectively, h i ·h j represents the dot product of the feature vectors of node i and node j, and ‖h i ‖ 2 and ‖h j‖ 2 They are the L2 norms of the feature vectors of node i and node j respectively, that is, the modulus lengths of the feature vectors;
[0095] S312, Graph convolution operation: Update the features of each node through graph convolution operation. The update of each node depends not only on its own features but also on the features of neighboring nodes. The adjacency matrix is used to determine the connection relationship between nodes. Introduce dynamic weight adjustment and immune cell type-specific weighting to capture the complex relationships in the tumor immune microenvironment, expressed as:
[0096]
[0097] where, is the feature of node i at the k+1 layer, N(i) is the set of neighboring nodes of node i, is the dynamic edge weight between node i and node j, d j is the degree of node j (i.e., the number of edges connected to node j), is the feature of node j at the k layer, W (k) is the weight matrix at the k layer, σ(·) is the ReLU activation function;
[0098] S313, Dynamic weight adjustment: In the graph convolution operation, dynamically adjust the information propagation between different nodes, use the feature differences between nodes to update the weights of the adjacency matrix, and obtain dynamic edge weights by calculating the similarity between node features, so as to more accurately reflect the relationship between nodes, expressed as:
[0099]
[0100] where, is the dynamic edge weight between node i and node j, are the features of node i and node j at the k layer respectively, ‖·‖ 2 is the Euclidean distance, used to measure the similarity between node features, and δ is the standard deviation parameter to control the sensitivity of the similarity;
[0101] S314, Immune cell-specific weighting: Since different immune cell types have different effects on the infiltration degree of tumors, by introducing the specific weighting coefficients of immune cells, different weights are assigned to different types of immune cells to optimize the prediction of the infiltration degree, expressed as:
[0102]
[0103] where, is the feature of node i at the k+1 layer, is the dynamic edge weight between node i and node j, w immuneis the weighted coefficient of immune cell types, representing the weight of this type of immune cell. is the feature of node j in the k-th layer, W (k) is the weight matrix of the k-th layer, and σ(·) is the ReLU activation function;
[0104] S315, infiltration degree prediction: Through the Sigmoid activation function, the features of each node (i.e., immune cell) are processed to obtain the final feature vector of each node, which is used to predict the infiltration degree of immune cells, expressed as:
[0105]
[0106] where, I infiltration is the predicted value of immune cell infiltration of node i, representing the probability or intensity of infiltration, W final is the weight matrix of the last layer, b final is the bias term of the last layer, and Sigmoid is the activation function, mapping the prediction result to the range of [0,1];
[0107] Through the above content, introducing a dynamic adjacency matrix based on cosine similarity can more accurately capture the relationships and similarities between different cells in the tumor immune microenvironment, thereby effectively improving the accuracy of tumor immune cell infiltration prediction. It can not only dynamically adjust the connection strength between nodes according to the characteristics of the tumor immune microenvironment, but also strengthen the influence of key immune cells through an adaptive similarity scale factor, avoiding the limitations brought by the static adjacency matrix, enhancing the prediction ability of the infiltration degree of tumor immune cells, and further providing a more accurate basis for the prognosis analysis and treatment plan formulation of chest tumors.
[0108] The prognostic impact analysis in S32 includes:
[0109] S321, evaluation of the impact of immune cell infiltration on prognosis: Based on the results of infiltration degree prediction, calculate the prognostic risk score Risk Score of the patient, expressed as:
[0110] RiskScore = w 1 ·I infiltration ;
[0111] where, Risk Score is the prognostic risk score of the patient, w 1 is the weighted coefficient of the immune cell infiltration degree on the prognostic risk score;
[0112] S322, impact assessment classification: Based on the calculated prognostic risk score Risk Score, compare it with the preset risk threshold T th for comparison, classify the impact of the patient's immune cell infiltration on prognosis. When Risk Score > Tth When it is, it is considered that immune cell infiltration has an impact on prognosis. When Risk Score ≤ T th When it is, it is considered that immune cell infiltration has no impact on prognosis;
[0113] Risk threshold T th is set based on historical data, specifically including:
[0114] Data collection and historical data preparation: Collect data of chest tumor patients with known prognosis information, including the degree of immune cell infiltration, clinical prognosis results (such as survival period, recurrence risk, etc.);
[0115] Patient grouping: According to the prognosis information in the historical data, use the Kaplan-Meier survival analysis method to divide the patients into a low-risk group and a high-risk group.
[0116] Analysis of the relationship between immune cell infiltration and prognosis: Analyze the relationship between the degree of immune cell infiltration and the prognosis of patients, use the survival analysis method to find the boundary value between the infiltration degree and the prognosis risk;
[0117] Threshold setting: Based on the analysis results, determine the optimal threshold T of immune cell infiltration by maximizing the difference in survival curves th , expressed as:
[0118]
[0119] Through the above content, not only the data processing flow is simplified, but also the prognosis of patients can be quantitatively analyzed through precise weight coefficient and threshold setting, ensuring that the prognosis judgment is more objective and consistent. At the same time, as a key indicator of tumor immune escape and treatment response, immune cell infiltration makes this analysis have high practical application value in tumor immunotherapy and recurrence risk prediction, improving the accuracy of the formulation of individualized treatment plans.
[0120] The prognosis analysis and prediction in S4 include:
[0121] S41, Data integration: According to the results of tumor immune cell infiltration prediction, integrate the results of infiltration degree prediction, genomic data, and clinical medical record data to generate comprehensive data;
[0122] S42, Prognosis result generation: Based on the generated comprehensive data, use the multi-layer perceptron (MLP) algorithm to predict the comprehensive data and generate the prognosis results of patients, including the metastasis risk of tumors, recurrence probability, and immune therapy response;
[0123] S43, Prognosis Outcome Assessment: Assess the prognosis outcome of patients, classify the patient risk levels, including low risk, medium risk, and high risk, to help clinicians formulate personalized treatment plans according to the specific risks of patients and improve the treatment effect;
[0124] Through the above content, the health status of patients can be comprehensively and accurately evaluated. After generating the prognosis outcome, by using the multi-layer perceptron algorithm, key indicators such as the risk of tumor metastasis, recurrence probability, and immunotherapy response can be accurately predicted, providing decision support for personalized treatment for clinicians. Based on the assessment of the prognosis outcome, the patient risk levels can be further classified, effectively helping doctors formulate specific treatment plans for patients with different risk levels, thereby improving the accuracy and effect of treatment and optimizing the overall treatment path of patients.
[0125] The data integration in S41 includes:
[0126] S411, Judge whether the immune cell infiltration prediction value has an impact: Through the impact assessment and classification of the immune cell infiltration degree of patients, judge whether to include the infiltration prediction value in the data integration. If it is judged to have an impact, fuse the immune cell infiltration prediction value with the genomic data and clinical medical record data. If it is judged to have no impact, only integrate the genomic data and clinical medical record data, expressed as:
[0127]
[0128] S412, Data fusion: Adopt the weighted average method to perform data fusion according to the result of judging whether the immune cell infiltration prediction value has an impact, and generate the comprehensive data D 综合 , expressed as:
[0129]
[0130] Among them, I infiltration is the immune cell infiltration prediction value, D 基因组数据 and D 临床病历数据 are the genomic data and clinical medical record data respectively, and w 2 , w 3 , w 4 are the weights of each data source respectively;
[0131] Based on the above, according to the prediction results of tumor immune cell infiltration, combined with genomic data and clinical record data, comprehensive data is generated, thus providing more comprehensive and accurate information for prognosis analysis. Through the judgment based on impact assessment classification, the predicted value of immune cell infiltration is only added when the infiltration degree affects the prognosis of the patient. This selective integration can effectively reduce the interference of redundant data, improve the prediction efficiency and accuracy of the model. At the same time, it can combine the advantages of multiple data sources, better reflect the overall health status of the patient and the biological characteristics of the tumor, and further improve the reliability and precision of prognosis analysis.
[0132] The generation of the prognosis result in S42 includes:
[0133] S421, input layer: Input the comprehensive data D 综合 as the input data into the input layer of the multi-layer perceptron algorithm. Let the comprehensive data vector be X = {x 1 , x 2 ,..., x n}, where x i represents each feature (clinical record, genomic information, predicted value of immune cell infiltration);
[0134] S422, hidden layer: Perform a non-linear transformation on the input data of the input layer through the hidden layer, and gradually extract the high-order features in the input data using the ReLU activation function. The output of the hidden layer is expressed as:
[0135] h = f(WX + b);
[0136] Among them, h is the output of the hidden layer, W is the weight matrix, b is the bias, and f is the ReLU activation function, that is, f(x) = max(0, x);
[0137] S423, output layer: Generate the final prognosis result through the output layer, including the metastasis risk of the tumor, the recurrence probability, and the probability of immune therapy response. Use the Sigmoid function to map the output to the probability space. Let the output be Y = {y 1 , y 2 , y 3}, where y 1 is the risk probability of tumor metastasis, y 2 is the recurrence probability of the tumor, y 3 is the probability of immune therapy response. The output layer is expressed as:
[0138]
[0139] Among them, W h is the weight matrix of the output layer, b h is the bias of the output layer, is the Sigmoid activation function, that is
[0140] S424, Loss function: The cross - entropy loss function is used to measure the difference between the prediction results of the multi - layer perceptron algorithm and the true labels. Let the true labels be \(T = \{t 1 , t 2 , t 3 \}\), then the loss function is expressed as:
[0141]
[0142] where \(y i \) is the probability value predicted by the multi - layer perceptron algorithm, and \(t i \) is the true label;
[0143] Through the above content, it is possible to accurately synthesize various clinical, genomic, and immune cell infiltration prediction results, achieve comprehensive and accurate risk assessment. The MLP model can automatically extract hidden features from complex multi - modal data, capture the potential correlations between different data sources, thereby providing more accurate predictions of tumor metastasis risk, recurrence probability, and immunotherapy response. It not only improves the accuracy of prediction but also provides personalized treatment guidance for clinicians, helps optimize treatment decisions, and improves the treatment effect and survival expectancy of patients.
[0144] The prognostic result evaluation in S43 includes:
[0145] S431, Risk level classification: According to the generated prognostic results of the patient, the weighted scoring method is used to perform weighted scoring on each prognostic factor (tumor metastasis risk, recurrence probability, and immunotherapy response), and the comprehensive risk score \(R score \) is calculated, expressed as:
[0146] R score = w 5 ×T score + w 6 ×P score + w 7 ×I score ;
[0147] where \(T score \) is the tumor metastasis risk score, \(P score \) is the recurrence probability score, \(I score \) is the immunotherapy response score, and \(w 5 , w 6 , w 7 \) are the weights of each score respectively, satisfying \(w 5 + w 6 + w 7 = 1;
[0148] S432, Risk level classification: Compare the comprehensive risk score R score with the upper limit T high of the risk score threshold and the lower limit T low of the risk score threshold to classify the patient's risk level. When R score ≤T low , it indicates that the patient's risk level is low risk. When T low <R score ≤T high , it indicates that the patient's risk level is medium risk. When R score >T high , it indicates that the patient's risk level is high risk;
[0149] The upper limit T high of the risk score threshold and the lower limit T low of the risk score threshold are set through historical data, specifically including:
[0150] Collect and process historical data: Collect the historical data of patients and calculate the comprehensive risk score R score ;
[0151] Calculate the mean and standard deviation: Conduct statistical analysis on the comprehensive risk scores R score of all patients, and calculate the mean and standard deviation of this dataset, expressed as:
[0152]
[0153] where N is the total number of samples, Rscore i is the risk score of the i-th patient, and μ is the mean;
[0154]
[0155] where θ is the standard deviation;
[0156] Set the risk score threshold: Based on the calculated mean and standard deviation, set the upper limit T high of the risk score threshold and the lower limit T low of the risk score threshold, expressed as:
[0157] T high = μ + θ;
[0158] T low = μ - θ;
[0159] S433, Result output: Provide personalized treatment recommendations and management plans for each patient according to the classified risk levels. For high-risk patients, it is recommended to increase follow-up and monitoring. For low-risk patients, it is recommended to reduce the monitoring frequency;
[0160] Through the above, the accurate classification of patient risks can help doctors formulate personalized treatment strategies according to different risk levels, improve treatment effects and reduce unnecessary treatment interventions. In addition, the thresholds set based on statistical analysis provide a quantifiable and standardized basis for the management of patient groups, which helps to conduct consistency evaluation and decision-making among different medical institutions.
[0161] As Figure 2 shown, a prognostic prediction system for chest tumors based on data analysis, which is used to implement the above-mentioned prognostic prediction method for chest tumors based on data analysis, includes the following modules:
[0162] Data acquisition module: Collect multi-modal data of chest tumor patients, including medical image data, genomic data, clinical record data, and liquid biopsy data;
[0163] Data preprocessing module: Preprocess the collected multi-modal data, including denoising, normalization, and standardization;
[0164] Tumor immune cell infiltration prediction module: Based on the preprocessed multi-modal data, predict the infiltration degree of tumor immune cells and evaluate its impact on the patient's prognosis;
[0165] Prognosis analysis and prediction module: Combine the results of tumor immune cell infiltration prediction, genomic data, and clinical record data of the patient, and conduct comprehensive analysis through the multi-layer perceptron algorithm to generate the prognostic prediction results of tumor patients, including the metastasis risk, recurrence probability, and immunotherapy response of the tumor.
[0166] The present invention covers any alternatives, modifications, equivalent methods, and solutions made on the essence and scope of the present invention. In order to enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention without the description of these details. In addition, well-known methods, processes, procedures, components, and circuits are not described in detail to avoid unnecessary confusion to the essence of the present invention.
[0167] The above is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for predicting the prognosis of thoracic tumors based on data analysis, characterized in that: The following steps are involved: S1, data collection: collect multimodal data of patients with thoracic tumors, including medical imaging data, genomic data, clinical medical record data, and liquid biopsy data; S2, data preprocessing: preprocessing the collected multimodal data, including denoising, normalization and standardization; S3, Tumor immune cell infiltration prediction: Based on the pre-processed multimodal data, predict the infiltration degree of tumor immune cells and evaluate its impact on patient prognosis, including: S31, prediction of infiltration degree: based on the pre-processed multimodal data, a tumor immune microenvironment model is constructed to predict the infiltration degree of tumor immune cells; S32, prognostic impact analysis: based on the predicted degree of immune cell infiltration, evaluate the impact of the infiltration degree on the patient's prognosis; S4, Prognostic analysis and prediction: Combining the patient's tumor immune cell infiltration prediction results, genomic data and clinical medical record data, a comprehensive analysis is performed through a multi-layer perceptron algorithm to generate prognostic prediction results for tumor patients, including the tumor's metastasis risk, recurrence probability and immunotherapy response.
2. A method for predicting the prognosis of a thoracic tumor based on data analysis according to claim 1, characterized in that: The data collection in S1 includes: S11, medical imaging data acquisition: Collect the patient's chest tumor imaging data through CT scan and MRI scan, including the size, shape, location, and boundary characteristics of the tumor; S12, genomic data collection: Collect the patient's genomic data through gene sequencing technology, including gene expression profiles, mutation data, and copy number variations of tumor tissues and normal tissues; S13, clinical medical record data collection: collect the patient's clinical medical record data, including the patient's age, gender, family history, past medical history, treatment history, and clinical symptoms; S14, Liquid biopsy data collection: Collect patients' liquid biopsy data through liquid biopsy technology, including circulating tumor DNA, circulating tumor cells, and exosomes.
3. The method for predicting the prognosis of a thoracic tumor based on data analysis according to claim 1, characterized in that: The data preprocessing in S2 includes: S21, data denoising: Gaussian filtering algorithm is used to denoise medical imaging data and genomic data to remove noise and interference signals; S22, data normalization: using the minimum-maximum normalization method to normalize the collected multimodal data; S23, data standardization: The collected multimodal data were standardized using the z-score standardization method.
4. The method for predicting the prognosis of a thoracic tumor based on data analysis according to claim 1, characterized in that: The tumor immune microenvironment model in S31 adopts a graph convolutional network model, and the graph convolutional network model includes: S311, data conversion into graph structure: Convert multimodal data into a graph structure, where each node represents a tumor immune microenvironment feature. Node features include medical imaging data, genomic data, clinical medical record data, and liquid biopsy data. The edges in the graph represent the relationship between nodes. S312, graph convolution operation: introduces dynamic weight adjustment and immune cell type-specific weighting to capture the complex relationships in the tumor immune microenvironment; S313, dynamic weight adjustment: In the graph convolution operation, the information propagation between different nodes is dynamically adjusted, the feature differences between nodes are used to update the weights of the adjacency matrix, and the dynamic edge weights are obtained by calculating the similarity between node features; S314, immune cell specific weighting: by introducing the specific weighting coefficient of immune cells, different types of immune cells are given different weights to optimize the prediction of infiltration degree; S315, infiltration degree prediction: The features of each node are processed through the Sigmoid activation function to obtain the final feature vector of each node, which is used to predict the infiltration degree of immune cells.
5. A method for predicting the prognosis of a thoracic tumor based on data analysis according to claim 4, characterized in that: The prognostic impact analysis in S32 includes: S321, Evaluation of the impact of immune cell infiltration on prognosis: Based on the results of the infiltration degree prediction, calculate the patient's prognostic risk score Risk Score; S322, Impact Assessment Classification: Based on the calculated prognostic risk score Risk Score and the preset risk threshold T th For comparison, the effect of immune cell infiltration on prognosis was classified. When Risk Score>T th When Risk Score≤T th When the number of immune cells infiltration is less than 20%, it is considered that immune cell infiltration has no effect on prognosis.
6. A method for predicting the prognosis of a breast tumor based on data analysis according to claim 5, characterized in that: The prognostic analysis and prediction in S4 include: S41, data integration: based on the results of tumor immune cell infiltration prediction, the results of infiltration degree prediction, genomic data and clinical medical record data were integrated to generate comprehensive data; S42, prognosis result generation: based on the generated comprehensive data, a multi-layer perceptron algorithm is used to predict the comprehensive data and generate the patient's prognosis, including the risk of tumor metastasis, probability of recurrence, and response to immunotherapy; S43, prognosis evaluation: Evaluate the patient's prognosis and classify the patient's risk level, including low risk, medium risk and high risk.
7. A method for predicting the prognosis of a thoracic tumor based on data analysis according to claim 6, characterized in that: The data integration in S41 includes: S411, determine whether the immune cell infiltration prediction value has an impact: by classifying the patient's immune cell infiltration degree, determine whether to include the infiltration prediction value in the data integration. If it is determined to have an impact, the immune cell infiltration prediction value is integrated with the genome data and clinical medical record data. If it is determined to have no impact, only the genome data and clinical medical record data are integrated. S412, data fusion: Use the weighted average method to perform data fusion based on the results of judging whether the predicted value of immune cell infiltration has an impact, and generate comprehensive data D 综合 .
8. The method for predicting the prognosis of a thoracic tumor based on data analysis according to claim 7, characterized in that: The generation of the prognosis result in S42 includes: S421, input layer: the comprehensive data D 综合 As input data to the input layer of the multi-layer perceptron algorithm; S422, hidden layer: the input data of the input layer is transformed nonlinearly through the hidden layer, and the high-order features in the input data are gradually extracted using the ReLU activation function; S423, output layer: the final prognostic results are generated through the output layer, including the risk of tumor metastasis, probability of recurrence, and probability of immunotherapy response; S424, Loss function: The cross entropy loss function is used to measure the difference between the prediction results of the multilayer perceptron algorithm and the true label.
9. The method for predicting the prognosis of a thoracic tumor based on data analysis according to claim 8, characterized in that: The prognostic outcome assessment in S43 includes: S431, Risk level classification: Based on the generated patient prognosis results, use the weighted scoring method to weight each prognostic factor and calculate the comprehensive risk score R score ; S432, Risk level classification: The comprehensive risk score R score and the upper threshold of risk score T high , Risk score threshold lower limit T low Compare and classify patients into risk levels. score ≤T low When T low <R score ≤T high When R score >T high When , it means the patient’s risk level is high risk; S433, result output: Provide personalized treatment recommendations and management plans for each patient based on the risk level. For high-risk patients, it is recommended to increase follow-up and monitoring, and for low-risk patients, it is recommended to reduce the monitoring frequency.
10. A breast tumor prognosis prediction system based on data analysis, used to implement a breast tumor prognosis prediction method based on data analysis as claimed in any one of claims 1 to 9, characterized in that: Includes the following modules: Data collection module: collects multimodal data of chest tumor patients, including medical imaging data, genomic data, clinical medical record data, and liquid biopsy data; Data preprocessing module: preprocess the collected multimodal data, including denoising, normalization and standardization; Tumor immune cell infiltration prediction module: Based on pre-processed multimodal data, predict the degree of tumor immune cell infiltration and evaluate its impact on patient prognosis; Prognostic analysis and prediction module: Combines the patient's tumor immune cell infiltration prediction results, genomic data and clinical medical record data, and performs comprehensive analysis through a multi-layer perceptron algorithm to generate prognostic prediction results for tumor patients, including the tumor's metastasis risk, recurrence probability and immunotherapy response.
Citation Information
Cited By
Postoperative immune state monitoring and prognosis evaluation method for patient with thymoma
CN120496856A
Tumor recurrence risk prediction method and system based on electronic medical record data
CN120544908A
Early warning method and system for pheochromocytoma and paraganglioma
CN120895241A