A tumor prognosis prediction method based on text recognition of tumor databases

By integrating multimodal data and using graph neural networks and multi-layer neural network models, the problem of data integration and semantic correlation in tumor prognosis prediction is solved, and high-accurate tumor metastasis path prediction and personalized risk assessment are achieved, providing a scientific basis for tumor treatment.

CN119480125BActive Publication Date: 2025-05-27THE FIRST AFFILIATED HOSPITAL OF XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510074327.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-27
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently integrate multimodal data in tumor prognosis prediction, and traditional methods are difficult to capture semantic associations between tumor type, treatment regimen and metastasis location, resulting in insufficient prediction accuracy.

Method used

By collecting clinical text, image and genomic data from multiple tumor databases, preprocessing and text information extraction, the graph structure is constructed using the graph neural network model, the metastasis path rules are extracted, and prognostic evaluation and risk prediction are combined with multi-layer neural networks.

Benefits of technology

It significantly improves the accuracy and interpretability of tumor metastasis path prediction, provides quantitative five-year survival rate and recurrence risk prediction results, realizes accurate grading and personalized management, and provides a scientific basis for the formulation of tumor treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119480125B_ABST
    Figure CN119480125B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of medical care, and particularly relates to a tumor prognosis prediction method based on tumor database text recognition, comprising the following steps: S1, data collection and preprocessing: collecting tumor-related medical data from multiple tumor databases; S2, text information extraction: extracting information from clinical text data; S3, tumor metastasis path prediction: using a graph neural network model to predict the metastasis path, predicting the potential metastasis path of the tumor, and analyzing the impact of different metastasis paths on the prognosis of patients; S4, prognosis evaluation and risk prediction: combining the medical data of patients to predict the tumor prognosis risk, and comprehensively evaluating the predicted value of the five-year survival rate and the predicted value of the recurrence risk of patients. The present invention can assist medical institutions in optimizing resource allocation, formulating scientific personalized diagnosis and treatment plans, and significantly improving the survival quality and treatment effect of patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of healthcare, and particularly to a tumor prognosis prediction method based on tumor database text recognition. Background Art

[0002] With the rapid accumulation of medical data and the continuous development of artificial intelligence technology, tumor prognosis prediction based on big data has become one of the hotspots in precision medicine research. The prognosis assessment of tumor patients is of great significance for the formulation of treatment plans and the long-term management of patients. By analyzing the clinical text records, imaging data, and genomic data of patients, combined with data modeling techniques, not only can the five-year survival rate of patients be predicted, but also the recurrence risk can be evaluated. However, due to the diverse sources and heterogeneity of medical data, how to efficiently integrate multi-modal data and construct an accurate prediction model has become a key problem in current medical artificial intelligence research.

[0003] The existing technologies face the following main problems in the field of tumor prognosis prediction: First, there are differences in formats and feature expressions among multi-modal data (such as clinical text, imaging, and genomic data), and traditional data preprocessing methods have limited effects in eliminating noise information and retaining important features; Second, the prediction of tumor metastasis pathways depends on complex heterogeneous relationship modeling, and traditional methods are difficult to effectively capture the semantic associations among tumor types, treatment plans, and metastasis locations, resulting in insufficient prediction accuracy; In addition, existing prognosis assessment methods often fail to fully consider the comprehensive effects of individual characteristics and metastasis pathways, and it is difficult to provide targeted risk grading suggestions, which limits their application value in clinical practice.

[0004] The purpose of the present invention is to provide a tumor prognosis prediction method based on tumor database text recognition, which not only improves the accuracy and interpretability of tumor metastasis pathway prediction, but also can provide quantified five-year survival rate and recurrence risk prediction results for patients, further realizing precise grading and personalized management, and providing a scientific basis for the formulation of tumor treatment plans. Summary of the Invention

[0005] The present invention provides a tumor prognosis prediction method based on tumor database text recognition.

[0006] A tumor prognosis prediction method based on tumor database text recognition includes the following steps:

[0007] S1, data collection and preprocessing: Collect tumor-related medical data from multiple tumor databases, including clinical text data, imaging data, and genomic data, and perform denoising, word segmentation on the clinical text data, standardize the imaging data, and complete missing values and normalize the genomic data;

[0008] S2, Text Information Extraction: Use natural language processing techniques to extract information from clinical text data, including tumor type, treatment plan, and metastasis location;

[0009] S3, Tumor Metastasis Path Prediction: Based on the extracted tumor type, treatment plan, and metastasis location, use a graph neural network (GNN) model to predict the metastasis path, predict the potential metastasis path of the tumor, and analyze the impact of different metastasis paths on the patient's prognosis, specifically including:

[0010] S31, Atlas Construction: Based on the extracted tumor type, treatment plan, and metastasis location, construct a graph structure composed of nodes and edges;

[0011] S32, Metastasis Path Prediction: Based on the constructed graph structure, use a graph neural network (GNN) model to extract the potential rules of the tumor metastasis path and perform metastasis path prediction;

[0012] S33, Metastasis Path Evaluation: Based on the prediction results of the metastasis path, analyze the impact of different metastasis paths on the patient's prognosis;

[0013] S4, Prognosis Evaluation and Risk Prediction: Based on the results of tumor metastasis path prediction, combine the patient's medical data to predict the tumor prognosis risk and comprehensively evaluate the patient's five-year survival rate and recurrence risk.

[0014] Optionally, the data collection and preprocessing in S1 include:

[0015] S11, Data Collection: Collect tumor-related medical data from multiple tumor databases, including clinical text data, imaging data, and genomic data;

[0016] S12, Clinical Text Data Preprocessing: Denoise and tokenize the clinical text data;

[0017] S13, Imaging Data Preprocessing: Use the Z-score normalization method to normalize the imaging data;

[0018] S14, Genomic Data Preprocessing: Complete missing values and normalize the genomic data.

[0019] Optionally, the text information extraction in S2 includes:

[0020] S21, Tokenization and Text Cleaning: Tokenize the clinical text data through a maximum entropy model based on probability, and remove stop words and irrelevant characters;

[0021] S22, Named Entity Recognition: Use a bidirectional long short-term memory network (Bi-LSTM) combined with a conditional random field (CRF) model to identify entity information of tumor types, treatment plans, and metastasis locations in clinical text data;

[0022] Relationship Extraction: Based on the extracted named entities, use a relationship extraction model with an attention mechanism to identify the associations between tumor types, treatment plans, and metastasis locations.

[0023] Optionally, the atlas construction in S31 includes:

[0024] S311, Node Definition and Representation: According to the clinical text data, use the extracted tumor types, treatment plans, and metastasis locations as nodes in the graph, denoted as the node set , where represents the corresponding clinical text data, and the feature vector of each node is represented as ;

[0025] S312, Edge Definition and Weight Calculation: Construct an edge set according to the relationships between nodes , represents the association between nodes and , and calculate the weight of the edge through the relationship extraction model;

[0026] S313, Graph Structure Construction: Use the node set and the edge set to construct a graph structure . In the graph structure, nodes represent tumor types, treatment plans, or metastasis locations, and edges represent the association relationships between nodes (such as the relationship between treatment plans and tumor types, the association between tumor types and metastasis locations, etc.);

[0027] S314, Calculation of the Normalized Adjacency Matrix: Organize the weights of the edges into an adjacency matrix , and normalize the adjacency matrix.

[0028] Optionally, the metastasis path prediction in S32 includes:

[0029] S321, Input Representation: Use the node feature matrix of the constructed graph structure and the normalized adjacency matrix as the input of a graph neural network (GNN) model;

[0030] S322, Graph Convolutional Layer Calculation: Extract node features and update node representations through a graph convolutional network (GCN) model;

[0031] S323, Aggregation and Prediction: After multiple layers of graph convolution, extract the final node feature matrix represented as , where is the number of layers of the network. A fully connected layer is used to process the node features to predict the potential metastasis paths of tumors, and the prediction results are output through a classifier as the probabilities of each metastasis path;

[0032] S324, Metastasis Path Extraction and Sorting: According to the predicted probability distribution, extract the metastasis paths and sort them according to the probability size, and output the potential laws and paths of tumor metastasis.

[0033] Optionally, the metastasis path evaluation in S33 includes:

[0034] S331, Feature Extraction of Metastasis Path: According to the prediction results of the metastasis path, extract the node features related to the metastasis path from the final node feature matrix output by the graph neural network (GNN) model and aggregate all the node features on the path to generate a representation vector of the metastasis path ;

[0035] S332, Prognosis Score Calculation: Use a fully connected layer to calculate the prognosis score for the representation vector of each metastasis path ;

[0036] S333, Analysis of the Overall Impact of Paths on Prognosis: Combine all the predicted metastasis paths and their prognosis scores to calculate the overall prognosis impact value , representing the overall prognosis status of the patient.

[0037] Optionally, the prognosis evaluation and risk prediction in S4 include:

[0038] S41, Comprehensive Feature Fusion: Based on the calculated overall prognosis impact value , fuse it with the medical data of the patient to generate a comprehensive feature representation of the patient;

[0039] S42, Prognosis Risk Index Prediction: Based on the fused comprehensive feature representation of the patient, predict the prognosis risk indices of the patient through a multi-layer neural network, including the five-year survival rate and the recurrence risk;

[0040] S43, Comprehensive Evaluation and Grading Recommendation: Combine the five-year survival rate, recurrence risk, and overall prognosis impact value of the patient to comprehensively calculate the personalized prognosis evaluation value of the patient, which is used to comprehensively reflect the prognosis risk level of the patient and give corresponding prognosis recommendations.

[0041] Optionally, the comprehensive feature fusion in S41 includes:

[0042] S411, Feature Fusion: The overall prognosis impact value obtained from the prediction of the tumor metastasis path Perform splicing operations on the medical data of the patient to generate a comprehensive feature representation of the patient ;

[0043] S412, Feature normalization and dimensionality reduction: Normalize and reduce the dimensionality of the spliced comprehensive feature representation for processing.

[0044] Optionally, the prognostic risk index prediction in S42 includes:

[0045] S421, Neural network structure: Perform layer-by-layer feature extraction and non-linear transformation on the dimensionality-reduced comprehensive feature representation through a multi-layer fully-connected neural network, and output the five-year survival rate and recurrence risk of the patient, specifically including:

[0046] S4211, Calculation of the hidden layer: Calculate the layer hidden feature representation ;

[0047] S4212, Calculation of the output layer: The last layer outputs the predicted value of the five-year survival rate of the patient and the predicted value of the recurrence risk ;

[0048] S422, Loss function optimization: Use the binary cross-entropy loss function to optimize the prediction results of the multi-layer fully-connected neural network, and optimize them separately for the predicted value of the five-year survival rate and the predicted value of the recurrence risk.

[0049] Optionally, the comprehensive evaluation and grading suggestions in S43 include:

[0050] S431, Calculation of the personalized prognosis evaluation value: According to the predicted value of the five-year survival rate of the patient , the predicted value of the recurrence risk and the overall prognosis impact value , comprehensively calculate the personalized prognosis evaluation value of the patient ;

[0051] S432, Risk level division: According to the personalized prognosis evaluation value , grade the risk level of the patient and determine the corresponding prognosis suggestions, specifically including:

[0052] High risk: When , it indicates high risk, the patient has a poor prognosis, and it is recommended to carry out key interventions, including strengthening the treatment plan, intensive monitoring, and auxiliary examinations;

[0053] Medium risk: When , it indicates medium risk, the patient has a medium prognosis, and it is recommended to carry out routine treatment while strengthening regular reexaminations and monitoring;

[0054] Low risk: When When it is, it indicates low risk, the patient's prognosis is normal, and it is recommended to follow the treatment plan and pay attention to follow-up.

[0055] Advantages of the present invention:

[0056] In the present invention, by integrating clinical text data, imaging data and genomic data, combining a variety of preprocessing algorithms and data processing technologies, the consistency, comparability and robustness of the data are significantly improved. In the preprocessing stage, not only the noise information in the data is removed, but also the semantic and feature information of the medical data is retained, solving the problem of difficult to process multi-modal medical data simultaneously, and greatly improving the prediction accuracy and generalization ability of the model.

[0057] In the present invention, through the graph neural network model, the relationship between tumor type, treatment plan and metastasis location is deeply modeled and the metastasis path is predicted. It effectively captures the complex heterogeneous medical relationship structure. Through a multi-step process of constructing a graph, extracting graph features, path prediction and ranking, it can accurately predict the tumor metastasis path and analyze its impact on the patient's prognosis. At the same time, using the normalized adjacency matrix of the graph structure and multi-layer graph convolution operations, not only the numerical stability of the model is improved, but also the interpretability of the path prediction is enhanced, making up for the deficiency of traditional models in modeling complex medical data relationships, and providing an important reference for the formulation of personalized tumor treatment plans.

[0058] In the present invention, through the prognosis evaluation and risk prediction framework, the overall prognosis impact value of the tumor metastasis path and the patient's individualized multi-modal data are comprehensively integrated. Combining a multi-layer neural network to predict the five-year survival rate and recurrence risk, and through the quantified personalized prognosis evaluation value and clear risk grading criteria, accurate risk assessment and grading suggestions are provided for patients, solving the problem of insufficient consideration of individualized characteristics in prognosis evaluation, greatly improving the accuracy and clinical applicability of prognosis evaluation. Finally, it can assist medical institutions to optimize resource allocation, formulate scientific personalized diagnosis and treatment plans, and significantly improve the survival quality and treatment effect of patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the embodiments or the description of the prior art. Obviously, the drawings in the following description are only of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0060] Figure 1 It is a schematic flow chart of the prediction method according to an embodiment of the present invention;

[0061] Figure 2 It is a schematic diagram of tumor metastasis path prediction according to an embodiment of the present invention. Detailed implementation mode

[0062] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the accompanying drawings are only for more specifically describing the embodiments and are not intended to specifically limit the present invention.

[0063] As Figure 1 - Figure 2 shown, a tumor prognosis prediction method based on tumor database text recognition includes the following steps:

[0064] S1, data collection and preprocessing: Collect tumor-related medical data from multiple tumor databases, including clinical text data, image data, and genomic data, denoise and segment the clinical text data, standardize the image data, and complete missing values and normalize the genomic data;

[0065] S2, text information extraction: Use natural language processing technology to extract information from clinical text data, including tumor type, treatment plan, and metastasis location;

[0066] S3, tumor metastasis path prediction: Based on the extracted tumor type, treatment plan, and metastasis location, use a graph neural network (GNN) model to predict the metastasis path, predict the potential metastasis path of the tumor, and analyze the impact of different metastasis paths on the patient's prognosis, specifically including:

[0067] S31, graph construction: Based on the extracted tumor type, treatment plan, and metastasis location, construct a graph structure composed of nodes and edges;

[0068] S32, metastasis path prediction: Based on the constructed graph structure, use a graph neural network (GNN) model to extract the potential rules of the tumor metastasis path and perform metastasis path prediction;

[0069] S33, metastasis path evaluation: Based on the prediction results of the metastasis path, analyze the impact of different metastasis paths on the patient's prognosis;

[0070] S4, prognosis evaluation and risk prediction: Based on the results of tumor metastasis path prediction, combine the patient's medical data to predict the tumor prognosis risk, and comprehensively evaluate the patient's five-year survival rate and recurrence risk;

[0071] Through the above, it is possible to effectively integrate data from different sources, provide accurate analysis of tumor metastasis pathways and personalized prognosis prediction. This not only improves the accuracy of tumor metastasis prediction, but also can predict the survival rate and recurrence risk based on the specific situation of the patient, providing more scientific and data-driven support for the formulation of tumor treatment plans.

[0072] The data collection and preprocessing in S1 include:

[0073] S11, data collection: Collect tumor-related medical data from multiple tumor databases, including clinical text data, imaging data, and genomic data;

[0074] S12, clinical text data preprocessing: Denoise and segment the clinical text data, specifically including:

[0075] Denoising process: Use the TF-IDF (term frequency-inverse document frequency) method to denoise the clinical text data and remove low-frequency irrelevant words, expressed as:

[0076] ;

[0077] Among them, is the importance (weight) of the word in the document , is the frequency of the word in the document , is the number of documents containing the word , is the total number of documents;

[0078] Word segmentation operation: Use the Word2Vec embedding model to convert the clinical text data into word vector representations, retain meaningful words, and remove stop words and irrelevant high-frequency words;

[0079] S13, imaging data preprocessing: Use the Z-score normalization method to normalize the imaging data, expressed as:

[0080] ;

[0081] Among them, is the value after normalization, is the original value of the imaging data, is the mean of the imaging data, is the standard deviation of the imaging data;

[0082] S14, genomic data preprocessing: Complete missing values and normalize the genomic data, specifically including:

[0083] Missing value imputation: Use the KNN interpolation method to fill in the missing values, and fill them by selecting the known values of neighboring samples, which is expressed as:

[0084] ;

[0085] Among them, is the imputation of the missing value, is the number of neighboring samples selected, is the known value of the neighboring samples;

[0086] Normalization: Use the Min-Max normalization method to normalize the genomic data to the interval, which is expressed as:

[0087] ;

[0088] Among them, is the original genomic data, and are the minimum and maximum values of this feature respectively, is the result after normalization;

[0089] Through the above content, the consistency and comparability of different data sources are ensured, noise information is removed, and the accuracy and effectiveness of the data are improved. In addition, by combining algorithms such as TF-IDF noise reduction, Word2Vec word embedding, Z-score standardization, and KNN interpolation, the semantic and feature information of the data is fully retained, and the robustness and prediction ability of the model are greatly improved.

[0090] The text information extraction in S2 includes:

[0091] S21, Word segmentation and text cleaning: Segment the clinical text data through the maximum entropy model based on probability, and remove stop words and irrelevant characters, which is expressed as:

[0092] ;

[0093] Among them, represents the current word, represents the context, is the feature function, is the feature weight, is the normalization factor, is the total number of feature functions;

[0094] ;

[0095] S22, Named Entity Recognition: Using a bidirectional long short-term memory network (Bi-LSTM) combined with a conditional random field (CRF) model to identify entity information of tumor types, treatment plans, and metastasis locations in clinical text data. The objective function of entity recognition is expressed as:

[0096] ;

[0097] Among them, is the conditional probability of the output sequence given the input sequence . is the current annotation label, is the input sequence, clinical text data, is at time step , the joint feature score of the label and the previous label , depending on the input sequence . is all possible label sequences, is the input sequence , that is, the total number of time steps, is the sum of the scores of all possible output sequences, used to normalize the probability;

[0098] Relation Extraction: Based on the extracted named entities, using a relation extraction model with an attention mechanism to identify the associations between tumor types, treatment plans, and metastasis locations, expressed as:

[0099] ;

[0100] ;

[0101] Among them, is the attention weight, , , are the query, key, and value vectors respectively, is the context vector, representing the semantic relationship between tumor types, treatment plans, and metastasis locations, is the number of all feature vectors or vocabulary existing in the context, is the vector transpose symbol;

[0102] Through the above content, the accurate extraction and structured representation of key information such as tumor types, treatment plans, and metastasis locations are achieved. It not only solves the problem that traditional methods are difficult to efficiently process unstructured medical texts but also improves the accuracy and semantic integrity of extraction.

[0103] The atlas construction in S31 includes:

[0104] S311, Node Definition and Representation: Based on the clinical text data, the extracted tumor types, treatment plans, and metastasis locations are used as nodes in the graph, denoted as the node set , where represents the corresponding clinical text data, and the feature vector of each node is represented as , generated by the Word2Vec model;

[0105] S312, Edge Definition and Weight Calculation: Construct the edge set according to the relationships between nodes , represents the node and are associated, and the weight of the edge is calculated through a relationship extraction model, expressed as:

[0106] ;

[0107] where, represents the semantic similarity between the nodes and , is the cosine similarity, measuring the strength of the relationship between two nodes;

[0108] S313, Construction of the Graph Structure: Use the node set and the edge set to construct the graph structure . In the graph structure, the nodes represent tumor types, treatment plans, or metastasis locations, and the edges represent the associated relationships between the nodes (such as the relationship between the treatment plan and the tumor type, the association between the tumor type and the metastasis location, etc.);

[0109] S314, Calculation of the Normalized Adjacency Matrix: Organize the weights of the edges into an adjacency matrix , and normalize the adjacency matrix for subsequent graph neural network calculations, expressed as:

[0110] ;

[0111] where, is the node degree matrix, , is the normalized adjacency matrix;

[0112] Through the above content, the complex relationships between clinical data are accurately expressed, the semantic relevance between features is fully retained, and through the processing of the normalized adjacency matrix, the numerical stability and model adaptability of the graph structure are enhanced. It not only solves the problem of unstructured clinical data modeling but also improves the accuracy and interpretability of tumor metastasis path prediction.

[0113] The metastasis path prediction in S32 includes:

[0114] S321, Input representation: Use the constructed node feature matrix of the graph structure and the normalized adjacency matrix as the input of the graph neural network (GNN) model;

[0115] S322, Graph convolution layer calculation: Extract node features and update node representations through the graph convolutional network (GCN) model, expressed as:

[0116] ;

[0117] where, is the node feature matrix of the -th layer, is the trainable weight matrix of the -th layer, is the ReLU non-linear activation function, is the node feature matrix of the -th layer;

[0118] S323, Aggregation and prediction: After multiple layers of graph convolution, extract the final node feature matrix expressed as , where is the number of layers of the network. Use a fully connected layer to process the node features and predict the potential metastasis paths of the tumor. The prediction results are output through a classifier as the probability of each metastasis path, expressed as:

[0119] ;

[0120] where, is the conditional probability distribution, indicating the probability that the predicted class (i.e., the possible metastasis path) is after being processed by the -th layer of the graph neural network model, is the weight matrix of the output layer, is the number of classes (i.e., the number of possible metastasis paths);

[0121] S324, Metastasis path extraction and sorting: Extract the metastasis paths according to the predicted probability distribution, sort them by probability size, and output the potential laws and paths of tumor metastasis;

[0122] Through the above content, the complex relationship between tumor type, treatment plan and metastasis location can be efficiently captured. Combined with the feature extraction of multi-layer graph convolution and the probability output of softmax classifier, it can not only accurately explore the potential rules of tumor metastasis pathways, but also quantify the predicted probability of each pathway, providing scientific support for personalized tumor treatment plans, and effectively solving the problem that traditional methods are difficult to deal with multi-dimensional heterogeneous relationships, greatly improving the accuracy and interpretability of metastasis path prediction.

[0123] The transfer path assessment in S33 includes:

[0124] S331, feature extraction of transfer path: Based on the transfer path prediction results, the final node feature matrix output from the graph neural network (GNN) model Extract the node features related to the transfer path, aggregate all the node features on the path, and generate the representation vector of the transfer path , expressed as:

[0125] ;

[0126] in, For the predicted The set of nodes on the transfer path, For Node The characteristic vector of is a feature aggregation function, such as average or maximum value;

[0127] S332, prognostic score calculation: representation vector for each transfer pathway The prognostic score is calculated using a fully connected layer, expressed as:

[0128] ;

[0129] in, For path The prognostic score ranges from , is the weight vector of the prognostic evaluation model, is the activation function;

[0130] S333, analysis of the overall impact of the pathway on prognosis: All predicted metastatic pathways and their prognostic scores were combined to calculate the overall prognostic impact value , represents the overall prognosis of the patient, expressed as:

[0131] ;

[0132] in, is the total number of predicted transfer paths, For path The weight of

[0133] Through the above content, the impact of each path on the patient's prognosis is accurately quantified, and the overall prognosis of the patient is comprehensively evaluated. By using the node features extracted by the graph neural network (GNN) and combining feature aggregation and fully connected layer scoring, the risk levels of different metastatic paths can be effectively distinguished, providing a scientific basis for personalized diagnosis and treatment. This not only improves the accuracy of prognosis evaluation but also enhances clinical interpretability, making medical decisions more intelligent and data-driven.

[0134] The prognosis evaluation and risk prediction in S4 include:

[0135] S41, comprehensive feature fusion: Based on the calculated overall prognosis impact value , fuse it with the patient's medical data to generate a comprehensive patient feature representation;

[0136] S42, prognosis risk index prediction: Based on the fused comprehensive patient feature representation, predict the patient's prognosis risk indices through a multi-layer neural network, including five-year survival rate and recurrence risk;

[0137] S43, comprehensive evaluation and grading recommendation: Combine the patient's five-year survival rate, recurrence risk, and overall prognosis impact value , comprehensively calculate the personalized prognosis evaluation value of the patient, which is used to comprehensively reflect the patient's prognosis risk level and give corresponding prognosis suggestions;

[0138] Through the above content, comprehensively integrate the metastatic path information and patient individual characteristics, generate a high-quality comprehensive feature representation, predict the patient's five-year survival rate and recurrence risk, provide accurate individualized prognosis risk indices, and calculate the personalized prognosis evaluation value based on these indices, providing a scientific basis for the comprehensive risk assessment of the patient. This can not only accurately reflect the overall prognosis risk level of the patient but also put forward targeted treatment and management suggestions according to the specific grading results, significantly improving the accuracy, comprehensiveness, and clinical practicality of the assessment.

[0139] The comprehensive feature fusion in S41 includes:

[0140] S411, feature fusion: Concatenate the overall prognosis impact value predicted by the tumor metastasis path and the patient's medical data to generate a comprehensive patient feature representation , expressed as:

[0141] ;

[0142] where , , are the clinical, imaging, and genomic features of the patient respectively, It is a feature splicing operation that integrates all features into a unified representation;

[0143] S412, Feature normalization and dimensionality reduction: Perform normalization and dimensionality reduction on the concatenated comprehensive feature representation to improve computational efficiency and model stability, specifically including:

[0144] Normalization processing: Use Min - Max normalization to map the feature values to , expressed as:

[0145] ;

[0146] where, and are the minimum and maximum values of the feature matrix respectively, and is the normalized comprehensive feature representation;

[0147] Dimensionality reduction processing: Use principal component analysis (PCA) for dimensionality reduction to map the dimension of the normalized comprehensive feature representation from a high - dimensional space to a low - dimensional space , expressed as:

[0148] ;

[0149] where, is the comprehensive feature representation after dimensionality reduction, and is the projection matrix of PCA;

[0150] Through the above content, a comprehensive patient feature representation is generated. After fusion, the data dimension is reduced through normalization and dimensionality reduction operations, improving computational efficiency while retaining high - quality feature information, providing reliable input for subsequent prognosis risk prediction, effectively integrating the global impact of the metastasis path on the patient's prognosis and the local information of individual characteristics, and significantly enhancing the comprehensiveness, accuracy, and practicality of prognosis assessment.

[0151] The prognosis risk index prediction in S42 includes:

[0152] S421, Neural network structure: Perform layer - by - layer feature extraction and non - linear transformation on the comprehensive feature representation after dimensionality reduction through a multi - layer fully - connected neural network, and output the five - year survival rate and recurrence risk of the patient, specifically including:

[0153] S4211, Calculation of the hidden layer: Calculate the hidden feature representation of the th layer , expressed as:

[0154] ;

[0155] where, is the layer hidden feature representation, is the comprehensive feature representation after dimensionality reduction of the initial input, is the layer weight matrix, is the layer bias term, is the ReLU activation function;

[0156] S4212, calculation of the output layer: The output of the last layer is the predicted value of the patient's five-year survival rate and the predicted value of the recurrence risk , expressed as:

[0157] ;

[0158] ;

[0159] where, is the feature representation of the last hidden layer, , are the weight matrices for five-year survival rate and recurrence risk prediction respectively, , are the bias terms corresponding to the output layer respectively, is the Sigmoid activation function;

[0160] S422, loss function optimization: Use the binary cross-entropy loss function to optimize the prediction results of the multi-layer fully connected neural network, and optimize for the predicted value of the five-year survival rate and the predicted value of the recurrence risk respectively, expressed as:

[0161] ;

[0162] ;

[0163] ;

[0164] where, is the loss function of the five-year survival rate, indicating the predicted five-year survival rate and the actual label the error between, is the loss function of the recurrence risk, indicating the predicted recurrence risk and the actual label the error between, is the total loss function, is the total number of samples;

[0165] Through the above, by extracting deep features layer by layer based on a multi-layer neural network and combining the modeling ability of the activation function for non-linear relationships, it is possible to more accurately predict the five-year survival rate and recurrence risk of patients. Optimized using the binary cross-entropy loss function, it ensures the reliability and accuracy of the prediction results. Integrating the comprehensive impact of multi-modal data (such as clinical, imaging, genomic features) and the tumor metastasis pathway provides a scientific basis for personalized risk assessment, significantly improving the accuracy, comprehensiveness, and clinical utility of prognosis prediction.

[0166] The comprehensive evaluation and grading recommendations in S43 include:

[0167] S431, Calculation of the personalized prognosis evaluation value: According to the predicted value of the five-year survival rate of the patient , the predicted value of the recurrence risk and the overall prognosis impact value , comprehensively calculate the personalized prognosis evaluation value of the patient , expressed as:

[0168] ;

[0169] where, is the personalized prognosis evaluation value of the patient, and the higher the value, the better the prognosis, is the overall prognosis impact value, indicating the overall impact of the tumor metastasis pathway on the patient's prognosis, is the predicted value of the five-year survival rate, indicating the probability of the patient's survival, is the predicted value of the recurrence risk, indicating the probability of the patient's tumor recurrence, , , are the corresponding weight parameters respectively, satisfying ;

[0170] S432, Risk level classification: According to the personalized prognosis evaluation value , classify the patient's risk level and determine the corresponding prognosis recommendations, specifically including:

[0171] High risk: When , it indicates high risk, the patient has a poor prognosis, and it is recommended to carry out key interventions, including strengthening the treatment plan, intensive monitoring, and auxiliary examinations;

[0172] Medium risk: When , it indicates medium risk, the patient has a medium prognosis, and it is recommended to carry out routine treatment while strengthening regular reexaminations and monitoring;

[0173] Low risk: When , it indicates low risk, the patient has a normal prognosis, and it is recommended to follow the treatment plan and pay attention to follow-up;

[0174] Through the above, it accurately reflects the overall prognostic risk level of the patient, provides a scientific quantitative basis, ensures the personalization and clinical applicability of the evaluation results. At the same time, through clear risk classification criteria and specific grading suggestions, it can formulate targeted diagnosis, treatment and management plans for patients with different risk levels, which helps to optimize the allocation of medical resources and improve the treatment effect and quality of life of patients.

[0175] This invention covers any substitutions, modifications, equivalent methods and solutions made on the essence and scope of this invention. To enable the public to have a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments of this invention. However, those skilled in the art can fully understand this invention without the description of these details. In addition, well-known methods, processes, procedures, components and circuits, etc. are not described in detail to avoid unnecessary confusion to the essence of this invention.

[0176] The above are only the preferred embodiments of this invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of this invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this invention.

Claims

1. A tumor prognosis prediction method based on tumor database text recognition, characterized in that: The following steps are involved: S1, data collection and preprocessing: collect tumor-related medical data from multiple tumor databases, including clinical text data, imaging data and genomic data, and perform denoising and word segmentation on clinical text data, standardize imaging data, and fill in missing values ​​and normalize genomic data; S2, text information extraction: natural language processing technology is used to extract information from clinical text data, including tumor type, treatment plan, and metastasis location; S3, Tumor metastasis path prediction: Based on the extracted tumor type, treatment plan and metastasis location, the graph neural network model is used to predict the potential metastasis path of the tumor and analyze the impact of different metastasis paths on patient prognosis, including: S31, graph construction: construct a graph structure consisting of nodes and edges based on the extracted tumor types, treatment options, and metastasis locations; S32, metastasis pathway prediction: Based on the constructed graph structure, the graph neural network model is used to extract the potential rules of tumor metastasis pathways and predict the metastasis pathways; S33, metastatic pathway assessment: based on the prediction results of metastatic pathways, analyze the impact of different metastatic pathways on patient prognosis; S4, Prognosis assessment and risk prediction: Based on the results of tumor metastasis path prediction, combined with the patient's medical data, the tumor prognosis risk is predicted, and the patient's five-year survival rate and recurrence risk are comprehensively evaluated.

2. The method for predicting tumor prognosis based on tumor database text recognition according to claim 1, characterized in that: The data collection and preprocessing in S1 include: S11, data collection: collect tumor-related medical data from multiple tumor databases, including clinical text data, imaging data, and genomic data; S12, clinical text data preprocessing: denoising and word segmentation of clinical text data; S13, image data preprocessing: image data were standardized using the Z-score standardization method; S14, genomic data preprocessing: missing value filling and normalization of genomic data.

3. The method for predicting tumor prognosis based on tumor database text recognition according to claim 1, characterized in that: The text information extraction in S2 includes: S21, word segmentation and text cleaning: clinical text data is segmented using a probability-based maximum entropy model to remove stop words and irrelevant characters; S22, named entity recognition: using a bidirectional long short-term memory network combined with a conditional random field model to identify entity information of tumor type, treatment regimen, and metastasis location in clinical text data; Relation extraction: Based on the extracted named entities, a relation extraction model with attention mechanism is used to identify the associations between tumor type, treatment regimen, and metastasis location.

4. The method for predicting tumor prognosis based on tumor database text recognition according to claim 1, characterized in that: The map construction in S31 includes: S311, Node definition and representation: Based on clinical text data, the extracted tumor types, treatment plans and metastasis locations are used as nodes in the graph, recorded as a node set ,in Represents the corresponding clinical text data, and the feature vector of each node is expressed as ; S312, Edge definition and weight calculation: Construct edge sets based on the relationships between nodes , Representation Node and The relationship between them is extracted and the weight of the edge is calculated through the relationship extraction model; S313, Construction of graph structure: Using node sets and edge set Building the graph structure ,In the graph structure, nodes represent tumor types, treatment ,schemes, or metastasis locations, and edges represent the ,association between nodes; S314, Calculation of normalized adjacency matrix: Organize edge weights into an adjacency matrix , and normalize the adjacency matrix.

5. The method for predicting tumor prognosis based on tumor database text recognition according to claim 4, characterized in that: The transfer path prediction in S32 includes: S321, input representation: the graph structure to be constructed The node feature matrix and normalized adjacency matrix as input to the graph neural network model; S322, graph convolution layer calculation: extract node features and update node representation through the graph convolution network model; S323, aggregation and prediction: After multiple layers of graph convolution, the final node feature matrix is ​​extracted and expressed as ,in is the number of layers in the network. The fully connected layer is used to process the node features and predict the potential metastasis path of the tumor. The prediction result is output as the probability of each metastasis path through the classifier. S324, metastasis path extraction and sorting: extract the metastasis path based on the predicted probability distribution, sort it by probability, and output the potential rules and paths of tumor metastasis.

6. The method for predicting tumor prognosis based on tumor database text recognition according to claim 5, characterized in that: The transfer path evaluation in S33 includes: S331, feature extraction of transfer path: Based on the transfer path prediction results, the final node feature matrix output from the graph neural network model Extract the node features related to the transfer path, aggregate all the node features on the path, and generate the representation vector of the transfer path ; S332, prognostic score calculation: representation vector for each transfer pathway The prognostic score was calculated using a fully connected layer; S333, analysis of the overall impact of the pathway on prognosis: All predicted metastatic pathways and their prognostic scores were combined to calculate the overall prognostic impact value , indicating the overall prognosis of the patient.

7. The method for predicting tumor prognosis based on tumor database text recognition according to claim 6, characterized in that: The prognostic assessment and risk prediction in S4 include: S41, Comprehensive feature fusion: Calculated overall prognostic impact value , which is fused with the patient’s medical data to generate a comprehensive feature representation of the patient; S42, prediction of prognostic risk indicators: Based on the fused comprehensive characteristics of the patient, the patient's prognostic risk indicators, including five-year survival rate and recurrence risk, are predicted through a multi-layer neural network; S43, Comprehensive assessment and grading recommendations: Combined with the patient's five-year survival rate, recurrence risk and overall prognostic impact value , comprehensively calculate the patient's personalized prognostic assessment value, which is used to comprehensively reflect the patient's prognostic risk level and give corresponding prognostic recommendations.

8. The method for predicting tumor prognosis based on tumor database text recognition according to claim 7, characterized in that: The comprehensive feature fusion in S41 includes: S411, Feature Fusion: The overall prognostic impact value obtained by predicting the tumor metastasis pathway And the patient's medical data are spliced ​​to generate a comprehensive feature representation of the patient ; S412, Feature Normalization and Dimensionality Reduction: Comprehensive Feature Representation after Concatenation Perform normalization and dimensionality reduction.

9. The method for predicting tumor prognosis based on tumor database text recognition according to claim 8, characterized in that: The prediction of the prognostic risk indicator in S42 includes: S421, Neural network structure: Through a multi-layer fully connected neural network, the comprehensive feature representation after dimensionality reduction is extracted layer by layer and nonlinearly transformed to output the patient's five-year survival rate and recurrence risk, including: S4211, calculation of hidden layer: Calculate the Layer hidden feature representation ; S4212, calculation of the output layer: the last layer outputs the patient's five-year survival rate prediction value and recurrence risk prediction value ; S422, loss function optimization: Use the binary cross entropy loss function to optimize the prediction results of the multi-layer fully connected neural network, and optimize the five-year survival rate prediction value and recurrence risk prediction value respectively.

10. The method for predicting tumor prognosis based on tumor database text recognition according to claim 9, characterized in that: The comprehensive assessment and grading recommendations in S43 include: S431, Calculation of personalized prognostic assessment value: based on the patient's five-year survival rate prediction value , recurrence risk prediction value and overall prognostic impact value , comprehensively calculate the patient's personalized prognostic assessment value ; S432, Risk level classification: based on personalized prognostic assessment values , classify the patient's risk level and determine the corresponding prognostic recommendations, including: High risk: When When , it indicates high risk and poor prognosis for the patient, and focused intervention is recommended, including intensive treatment plans, intensive monitoring, and auxiliary examinations; Medium risk: When , it indicates medium risk and the patient has a moderate prognosis. Conventional treatment is recommended, while regular review and monitoring are strengthened; Low risk: When When , it indicates low risk and the patient’s prognosis is normal. It is recommended to follow the treatment plan and pay attention to follow-up.

Citation Information

Patent Citations

  • Distance metastasis identification method based on gene interaction mode optimization graph representation

    CN114141306A

  • Liver transplantation postoperative tumor recurrence prediction method based on graph convolutional network

    CN115620913A