A digital tumor screening platform and method
By integrating imaging, clinical, genetic and other data through a multimodal fusion model, a tumor risk prediction model is constructed, which solves the problems of missed diagnosis, misdiagnosis and data silos in traditional tumor screening, and realizes efficient and personalized tumor screening and graded diagnosis and treatment.
Patent Information
- Application Number
- CN202511108521.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-08
AI Technical Summary
Traditional tumor screening relies on doctors' experience and single-modality data analysis, resulting in high rates of missed diagnosis and misdiagnosis, and serious medical data silos, making it difficult to achieve comprehensive consideration of multi-dimensional information and data interconnection.
A multimodal fusion model is constructed using a deep learning framework, integrating multi-source data such as imaging, clinical, genetic, and lifestyle habits. Features are extracted through convolutional neural networks and recurrent neural networks, and machine learning algorithms are used to train tumor risk prediction models, set stratified risk thresholds, and generate personalized treatment plans.
It significantly reduced the early missed diagnosis rate of common tumors, shortened the screening process time, improved the diagnostic compliance rate and doctor-patient communication efficiency, and realized personalized tumor screening and graded diagnosis and treatment.
Smart Images

Figure CN120600322B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital tumor screening technology, and in particular to a digital tumor screening platform and method. Background Art
[0002] Traditional tumor screening methods rely primarily on physician experience combined with various individual testing technologies. In imaging examinations, such as X-rays, CT scans, and MRIs, doctors need to manually read the films. Faced with massive amounts of imaging data, this not only consumes a significant amount of time and energy, but is also prone to missed diagnoses and misdiagnoses due to factors such as visual fatigue and differences in experience. In terms of laboratory testing, while common tumor marker tests can provide some reference, their specificity and sensitivity are poor. For example, when carcinoembryonic antigen (CEA) is used for colorectal cancer screening, some patients with early-stage colorectal cancer do not show a significant increase in this indicator, while some benign diseases such as colitis and pancreatitis may produce false positive results, increasing uncertainty in clinical diagnosis.
[0003] Currently, some medical institutions have introduced AI-based screening and diagnostic systems, which use deep learning algorithms to perform preliminary image analysis and label suspicious lesions. However, these systems mostly utilize single-modality data, analyzing only imaging or laboratory data independently and lacking comprehensive consideration of the patient's multidimensional information. For example, they fail to consider other factors that can influence screening results, such as the patient's family medical history, lifestyle habits (smoking history, occupational exposures, etc.), and genetic test results. Existing tumor screening platforms can automatically categorize and group patients using a pre-defined screening rule engine, combining basic patient information and questionnaire responses. However, overall data integration remains a significant problem, with medical data silos. Independent information systems across different medical institutions and departments operate with varying data formats and standards, making interoperability difficult. This hinders data acquisition for digital screening systems, preventing them from integrating massive, comprehensive case data for model training and optimization, limiting the advancement of the system's intelligent diagnostic capabilities. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a digital tumor screening platform and method. The following technical solutions are adopted:
[0005] A digital tumor screening method comprises the following steps:
[0006] Step 1: Based on a deep learning framework, a multimodal fusion model integrating convolutional neural networks and recurrent neural networks is constructed. The collected comprehensive patient data is input into the multimodal fusion model to extract imaging features and process time series clinical data. Through feature splicing and attention mechanism strategies, the different modal data are deeply fused to output comprehensive tumor correlation features.
[0007] Step 2: Utilize machine learning algorithms and large-scale tumor data for training to build a tumor risk prediction model, input comprehensive tumor correlation features, and output the probability of different tumors;
[0008] Step 3: Set stratified risk thresholds, stratify the probability of disease occurrence, and trigger corresponding warning actions for each stratum;
[0009] Step 4: Based on patient risk stratification, basic information, and tumor screening guidelines, an intelligent recommendation algorithm is used to generate treatment recommendations and expected outcomes.
[0010] Step 5: Use data visualization technology to generate a screening report, which includes comprehensive patient data, tumor type and incidence probability prediction, treatment recommendations, and expected results.
[0011] By adopting the above technical solutions, standardized interfaces are connected to hospital HIS, PACS, LIMS, and other systems to collect patient imaging data (such as CT and MRI), time-series clinical data (such as tumor marker changes and blood pressure fluctuations), text data (such as medical records and lifestyle questionnaires), and genetic data in real time. Image data is subjected to noise reduction and normalization, time-series data is imputed with missing values and normalized, and text data is segmented and converted into vectors.
[0012] Convolutional neural network (CNN) was used to extract lesion features from imaging data, and recurrent neural network (RNN) was used to process time series clinical data.
[0013] Through feature splicing, imaging features, time series features, text vectors and genetic data are merged into initial fusion features; through the attention mechanism, weights are dynamically assigned to achieve deep correlation of cross-modal features, and finally output comprehensive tumor correlation features (comprehensive feature vectors covering imaging, clinical, genetic and lifestyle habits).
[0014] Based on large-scale annotated data (including pathologically confirmed cancer patients, healthy individuals, and patients with benign diseases), we employed machine learning algorithms such as random forests and XGBoost, using comprehensive tumor-associated features as input and the presence or absence of tumors as the label for training. Through iterative optimization, the model achieved clinically applicable accuracy and recall on the validation set.
[0015] Probability output: For new patients, the fused comprehensive tumor association features are input. The model calculates and outputs the risk probability of common tumors such as lung cancer, gastric cancer, and intestinal cancer through trained parameters, enabling simultaneous assessment of multiple tumor types.
[0016] Based on clinical guidelines (such as the National Cancer Center's "Cancer Prevention and Screening Guidelines") and historical data statistics, risk stratification thresholds are set: high risk (probability of incidence ≥70%), intermediate risk (30%-70%), and low risk (<30%).
[0017] The system automatically matches the patient risk probability with the threshold and triggers the corresponding action:
[0018] Push to the medical side in real time for priority processing, synchronously link the patient's holographic files (images, test results, etc.), and prompt emergency examination suggestions;
[0019] Generate follow-up tasks, push them to the doctor's schedule, and send review reminders to the patient;
[0020] Low-risk patients: Personalized health intervention plans (such as smoking cessation guidelines and dietary recommendations) are pushed, and annual screening plans are automatically generated.
[0021] Based on patient risk stratification, basic information (age, gender, family history, etc.) and tumor screening guidelines, an intelligent recommendation algorithm using rule-based reasoning + collaborative filtering is used:
[0022] Match the examination or treatment items in the guidelines with the patient's characteristics;
[0023] Refer to the diagnosis and treatment pathways of 1,000+ similar cases (same age, same risk, same lifestyle) to optimize recommendation priorities.
[0024] As the patient's health data is updated (such as physical examination results and changes in lifestyle habits), the system recalculates the risk probability in real time and adjusts the recommended plan synchronously (such as extending the interval for lung cancer screening from six months to one year after quitting smoking).
[0025] Expected effect evaluation: Based on the statistics of intervention effects on similar patients in historical data, personalized expected effects are output.
[0026] Data standardization and mapping: Standardize data such as tumor risk probability and abnormality indicators (such as Z-score conversion) and map the values into visual dimensions (color, size, and position).
[0027] Multi-dimensional chart presentation:
[0028] Use a radar chart to display the risk probability of each tumor (high-risk items are marked in red, medium-risk items are marked in yellow);
[0029] Use heatmaps to present abnormal indicators (the larger the absolute value of the Zscore, the darker the color);
[0030] Use a flowchart to show the chronological order of recommended solutions;
[0031] Use a bar graph to compare expected effects before and after the intervention.
[0032] Through multimodal fusion models, we can integrate multi-source data such as imaging, clinical, genetic, and lifestyle habits to comprehensively capture tumor characteristics, which can significantly reduce the early missed diagnosis rate of common tumors.
[0033] Through real-time access to multi-source data through standardized interfaces, multimodal models can complete feature extraction and risk prediction in a short time, significantly shortening the entire screening process.
[0034] The tiered early warning mechanism enables doctors to prioritize high-risk patients, reduce the waste of ineffective screening resources, improve the initial tumor screening capabilities of primary medical institutions, and facilitate the implementation of tiered diagnosis and treatment.
[0035] Break through the traditional one-size-fits-all screening model and generate personalized plans based on patient risk stratification, family history, lifestyle habits, etc.
[0036] For men with a family history of lung cancer and long-term smoking, low-dose spiral CT plus lung cancer marker testing is recommended every six months;
[0037] For women with BRCA mutations, increase the frequency of breast MRI + genetic review to twice a year.
[0038] The system dynamically adjusts the plan as the patient's health data is updated (such as extending the screening interval after quitting smoking), ensuring that the intervention is always appropriate to the patient's current condition.
[0039] Through standardized interfaces and CDR data models, more than 80% of common tumor-related data (clinical, imaging, genetic, etc.) are integrated to solve the problem of medical data silos and build a patient-centered holographic archive.
[0040] The visual report presents the results in intuitive charts (heat maps, flow charts), accompanied by plain text explanations, which improves patients' understanding of the screening results; doctors can quickly obtain full-dimensional information about patients through holographic files, improve diagnostic compliance, and significantly improve the efficiency of doctor-patient communication.
[0041] Optionally, step 1 includes the following sub-steps:
[0042] Step 11: For the image data, extract the lesion features through the convolution layer and the pooling layer, and output a 256-dimensional image feature vector;
[0043] Step 12: For the time series clinical data, LSTM is used to process the changing trend of tumor markers over time and output a 128-dimensional time series feature vector;
[0044] Step 13: Concatenate the image features output by the convolutional neural network, the temporal features output by the recurrent neural network, the text vector, and the gene data vector into a 512-dimensional initial fusion vector;
[0045] Step 14: Dynamically assign weights through the attention layer;
[0046] Step 15: After processing by the fully connected layer, a 1024-dimensional tumor comprehensive correlation feature vector is generated.
[0047] By adopting the above technical solutions, the model is compatible with multi-source data of various tumor types (lung cancer, gastric cancer, breast cancer, etc.). Through the universal framework design at the feature level, it can be quickly migrated to different tumor screening scenarios (such as adapting the fusion logic of the lung cancer model to breast cancer by only adjusting the attention weight parameters), reducing the cost of repeated development and significantly shortening the system's adaptation cycle for new tumor types.
[0048] Step 1 comprehensively captures the imaging characteristics, progression patterns, and associated factors of tumors through precise extraction, dynamic fusion, and high-dimensional representation of multimodal features, laying the core foundation for subsequent risk prediction and personalized screening, and significantly improving the accuracy and comprehensiveness of early tumor screening.
[0049] Optionally, in step 2, an ensemble learning algorithm is used to input the tumor comprehensive association feature vector generated in step 1, and training is performed with whether or not the tumor occurs as the dependent variable.
[0050] Optionally, in step 2, a random forest is used to construct a set number of decision trees, and the Gini coefficient is used to evaluate the importance of features and rank them according to their importance;
[0051] The depth and number of leaf nodes of the decision tree were adjusted through grid search, and 5-fold cross validation was used to avoid overfitting.
[0052] By adopting the above technical solution, random forest can handle multiple classification tasks simultaneously. After inputting a 1024-dimensional tumor comprehensive correlation feature vector, it can simultaneously output the incidence probability of multiple tumors such as lung cancer, gastric cancer, and intestinal cancer. There is no need to train a separate model for each tumor, thereby improving screening efficiency.
[0053] Combined with the feature importance ranking, the key risk factors for different tumors (such as smoking history for lung cancer and BRCA mutation for breast cancer) can be identified at the same time, providing a multi-dimensional basis for subsequent stratified warnings and personalized plans.
[0054] Optionally, in step 3, based on clinical data statistics, the probability of disease progression is divided into high risk, medium risk, low risk and no risk, where high risk corresponds to a probability of disease progression greater than or equal to 70%, medium risk corresponds to a probability of disease progression greater than or equal to 30% and less than 70%, low risk corresponds to a probability of disease progression less than 30%, and no risk corresponds to a probability of disease progression less than 5%.
[0055] Optionally, step 4 includes the following sub-steps:
[0056] Step 41: extract the core characteristics of the patient, including risk stratification, basic information, lifestyle habits, and genetic test results;
[0057] Step 42: extract the recommended examination or treatment plan, recommended frequency Fq, and applicable population condition C from the tumor screening guidelines for different tumor types;
[0058] Step 43, using a hybrid recommendation algorithm combining tumor screening guideline matching and similar case collaborative filtering to calculate the recommended treatment plan score;
[0059] Step 44: Filter projects with recommendation scores greater than a set score threshold, sort them from high to low by score, and form a core recommendation plan;
[0060] Step 45: Based on the statistics of intervention effects on similar patients in historical data, the expected effect corresponding to the core recommended plan is output.
[0061] Optionally, in step 43, the formula for calculating the recommended treatment plan score is:
[0062] ;
[0063] in It is the recommended treatment option The score, and is the weight coefficient, It is a treatment plan Compatibility with patient P's cancer screening guidelines, It is a treatment plan Similarity of recommendations in detailed cases.
[0064] Optional treatment options Compatibility with patient P's cancer screening guidelines The calculation formula is:
[0065] ;
[0066] in Treatment options in cancer screening guidelines The kth applicable condition of is the kth feature of patient P, is the matching indicator function, is the weight of condition k, which is set by expert consensus;
[0067] Treatment options Similarity of recommendations in detailed cases The calculation formula is:
[0068] ;
[0069] in, is the jth historical case with similar characteristics to patient P, For patients P and The feature similarity of For cases Whether treatment plans have been used , if yes, it is 1, otherwise, it is 0.
[0070] Optionally, the expected effect index is calculated based on the statistical data of the intervention effect of the corresponding treatment plan for similar patients in historical data. The calculation formula is:
[0071] ;
[0072] in is the expected effect index, is the total number of people historically with similar characteristics to the current patients who adopted this regimen; It is the number of people diagnosed early or with reduced risk in the group using this regimen.
[0073] By adopting the above technical solutions, a personalized recommendation logic based on guideline specifications, actual case verification, and quantitative effect evaluation was constructed through quantitative formulas (recommended treatment plan score, guideline matching degree, case recommendation similarity, and expected effect index). This solves the problems of traditional treatment plans relying on experience, being highly subjective, and having difficult to quantify effects, improves the accuracy and standardization of treatment plans, and reduces unreasonable recommendations.
[0074] The guideline-matching formula aims to emphasize core criteria crucial for cancer screening by using weights set by expert consensus (e.g., family history weighted at 0.3, specific tumor marker weighted at 0.25), while downplaying the influence of secondary criteria (e.g., body mass index weighted at 0.05). For example, in lung cancer screening, long-term smoking history and CT nodule characteristics are weighted significantly higher than dietary habits. This ensures that the plan strictly aligns with authoritative standards such as the "Cancer Prevention and Screening Guidelines," avoiding over- or under-recommendations due to interference from irrelevant criteria.
[0075] Compared with the traditional method of mechanically applying guidelines, this formula improves the matching accuracy between the plan and the guidelines, greatly increases the proportion of compliance with medical standards, and reduces the medical risks caused by the plan deviating from the guidelines.
[0076] Based on historical data on the early diagnosis rate or risk reduction rate of similar patients using the plan, this provides quantitative evidence for the plan's effectiveness (e.g., this plan increased the early diagnosis rate for similar patients to 82%), rather than vague statements (e.g., "probably effective"). This improves patients' understanding of the plan's effectiveness, allowing doctors to explain its value to patients based on quantitative data, shortening doctor-patient communication time. Furthermore, quantitative results provide objective reference for doctors to adjust the plan, reducing subjective decision-making bias.
[0077] A digital tumor screening platform is used to implement a digital tumor screening method. The tumor screening platform includes a patient information entry module, multiple information docking modules, a data analysis module, a tumor risk prediction module, a treatment plan recommendation module and a tumor screening result output module. The patient information entry module is used to enter comprehensive patient data. The multiple information docking modules are respectively connected to the hospital information system, image archiving and communication system and laboratory information management system. The data analysis module builds a multimodal fusion model and communicates with the patient information entry module and the multiple information docking modules respectively to output comprehensive tumor correlation features. The tumor risk prediction module builds a tumor risk prediction model and communicates with the data analysis module to input comprehensive tumor correlation features and output the probability of different tumors. The treatment plan recommendation module communicates with the data analysis module and the tumor risk prediction module respectively to output the core recommended plan and the corresponding expected effect. The tumor screening result output module communicates with the data analysis module, the tumor risk prediction module and the treatment plan recommendation module to generate a screening report using data visualization technology.
[0078] In summary, the present invention includes at least one of the following beneficial technical effects:
[0079] The present invention can provide a digital tumor screening platform and method. Through a multimodal fusion model, it integrates multi-source data such as imaging, clinical, genetic, and lifestyle habits to comprehensively capture tumor characteristics, which can significantly reduce the early missed diagnosis rate of common tumors.
[0080] Through real-time access to multi-source data through standardized interfaces, multimodal models can complete feature extraction and risk prediction in a short time, significantly shortening the entire screening process.
[0081] The tiered early warning mechanism enables doctors to prioritize high-risk patients, reduce the waste of ineffective screening resources, improve the initial tumor screening capabilities of primary medical institutions, and facilitate the implementation of tiered diagnosis and treatment.
[0082] Break through the traditional one-size-fits-all screening model and generate personalized plans based on patient risk stratification, family history, lifestyle habits, etc.
[0083] The system dynamically adjusts the plan as the patient's health data is updated to ensure that the intervention is always appropriate to the patient's current condition.
[0084] Through standardized interfaces and CDR data models, more than 80% of common tumor-related data are integrated, solving the problem of medical data silos and building a patient-centered holographic archive.
[0085] The visual report presents the results in intuitive charts and accompanied by plain text explanations, which improves patients' understanding of the screening results; doctors can quickly obtain full-dimensional information about patients through holographic files, which increases the diagnostic compliance rate and significantly improves the efficiency of doctor-patient communication. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 This is a flow chart of a digital tumor screening method of the present invention. DETAILED DESCRIPTION
[0087] The present invention will be further described in detail below with reference to the accompanying drawings.
[0088] The embodiments of the present invention disclose a digital tumor screening platform and method.
[0089] Reference Figure 1 , Example 1, a digital tumor screening method, comprising the following steps:
[0090] Step 1: Based on a deep learning framework, a multimodal fusion model integrating convolutional neural networks and recurrent neural networks is constructed. The collected comprehensive patient data is input into the multimodal fusion model to extract imaging features and process time series clinical data. Through feature splicing and attention mechanism strategies, the different modal data are deeply fused to output comprehensive tumor correlation features.
[0091] Step 2: Utilize machine learning algorithms and large-scale tumor data for training to build a tumor risk prediction model, input comprehensive tumor correlation features, and output the probability of different tumors;
[0092] Step 3: Set stratified risk thresholds, stratify the probability of disease occurrence, and trigger corresponding warning actions for each stratum;
[0093] Step 4: Based on patient risk stratification, basic information, and tumor screening guidelines, an intelligent recommendation algorithm is used to generate treatment recommendations and expected outcomes.
[0094] Step 5: Use data visualization technology to generate a screening report, which includes comprehensive patient data, tumor type and incidence probability prediction, treatment recommendations, and expected results.
[0095] Standardized interfaces connect to hospital HIS, PACS, LIMS, and other systems, enabling real-time collection of patient imaging data (CT, MRI, etc.), time-series clinical data (tumor marker changes, blood pressure fluctuations, etc.), text data (medical records, lifestyle questionnaires), and genetic data. Image data is denoised and standardized, time-series data is filled with missing values and standardized, and text data is segmented and converted into vectors.
[0096] Convolutional neural network (CNN) was used to extract lesion features from imaging data, and recurrent neural network (RNN) was used to process time series clinical data.
[0097] Through feature splicing, imaging features, time series features, text vectors and genetic data are merged into initial fusion features; through the attention mechanism, weights are dynamically assigned to achieve deep correlation of cross-modal features, and finally output comprehensive tumor correlation features (comprehensive feature vectors covering imaging, clinical, genetic and lifestyle habits).
[0098] Based on large-scale annotated data (including pathologically confirmed cancer patients, healthy individuals, and patients with benign diseases), we employed machine learning algorithms such as random forests and XGBoost, using comprehensive tumor-associated features as input and the presence or absence of tumors as the label for training. Through iterative optimization, the model achieved clinically applicable accuracy and recall on the validation set.
[0099] Probability output: For new patients, the fused comprehensive tumor association features are input. The model calculates and outputs the risk probability of common tumors such as lung cancer, gastric cancer, and intestinal cancer through trained parameters, enabling simultaneous assessment of multiple tumor types.
[0100] Based on clinical guidelines (such as the National Cancer Center's "Cancer Prevention and Screening Guidelines") and historical data statistics, risk stratification thresholds are set: high risk (probability of incidence ≥70%), intermediate risk (30%-70%), and low risk (<30%).
[0101] The system automatically matches the patient risk probability with the threshold and triggers the corresponding action:
[0102] Push to the medical side in real time for priority processing, synchronously link the patient's holographic files (images, test results, etc.), and prompt emergency examination suggestions;
[0103] Generate follow-up tasks, push them to the doctor's schedule, and send review reminders to the patient;
[0104] Low-risk patients: Personalized health intervention plans (such as smoking cessation guidelines and dietary recommendations) are pushed, and annual screening plans are automatically generated.
[0105] Based on patient risk stratification, basic information (age, gender, family history, etc.) and tumor screening guidelines, an intelligent recommendation algorithm using rule-based reasoning + collaborative filtering is used:
[0106] Match the examination or treatment items in the guidelines with the patient's characteristics;
[0107] Refer to the diagnosis and treatment pathways of 1,000+ similar cases (same age, same risk, same lifestyle) to optimize recommendation priorities.
[0108] As the patient's health data is updated (such as physical examination results and changes in lifestyle habits), the system recalculates the risk probability in real time and adjusts the recommended plan synchronously (such as extending the interval for lung cancer screening from six months to one year after quitting smoking).
[0109] Expected effect evaluation: Based on the statistics of intervention effects on similar patients in historical data, personalized expected effects are output.
[0110] Data standardization and mapping: Standardize data such as tumor risk probability and abnormality indicators (such as Z-score conversion) and map the values into visual dimensions (color, size, and position).
[0111] Multi-dimensional chart presentation:
[0112] Use a radar chart to display the risk probability of each tumor (high-risk items are marked in red, medium-risk items are marked in yellow);
[0113] Use heatmaps to present abnormal indicators (the larger the absolute value of the Zscore, the darker the color);
[0114] Use a flowchart to show the chronological order of recommended solutions;
[0115] Use a bar graph to compare expected effects before and after the intervention.
[0116] Through multimodal fusion models, we can integrate multi-source data such as imaging, clinical, genetic, and lifestyle habits to comprehensively capture tumor characteristics, which can significantly reduce the early missed diagnosis rate of common tumors.
[0117] Through real-time access to multi-source data through standardized interfaces, multimodal models can complete feature extraction and risk prediction in a short time, significantly shortening the entire screening process.
[0118] The tiered early warning mechanism enables doctors to prioritize high-risk patients, reduce the waste of ineffective screening resources, improve the initial tumor screening capabilities of primary medical institutions, and facilitate the implementation of tiered diagnosis and treatment.
[0119] Break through the traditional one-size-fits-all screening model and generate personalized plans based on patient risk stratification, family history, lifestyle habits, etc.
[0120] For men with a family history of lung cancer and long-term smoking, low-dose spiral CT plus lung cancer marker testing is recommended every six months;
[0121] For women with BRCA mutations, increase the frequency of breast MRI + genetic review to twice a year.
[0122] The system dynamically adjusts the plan as the patient's health data is updated (such as extending the screening interval after quitting smoking), ensuring that the intervention is always appropriate to the patient's current condition.
[0123] Through standardized interfaces and CDR data models, more than 80% of common tumor-related data (clinical, imaging, genetic, etc.) are integrated to solve the problem of medical data silos and build a patient-centered holographic archive.
[0124] The visual report presents the results in intuitive charts (heat maps, flow charts), accompanied by plain text explanations, which improves patients' understanding of the screening results by 70%; doctors can quickly obtain full-dimensional information about patients through holographic files, increase the diagnostic compliance rate by 30%, and significantly improve the efficiency of doctor-patient communication.
[0125] In Example 2, step 1 includes the following sub-steps:
[0126] Step 11: For the image data, extract the lesion features through the convolution layer and the pooling layer, and output a 256-dimensional image feature vector;
[0127] Step 12: For the time series clinical data, LSTM is used to process the changing trend of tumor markers over time and output a 128-dimensional time series feature vector;
[0128] Step 13: Concatenate the image features output by the convolutional neural network, the temporal features output by the recurrent neural network, the text vector, and the gene data vector into a 512-dimensional initial fusion vector;
[0129] Step 14: Dynamically assign weights through the attention layer;
[0130] Step 15: After processing by the fully connected layer, a 1024-dimensional tumor comprehensive correlation feature vector is generated.
[0131] The model is compatible with multi-source data from various tumor types (lung cancer, gastric cancer, breast cancer, etc.). Through a universal framework design at the feature level, it can be quickly migrated to different tumor screening scenarios (for example, adapting the fusion logic of the lung cancer model to breast cancer only requires adjusting the attention weight parameters), reducing repeated development costs and significantly shortening the system's adaptation cycle for new tumor types.
[0132] Step 1 comprehensively captures the imaging characteristics, progression patterns, and associated factors of tumors through precise extraction, dynamic fusion, and high-dimensional representation of multimodal features, laying the core foundation for subsequent risk prediction and personalized screening, and significantly improving the accuracy and comprehensiveness of early tumor screening.
[0133] In Example 3, in step 2, an ensemble learning algorithm is used to input the tumor comprehensive correlation feature vector generated in step 1, and training is performed with whether the tumor has occurred or not as the dependent variable.
[0134] In Example 4, in step 2, a set number of decision trees are constructed using random forests, and the importance of features is evaluated using the Gini coefficient, and the importance is ranked;
[0135] The depth and number of leaf nodes of the decision tree were adjusted through grid search, and 5-fold cross validation was used to avoid overfitting.
[0136] Random forest can handle multiple classification tasks simultaneously. After inputting a 1024-dimensional tumor comprehensive correlation feature vector, it can simultaneously output the incidence probability of multiple tumors such as lung cancer, gastric cancer, and intestinal cancer. There is no need to train a separate model for each tumor, thereby improving screening efficiency.
[0137] Combined with the feature importance ranking, the key risk factors for different tumors (such as smoking history for lung cancer and BRCA mutation for breast cancer) can be identified at the same time, providing a multi-dimensional basis for subsequent stratified warnings and personalized plans.
[0138] In Example 5, in step 3, based on clinical data statistics, the probability of disease is divided into high risk, medium risk, low risk and no risk, where high risk corresponds to a probability of disease greater than or equal to 70%, medium risk corresponds to a probability of disease greater than or equal to 30% and less than 70%, low risk corresponds to a probability of disease less than 30%, and no risk corresponds to a probability of disease less than 5%.
[0139] In Example 6, step 4 includes the following sub-steps:
[0140] Step 41: extract the core characteristics of the patient, including risk stratification, basic information, lifestyle habits, and genetic test results;
[0141] Step 42: extract the recommended examination or treatment plan, recommended frequency Fq, and applicable population condition C from the tumor screening guidelines for different tumor types;
[0142] Step 43, using a hybrid recommendation algorithm combining tumor screening guideline matching and similar case collaborative filtering to calculate the recommended treatment plan score;
[0143] Step 44: Filter projects with recommendation scores greater than a set score threshold, sort them from high to low by score, and form a core recommendation plan;
[0144] Step 45: Based on the statistics of intervention effects on similar patients in historical data, the expected effect corresponding to the core recommended plan is output.
[0145] In Example 7, in step 43, the formula for calculating the recommended treatment plan score is:
[0146] ;
[0147] in It is the recommended treatment option The score, and is the weight coefficient, It is a treatment plan Compatibility with patient P's cancer screening guidelines, It is a treatment plan Similarity of recommendations in detailed cases.
[0148] Example 8, Treatment Plan Compatibility with patient P's cancer screening guidelines The calculation formula is:
[0149] ;
[0150] in Treatment options in cancer screening guidelines The kth applicable condition of is the kth feature of patient P, is the matching indicator function, is the weight of condition k, which is set by expert consensus;
[0151] Treatment options Similarity of recommendations in detailed cases The calculation formula is:
[0152] ;
[0153] in, is the jth historical case with similar characteristics to patient P, For patients P and The feature similarity of For cases Whether treatment plans have been used , if yes, it is 1, otherwise, it is 0.
[0154] Example 9: Based on the statistical data of the intervention effects of similar patients using the corresponding treatment plan in historical data, the expected effect index is calculated using the following formula:
[0155] ;
[0156] in is the expected effect index, is the total number of people historically with similar characteristics to the current patients who adopted this regimen; It is the number of people diagnosed early or with reduced risk in the group using this regimen.
[0157] Through quantitative formulas (recommended treatment plan score, guideline matching degree, case recommendation similarity, expected effect index), a personalized recommendation logic is constructed based on guideline specifications, actual case verification, and quantitative effect evaluation. This solves the problems of traditional treatment plans relying on experience, being highly subjective, and having difficult to quantify effects, improves the accuracy and standardization of treatment plans, and reduces unreasonable recommendations.
[0158] The guideline-matching formula aims to emphasize core criteria crucial for cancer screening by using weights set by expert consensus (e.g., family history weighted at 0.3, specific tumor marker weighted at 0.25), while downplaying the influence of secondary criteria (e.g., body mass index weighted at 0.05). For example, in lung cancer screening, long-term smoking history and CT nodule characteristics are weighted significantly higher than dietary habits. This ensures that the plan strictly aligns with authoritative standards such as the "Cancer Prevention and Screening Guidelines," avoiding over- or under-recommendations due to interference from irrelevant criteria.
[0159] Compared with the traditional method of mechanically applying guidelines, this formula improves the matching accuracy between the plan and the guidelines, greatly increases the proportion of compliance with medical standards, and reduces the medical risks caused by the plan deviating from the guidelines.
[0160] Based on historical data on the early diagnosis rate or risk reduction rate of similar patients using the plan, this provides quantitative evidence for the plan's effectiveness (e.g., this plan increased the early diagnosis rate for similar patients to 82%), rather than vague statements (e.g., "probably effective"). This improves patients' understanding of the plan's effectiveness, allowing doctors to explain its value to patients based on quantitative data, shortening doctor-patient communication time. Furthermore, quantitative results provide objective reference for doctors to adjust the plan, reducing subjective decision-making bias.
[0161] Example 10, a digital tumor screening platform, used to implement a digital tumor screening method, the tumor screening platform includes a patient information entry module, multiple information docking modules, a data analysis module, a tumor risk prediction module, a treatment plan recommendation module and a tumor screening result output module, the patient information entry module is used to enter comprehensive patient data, and the multiple information docking modules are respectively connected to the hospital information system, image archiving and communication system and laboratory information management system, the data analysis module builds a multimodal fusion model, and communicates with the patient information entry module and the multiple information docking modules respectively, and outputs comprehensive tumor correlation features, the tumor risk prediction module builds a tumor risk prediction model, and communicates with the data analysis module, inputs comprehensive tumor correlation features, and outputs the probability of occurrence of different tumors, the treatment plan recommendation module is respectively communicated with the data analysis module and the tumor risk prediction module, and outputs the core recommended plan and the corresponding expected effect, the tumor screening result output module is communicated with the data analysis module, the tumor risk prediction module and the treatment plan recommendation module, and generates a screening report using data visualization technology.
[0162] The following specific embodiments are used to illustrate the implementation principle of the present invention:
[0163] Scenario: Full process of lung cancer screening for 55-year-old men at high risk:
[0164] Patient Zhang, a 55-year-old male, has a 30-year smoking history (20 cigarettes per day). His father has lung cancer (positive family history). He recently developed coughing symptoms and went to a community hospital for cancer screening. The digital cancer screening platform conducted a full-process screening for him. The specific implementation process is as follows:
[0165] 1. Data collection and preprocessing (platform module: patient information entry module + information docking module);
[0166] 1. Data source and type;
[0167] Hospital system docking data:
[0168] The information docking module connects to the community hospital HIS system through a standardized interface (HL7FHIR) to obtain the patient's basic information (age 55 years old, male, family history of lung cancer = 1) and past medical history (no hypertension / diabetes);
[0169] Connect to the PACS system to obtain chest CT images (1mm slice thickness, 300 slices in total);
[0170] Connect to the LIMS system to obtain tumor marker detection data for the past three years (CEA: 5ng / mL in 2021→8ng / mL in 2022→15ng / mL in 2023; CYFRA211: 3.5ng / mL in 2023).
[0171] Patient data entry:
[0172] Patients filled out a lifestyle questionnaire through the platform's mobile app: smoking history for 30 years (20 cigarettes per day), drinking history (3 times a week, 50 mL each time), and exercise frequency (once a week);
[0173] Genetic testing data (sent to a third-party agency, EGFR gene wild type, no mutation).
[0174] 2. Data preprocessing;
[0175] Image data: The platform's data analysis module uses the OpenCV library to perform noise reduction (Gaussian filtering) and standardization (sizing to 512 × 512 pixels) on CT images.
[0176] Time series data: Missing values for CEA and CYFRA211 (data missing for June 2022) were filled using linear interpolation and normalized using the Z score (CEA 2023 Z score = 3.2, marked as a significant anomaly);
[0177] Text data: The 30-year smoking history and family history of lung cancer were converted into 128-dimensional vectors using Jieba word segmentation (smoking history vector value = 0.8, family history vector value = 0.9);
[0178] Gene data: EGFR wild-type was converted into a binary vector (0, no mutation).
[0179] 2. Multimodal fusion feature extraction (platform module: data analysis module);
[0180] The data analysis module builds a multimodal fusion model based on the TensorFlow framework and executes the sub-steps of step 1:
[0181] 1. Image feature extraction (step 11):
[0182] A three-layer CNN (3×3 convolution kernel, 2×2 pooling layer) was used to process CT images, identify an 8 mm nodule (with rough edges) in the right upper lobe of the lung, and output a 256-dimensional image feature vector (including nodule size, density, edge morphology, and other features).
[0183] 2. Temporal feature extraction (step 12):
[0184] The LSTM network is used to process the time series data of CEA (20212023) and CYFRA211 (2023) to capture the continuous upward trend of CEA (annual growth rate of 75%) and output a 128-dimensional time series feature vector.
[0185] 3. Feature splicing and attention mechanism (steps 13-15):
[0186] Concatenate 256-dimensional image features, 128-dimensional time series features, 128-dimensional text vector (smoking history + family history), and 64-dimensional gene vector (EGFR wild type) to form a 512-dimensional initial fusion vector;
[0187] The attention layer dynamically assigns weights: CT image nodule characteristics (weight 0.3), smoking history (weight 0.25), CEA time series trend (weight 0.2), family history (weight 0.15), and other characteristics (weight 0.1);
[0188] After processing by the fully connected layer, a 1024-dimensional tumor-related feature vector (covering comprehensive features of imaging, clinical data, and lifestyle habits) is output.
[0189] 3. Tumor risk prediction (Platform module: Tumor risk prediction module);
[0190] 1. Model training basis:
[0191] The platform uses a random forest algorithm to train a model based on 100,000 labeled data sets (including 50,000 lung cancer patients, 30,000 healthy people, and 20,000 patients with benign lung diseases). The decision tree depth (10 layers) and the number of leaf nodes (30) are adjusted through grid search. After 5-fold cross-validation, the validation set accuracy reached 92% and the recall rate reached 90%.
[0192] 2. Risk probability output:
[0193] Input Zhang's 1024-dimensional feature vector, and the model calculates and outputs the risk probability of multiple tumors:
[0194] Lung cancer: 85% (high risk); stomach cancer: 10% (low risk); colorectal cancer: 8% (low risk); liver cancer: 5% (no risk).
[0195] 4. Risk stratification and early warning (platform module: tumor risk prediction module linked to medical and nursing end);
[0196] 1. Layer determination:
[0197] Based on the thresholds set by the platform (high risk ≥ 70%, medium risk 30%-70%, low risk <30%, no risk <5%), Zhang's lung cancer risk of 85% was judged to be high risk.
[0198] 2. Warning action:
[0199] A real-time pop-up warning window appears on the medical side: Patient Zhang is at high risk of lung cancer (85%) and is recommended for priority treatment. The patient's holographic files (CT images, CEA time series data, family history, etc.) are also linked simultaneously.
[0200] Community doctors can receive early warnings on their mobile devices and directly view the original CT image data (DICOM format) and abnormal indicators (CEA=15ng / mL, Z=3.2).
[0201] 5. Personalized treatment plan recommendation (platform module: treatment plan recommendation module);
[0202] 1. Patient characteristics and guideline extraction (steps 41-42):
[0203] Core patient characteristics: high-risk for lung cancer, male, 55 years old, 30-year smoking history, positive family history of lung cancer, EGFR wild-type;
[0204] The applicable plan is extracted from the "National Cancer Center Lung Cancer Screening Guidelines": low-dose spiral CT (once every six months), CEA+CYFRA211 test (once every three months), applicable condition C: age ≥50 years + smoking history ≥20 years + positive family history.
[0205] 2. Calculation of solution scores (step 43):
[0206] Guideline matching:
[0207] Condition weight: age (0.2) + smoking history (0.3) + family history (0.3) + EGFR status (0.2);
[0208] Zhang completely matches all conditions, MS=0.2+0.3+0.3+0.2=1.0.
[0209] Similar cases recommended similarity SS:
[0210] 100 similar patients (around 55 years old, 30 years of smoking history, positive family history of lung cancer) were matched, 95 of whom adopted this scheme, and the similarity of patient characteristics was =0.9; =0.855.
[0211] 3. Plan and expected results (steps 44-45):
[0212] Core recommendation: low-dose spiral CT (once every six months) + CEA + CYFRA211 testing (once every three months);
[0213] Expected results: Based on historical data (1,000 similar patients using this regimen), =82% (early diagnosis rate 82%, an increase of 40% compared with the non-intervention group).
[0214] 6. Screening report generation (platform module: tumor screening result output module);
[0215] 1. Report content and visualization:
[0216] Risk Overview Page: A radar chart displays the risk of each tumor (lung cancer with an 85% risk is represented by a red sector, while other low-risk groups are represented by green);
[0217] Abnormal data page: Heat map showing the temporal changes of CEA (15 ng / mL in 2023 is marked in red, Z=3.2);
[0218] Recommended plan page: Flowchart showing Week 1: Low-dose CT → Month 3: CEA test → Month 6: CT review;
[0219] Expected results page: A bar chart comparing the adopted plan (early diagnosis rate 82%) and the non-adopted plan (42%).
[0220] 2. Differentiated display between doctors and patients:
[0221] Doctor-side report: Contains 3D reconstruction of CT images and feature importance ranking (smoking history weight 0.25, CT nodule weight 0.3);
[0222] Patient-side APP report: Simplified, your risk of lung cancer is high. It is recommended to have a CT scan every six months, which can greatly increase the probability of early detection. It comes with an appointment portal and a smoking cessation guidance video.
[0223] Zhang completed a low-dose CT scan within one week of the platform's screening and was diagnosed with stage I lung cancer (early stage). With timely surgical treatment, his five-year survival rate is expected to reach 70% (traditional screening may delay the disease to stage II, where the survival rate drops to 50%).
[0224] The platform takes only 4 hours from data collection to solution recommendation (the traditional manual process takes 3 days), and the screening efficiency is improved by 80%;
[0225] Patients’ understanding of the report reached 90% (far exceeding the 30% of traditional reports), and the rate of active cooperation with follow-up increased to 100%.
[0226] This example fully demonstrates the entire process of the digital tumor screening platform from multi-source data fusion to personalized intervention, verifying its core advantages of precision, efficiency, and personalization.
[0227] The above are all preferred embodiments of the present invention, and are not intended to limit the scope of protection of the present invention. Therefore, any equivalent changes made based on the structure, shape, and principle of the present invention should be included in the scope of protection of the present invention.
Claims
1. A digital tumor screening method, characterized by: The following steps are involved: Step 1: Based on a deep learning framework, a multimodal fusion model integrating convolutional neural networks and recurrent neural networks is constructed. The collected comprehensive patient data is input into the multimodal fusion model to extract imaging features and process time series clinical data. Through feature splicing and attention mechanism strategies, the different modal data are deeply fused to output comprehensive tumor correlation features. Step 2: Utilize machine learning algorithms and large-scale tumor data for training to build a tumor risk prediction model, input comprehensive tumor correlation features, and output the probability of different tumors; Step 3: Set stratified risk thresholds, stratify the probability of disease occurrence, and trigger corresponding warning actions for each stratum; Step 4: Based on patient risk stratification, basic information, and tumor screening guidelines, an intelligent recommendation algorithm is used to generate treatment recommendations and expected outcomes. Step 5: Use data visualization technology to generate a screening report, which includes comprehensive patient data, tumor type and disease probability prediction, treatment recommendations, and expected results. Step 4 includes the following sub-steps: Step 41: extract the core characteristics of the patient, including risk stratification, basic information, lifestyle habits, and genetic test results; Step 42: extract the recommended examination or treatment plan, recommended frequency Fq, and applicable population condition C from the tumor screening guidelines for different tumor types; Step 43, using a hybrid recommendation algorithm combining tumor screening guideline matching and similar case collaborative filtering to calculate the recommended treatment plan score; Step 44: Filter projects with recommendation scores greater than a set score threshold, sort them from high to low by score, and form a core recommendation plan; Step 45: Based on the statistics of intervention effects on similar patients in historical data, output the expected effect corresponding to the core recommended plan; In step 43, the formula for calculating the recommended treatment plan score is: ; in It is the recommended treatment option The score, and is the weight coefficient, It is a treatment plan Compatibility with patient P's cancer screening guidelines, It is a treatment plan similarity of recommendations in detailed cases; Treatment options Compatibility with patient P's cancer screening guidelines The calculation formula is: ; in Treatment options in cancer screening guidelines The kth applicable condition of is the kth feature of patient P, is the matching indicator function, is the weight of condition k, which is set by expert consensus; Treatment options Similarity of recommendations in detailed cases The calculation formula is: ; in, is the jth historical case with similar characteristics to patient P, For patients P and The feature similarity of For cases Whether treatment plans have been used , if yes, it is 1, otherwise, it is 0.
2. The digital tumor screening method according to claim 1, characterized in that: Step 1 includes the following sub-steps: Step 11: For the image data, extract the lesion features through the convolution layer and the pooling layer, and output a 256-dimensional image feature vector; Step 12: For the time series clinical data, LSTM is used to process the changing trend of tumor markers over time and output a 128-dimensional time series feature vector; Step 13: Concatenate the image features output by the convolutional neural network, the temporal features output by the recurrent neural network, the text vector, and the gene data vector into a 512-dimensional initial fusion vector; Step 14: Dynamically assign weights through the attention layer; Step 15: After processing by the fully connected layer, a 1024-dimensional tumor comprehensive correlation feature vector is generated.
3. The digital tumor screening method according to claim 2, characterized in that: In step 2, an ensemble learning algorithm is used to input the tumor comprehensive association feature vector generated in step 1, and training is performed with whether the tumor occurs or not as the dependent variable.
4. The digital tumor screening method according to claim 3, characterized in that: In step 2, random forest is used to construct a set number of decision trees, the Gini coefficient is used to evaluate the importance of features, and the importance is ranked; The depth and number of leaf nodes of the decision tree were adjusted through grid search, and 5-fold cross validation was used to avoid overfitting.
5. The digital tumor screening method according to claim 4, characterized in that: In step 3, based on clinical data statistics, the probability of disease is divided into high risk, medium risk, low risk and no risk, among which high risk corresponds to a probability of disease greater than or equal to 70%, medium risk corresponds to a probability of disease greater than or equal to 30% and less than 70%, low risk corresponds to a probability of disease less than 30%, and no risk corresponds to a probability of disease less than 5%.
6. The digital tumor screening method according to claim 5, characterized in that: In step 45, the expected effect index is calculated based on the statistical data of the intervention effect of the corresponding treatment plan on similar patients in the historical data. The calculation formula is: ; in is the expected effect index, is the total number of people historically with similar characteristics to the current patients who adopted this regimen; It is the number of people diagnosed early or with reduced risk in the group using this regimen.
7. A digital tumor screening platform, characterized by: A digital tumor screening method for implementing any one of claims 1-6, the tumor screening platform includes a patient information entry module, multiple information docking modules, a data analysis module, a tumor risk prediction module, a treatment plan recommendation module and a tumor screening result output module, the patient information entry module is used to enter comprehensive patient data, the multiple information docking modules are respectively connected to the hospital information system, the image archiving and communication system and the laboratory information management system, the data analysis module builds a multimodal fusion model, and communicates with the patient information entry module and the multiple information docking modules respectively, outputs comprehensive tumor correlation features, the tumor risk prediction module builds a tumor risk prediction model, and communicates with the data analysis module, inputs comprehensive tumor correlation features, and outputs the probability of occurrence of different tumors, the treatment plan recommendation module is respectively communicated with the data analysis module and the tumor risk prediction module, outputs the core recommended plan and the corresponding expected effect, the tumor screening result output module is communicated with the data analysis module, the tumor risk prediction module and the treatment plan recommendation module, and generates a screening report using data visualization technology.