Pancreatic cancer risk prediction model optimization method based on metabolic syndrome data

By using a pancreatic cancer risk prediction model based on metabolic syndrome data, which comprehensively integrates five physiological abnormalities and the course of diabetes, and uses logistic regression algorithm to generate a visual report, the model solves the problems of inaccurate prediction and lack of interpretability of existing models, and achieves efficient early screening and personalized management of pancreatic cancer.

CN121812150APending Publication Date: 2026-04-07GUANGDONG GENERAL HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing pancreatic cancer risk prediction models fail to fully integrate multiple abnormalities of metabolic syndrome, ignore the differences in the degree of abnormality of each component and their synergistic effects, and lack quantification and integration based on clinical diagnostic criteria, resulting in inaccurate predictions and a lack of interpretability, making them difficult to apply effectively in clinical settings.

Method used

A pancreatic cancer risk prediction model based on metabolic syndrome data was adopted. By acquiring and classifying five physiological abnormalities and combining them with the duration of diabetes, a logistic regression model was trained using machine learning algorithms to generate a visualized risk report and provide personalized intervention suggestions.

Benefits of technology

It significantly improves the accuracy and biological rationality of pancreatic cancer risk prediction, achieves interpretability and operability of the model, supports personalized health management and early screening, and enhances its practicality in clinical practice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121812150A_ABST
    Figure CN121812150A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical artificial intelligence and disease risk prediction, and discloses a pancreatic cancer risk prediction model optimization method based on metabolic syndrome data, and the method comprises the steps: a training stage: obtaining historical clinical data of five physiological abnormalities of a sample individual, carrying out the grading assignment of the historical clinical data based on a metabolic syndrome diagnosis standard, and carrying out the prediction of a pancreatic cancer risk prediction model; in the application stage, the clinical data of a target individual are subjected to same preprocessing and feature engineering and input into the risk prediction model, and the risk prediction model is obtained by combining a pancreatic cancer diagnosis tag and training a logistic regression model through a maximum likelihood estimation method. And calculating to obtain a pancreatic cancer risk assessment result of the target individual. According to the method, data modeling is carried out by innovatively integrating metabolic syndrome data and diabetes disease course information, accurate and early individualized prediction of the pancreatic cancer risk is realized, and an effective tool is provided for clinical screening and intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical artificial intelligence and disease risk prediction technology, and in particular to an optimization method for a pancreatic cancer risk prediction model based on metabolic syndrome data. Background Technology

[0002] In the field of artificial intelligence disease prediction technology, using computer models to mine and analyze massive amounts of clinical data has become an important technical approach to achieve early warning of diseases. Pancreatic cancer is highly malignant and has an extremely poor prognosis. Developing an automated prediction system that can accurately identify high-risk individuals from its related risk factors has significant social value.

[0003] Existing research indicates that metabolic syndrome and its core components are closely related to the occurrence and development of pancreatic cancer. Traditional risk prediction models often rely on single or a few risk factors and fail to systematically integrate metabolic syndrome, which is a whole pathological state containing multiple abnormalities, resulting in an incomplete and inaccurate characterization of individual risk.

[0004] In existing technologies, although there have been some studies that use clinical data to build predictive models, most of these studies only treat metabolic syndrome as a binary "yes / no" diagnostic label, ignoring the differences in the degree of abnormality of its internal components and their synergistic effects. Other studies, although they have introduced machine learning algorithms, have relatively crude feature engineering and lack quantification and integration based on clinical diagnostic criteria. In particular, they have failed to fully consider the dynamic interaction between the key time dimension factor of diabetes duration and metabolic abnormalities, which limits the accuracy and biological rationality of the models.

[0005] Furthermore, the decision-making process of many existing predictive models lacks interpretability, making it difficult for doctors to understand the basis of the model's conclusions, thereby reducing their clinical credibility and willingness to adopt them. At the same time, these models fail to form an effective closed loop between the predictive results and specific clinical follow-up, monitoring, and intervention recommendations, weakening their practical value in clinical scenarios. Summary of the Invention

[0006] The purpose of this invention is to propose an optimization method for pancreatic cancer risk prediction model based on metabolic syndrome data in order to solve the problems in the prior art.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: a method for optimizing a pancreatic cancer risk prediction model based on metabolic syndrome data, comprising a model training phase and a model application phase: The model training phase includes: Historical clinical data on five physiological abnormalities associated with metabolic syndrome were obtained from multiple individuals, including central obesity, hyperglycemia, hypertension, hypertriglyceridemia, and low high-density lipoprotein cholesterol. Based on the historical clinical data, five physiological abnormalities of each individual sample were graded and assigned values ​​according to the predefined diagnostic criteria for metabolic syndrome, resulting in a training feature queue. Based on the training feature queue, the duration of diabetes in the corresponding sample individuals, and the label of whether pancreatic cancer has been diagnosed, a risk prediction model is trained using a machine learning algorithm. The model application phase includes: Clinical data on five physiological abnormalities related to metabolic syndrome were collected from the target individual, and a five-dimensional quantitative vector of the target individual was generated based on the diagnostic criteria for metabolic syndrome in the same manner as in the training phase. Using the same approach as in the training phase, the five-dimensional quantized vector of the target individual is fused with the duration of diabetes into a comprehensive feature vector. The comprehensive pancreatic cancer risk assessment result of the target individual is then obtained through a pre-trained risk prediction model.

[0008] The beneficial effects of the technical solution provided by this invention include at least the following: This invention innovatively introduces interactive features to quantify five core physiological abnormality indicators of metabolic syndrome, enabling mathematical modeling of complex pathophysiological mechanisms. This allows for a more comprehensive capture of the synergistic effect of metabolic disorders and diabetes in driving pancreatic cancer risk, significantly improving the accuracy and biological rationality of risk prediction.

[0009] The risk prediction model of this invention is based on the logistic regression algorithm. Its weight coefficients clearly reflect the contribution of each risk factor. At the same time, it transforms the abstract score into an intuitive risk level through a visualized risk report, and provides differentiated follow-up monitoring and intervention suggestions. It directly supports individualized health management and precise early screening in clinical practice, realizes the transformation of scientific research results into clinical practice, and has good interpretability and operability.

[0010] This invention implements a complete data processing pipeline from raw data input to final decision support, greatly enhancing the practicality and operability of the data processing method in real-world applications, and providing an excellent methodological framework for solving similar complex risk prediction problems. In summary, this invention proposes a novel and systematic solution for predicting the risk of diabetes-related pancreatic cancer. It can not only improve the early diagnosis rate of pancreatic cancer, but also provide a generalizable methodological framework for clinical research on the association between metabolic diseases and tumors, with significant social benefits and clinical application prospects. Attached Figure Description

[0011] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart of a method provided in an embodiment of the present invention. Detailed Implementation

[0013] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a pancreatic cancer risk prediction model optimization method based on metabolic syndrome data proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0014] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0015] The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0016] The following description, in conjunction with the accompanying drawings, details the specific scheme of the pancreatic cancer risk prediction model optimization method based on metabolic syndrome data provided by this invention.

[0017] Please see Figure 1 The diagram illustrates a flowchart of a method for optimizing a pancreatic cancer risk prediction model based on metabolic syndrome data, according to an embodiment of the present invention. The method includes a model training phase and a model application phase, wherein: The model training phase includes: Historical clinical data on five physiological abnormalities associated with metabolic syndrome were obtained from multiple individuals. These abnormalities included central obesity, hyperglycemia, hypertension, hypertriglyceridemia, and low high-density lipoprotein cholesterol. Based on historical clinical data, five physiological abnormalities in each sample individual were graded and assigned values ​​according to predefined diagnostic criteria for metabolic syndrome, resulting in a training feature cohort. A risk prediction model is trained using a machine learning algorithm based on the training feature queue, the duration of diabetes in the corresponding sample individuals, and whether they have been diagnosed with pancreatic cancer. The model application phase includes: Clinical data on five physiological abnormalities related to metabolic syndrome were collected from the target individual, and a five-dimensional quantitative vector of the target individual was generated based on the diagnostic criteria for metabolic syndrome, using the same method as in the training phase. Using the same approach as in the training phase, the five-dimensional quantized vector of the target individual is fused with the duration of diabetes into a comprehensive feature vector. The comprehensive pancreatic cancer risk assessment result of the target individual is then obtained through a pre-trained risk prediction model.

[0018] In one embodiment of the present invention, the historical clinical data includes: abdominal fat accumulation data of individual samples as central obesity data, blood glucose level as hyperglycemia data, systolic and diastolic blood pressure as hypertension data, serum triglyceride concentration as hypertriglyceride data, and serum high-density lipoprotein cholesterol concentration as low high-density lipoprotein cholesterol data.

[0019] It should be noted that the degree of abdominal fat accumulation in an individual can be obtained by measuring their waist circumference, which is a key clinical indicator for assessing central obesity. During the measurement, the subject should stand with their feet about 25-30 cm apart. The measurement site is at the level of the midpoint of the line connecting the upper edge of the iliac crest and the lower edge of the twelfth rib (approximately at the same level as the navel). A soft measuring tape is used, close to the skin without compressing the soft tissue, and is wrapped around the abdomen. The reading is taken at the end of a normal exhalation and is accurate to 0.1 cm. This indicator is an important component in the diagnosis of metabolic syndrome and is closely related to insulin resistance and the risk of pancreatic cancer.

[0020] Fasting blood glucose testing requires subjects to fast for at least 8 hours before collecting venous blood. The plasma glucose concentration is measured using the glucose oxidase method to assess baseline blood glucose status. Glycated hemoglobin (HbA1c) testing measures the percentage of HbA1c in total hemoglobin in whole blood using high-performance liquid chromatography. Glycated hemoglobin reflects an individual's average blood glucose level over the past 2–3 months and is not affected by short-term fluctuations. Combining the two methods can provide a comprehensive assessment of an individual's blood glucose control.

[0021] Hypertension is a core component of metabolic syndrome, closely related to insulin resistance and cardiovascular metabolic risk, and also provides important evidence for pancreatic cancer risk assessment. Blood pressure is usually measured using a calibrated electronic sphygmomanometer or mercury sphygmomanometer, after the subject has been sitting and resting for 5 minutes. Systolic blood pressure and diastolic blood pressure correspond to the readings at phase I (when the Korotkoff sound begins) and phase V (when the sound disappears) of the sphygmomanometer, respectively. The measurement is repeated twice and the average value is taken.

[0022] Dyslipidemia is closely related to systemic inflammation and insulin resistance, and is an important biochemical indicator for assessing metabolic health and pancreatic cancer risk. The study collected venous blood samples from subjects after fasting for 12 hours and used an enzymatic method to measure their serum triglyceride (TG) and high-density lipoprotein cholesterol (HDL-C) concentrations. Subjects were required to avoid alcohol and high-fat diets before the test.

[0023] In one embodiment of the present invention, the step of classifying and assigning values ​​to five physiological abnormalities of each individual sample based on historical clinical data and predefined diagnostic criteria for metabolic syndrome to obtain a training feature queue includes: Identify missing values ​​in historical clinical data and impute them using mean imputation. Outliers in historical clinical data are identified and corrected using outlier detection methods based on reasonable medical ranges. Multiple grading threshold intervals corresponding to five physiological abnormalities were extracted from predefined diagnostic criteria for metabolic syndrome. The five physiological abnormality index values ​​of each individual sample were extracted from historical clinical data after processing missing and outlier values. An ordered numerical score was assigned to each index value based on the grading threshold range it fell into. The numerical scores corresponding to the five physiological abnormalities of each sample individual are concatenated into a five-dimensional quantization vector, and the five-dimensional training feature vectors of all sample individuals are concatenated into a training feature queue.

[0024] It should be noted that mean imputation is a commonly used technique for handling missing data. Specifically, it involves calculating the arithmetic mean of all non-missing values ​​of a specific clinical indicator in the dataset for a given clinical indicator with missing values, and then using this mean to fill in all missing positions of the indicator. This method is simple and efficient, and avoids information loss and statistical bias caused by directly deleting missing samples.

[0025] The outlier identification method based on medically reasonable ranges is an outlier identification strategy that combines clinical knowledge. This invention provides an example using box plots, which specifically includes: first, defining reasonable numerical ranges for various physiological indicators based on authoritative medical guidelines or clinical practice; then calculating the upper and lower quartiles (Q1 and Q3) and the interquartile range (IQR) of the data; considering values ​​less than Q1-1.5IQR or greater than Q3+1.5IQR as statistical outliers; cross-validating the identified outliers with the preset medically reasonable ranges; and only removing values ​​that exceed both limits and are clinically impossible. This method can preserve clinically significant real variation information while ensuring data quality.

[0026] Tiered scoring refers to converting continuous clinical indicators into ordered discrete scores to effectively connect medical diagnostic standards with machine learning models. It unifies five physiological indicators with different units and dimensions into a standardized, structured mathematical expression, including at least a two-level scoring system based on dichotomy and a three-level scoring system based on risk intervals, as illustrated below: Two-level assignment scheme: Multiple ordered levels are divided into two levels, with a value of 0 representing normal and a value of 1 representing abnormal. The specific assignment rules are as follows: Central obesity: Compare waist circumference measurements using the Chinese Diabetes Association standard. If the waist circumference is ≥90cm for men or ≥85cm for women, assign a value of 1; otherwise, assign a value of 0. High blood sugar: Compare fasting blood glucose data. If fasting blood glucose is ≥6.1mmol / L, assign a value of 1; otherwise, assign a value of 0. Hypertension: Compare systolic and diastolic blood pressure data. If systolic blood pressure ≥130 mmHg or diastolic blood pressure ≥85 mmHg, assign a value of 1; otherwise, assign a value of 0. High triglycerides: Compare triglyceride data. If triglycerides are ≥1.7 mmol / L, assign a value of 1; otherwise, assign a value of 0. Low HDL cholesterol: Compare HDL cholesterol data. If the level is <1.0 mmol / L for men or <1.3 mmol / L for women, assign a value of 1; otherwise, assign a value of 0.

[0027] The above five assignment results are combined in the order of [central obesity, hyperglycemia, hypertension, hypertriglyceridemia, and low high-density lipoprotein cholesterol] to generate a five-dimensional quantization vector consisting of 0 and 1, for example [1,0,1,1,0].

[0028] A three-tiered assignment scheme: Multiple ordered levels are represented by three levels, with a value of 0 indicating low risk, 1 indicating medium risk, and 2 indicating high risk. The specific assignment rules are as follows: Central obesity: If the waist circumference of a male is <90cm or the waist circumference of a female is <85cm, the value is assigned to 0; If a male's waist circumference is ≥90cm and <100cm, or a female's waist circumference is ≥85cm and <95cm, the value is assigned as 1; If a male's waist circumference is ≥100cm, or a female's waist circumference is ≥95cm, the value is assigned as 2.

[0029] High blood sugar: If fasting blood glucose is <6.1mmol / L, assign a value of 0; If the fasting blood glucose level is ≥6.1 mmol / L and <7.0 mmol / L, the value is assigned as 1; If fasting blood glucose is ≥7.0 mmol / L, assign a value of 2.

[0030] Hypertension: If systolic blood pressure is <130 mmHg and diastolic blood pressure is <85 mmHg, the value is assigned as 0; If the systolic blood pressure is ≥130 and <139 mmHg, or the diastolic blood pressure is ≥85 and <89 mmHg, the value is assigned as 1; If the systolic blood pressure is ≥140 mmHg or the diastolic blood pressure is ≥90 mmHg, the value is assigned as 2.

[0031] High triglycerides: If triglycerides <1.7 mmol / L, assign a value of 0; If triglyceride levels are ≥1.7 mmol / L and <2.3 mmol / L, assign a value of 1; If triglyceride level is ≥2.3 mmol / L, assign a value of 2.

[0032] Low high-density lipoprotein cholesterol: If HDL-C > 1.2 mmol / L for men or > 1.4 mmol / L for women, the value is assigned to 0. If a male's HDL-C is >1.0 and ≤1.2 mmol / L, or a female's HDL-C is >1.3 and ≤1.4 mmol / L, the value is assigned as 1. If the HDL-C value is ≤1.0 mmol / L for men or ≤1.3 mmol / L for women, the value is assigned as 2.

[0033] Combining the above five assignment results generates a five-dimensional quantized vector consisting of 0, 1, and 2, for example, [2, 1, 0, 2, 1].

[0034] In one embodiment of the present invention, the step of training a risk prediction model using a machine learning algorithm based on a training feature queue, the duration of diabetes in corresponding sample individuals, and whether or not pancreatic cancer has been diagnosed includes: The time from the initial diagnosis of diabetes to the diagnosis of pancreatic cancer in the sample individuals was obtained as the quantitative value of the duration of diabetes. The five-dimensional quantization vector of each sample in the training feature queue is fused with the quantization value of diabetes course into a training feature vector, and the pancreatic cancer diagnosis result is used as the binary classification label of the data record. Using the training feature vector as input features and the binary classification label as the target variable, a logistic regression model is trained using the maximum likelihood estimation algorithm to solve for the regression coefficients and intercept terms of each feature that maximize the likelihood function. The trained logistic regression model is evaluated using a validation dataset that was not involved in the training process. The logistic regression model that meets the preset performance threshold is used as the risk prediction model. At the same time, the final determined regression coefficients, intercept terms and data preprocessing rules are saved as a callable model file.

[0035] It should be noted that during model training, the maximum likelihood estimation algorithm is used to train the logistic regression model. Its core purpose is to find an optimal set of model parameters (i.e., the regression coefficients and intercept terms for each feature) that maximizes the probability of the observed sample label (i.e., whether or not the patient has pancreatic cancer) appearing under the current training data. Specifically: The logistic regression model first maps the linear combination of the input feature vector and its weight coefficients plus intercept term to the (0,1) interval using the Sigmoid function, outputting a predicted probability value (i.e., the predicted probability that the target sample has pancreatic cancer). Then, through maximum likelihood estimation, a likelihood function is constructed for all model parameters. This function quantifies the consistency between the actual labels and their predicted probabilities for all current training samples given the parameters, thus transforming the training process into a numerical optimization problem. Next, iterative algorithms such as gradient descent or quasi-Newton methods are used to continuously adjust the weight coefficients and intercept term to maximize this likelihood function or equivalently minimize its negative logarithm. When the optimization process converges, the solved regression coefficients clearly represent the direction and intensity of the corresponding feature's contribution to the risk of pancreatic cancer, while the intercept term represents the base risk log odds when all features take zero values. This parameter estimation method based on statistical theory gives the model good mathematical interpretability and statistical reliability.

[0036] An example of a judgment that satisfies a preset performance threshold is given below: The area under the receiver operating characteristic curve (AUC) of the logistic regression model on the independent validation set is no less than 0.85; The logistic regression model must satisfy the requirement that the area under the receiver operating characteristic curve is ≥0.85, while also achieving a precision ≥0.80 or a recall ≥0.75 on the independent validation set.

[0037] The area under the receiver operating characteristic curve (AUC) measures the overall performance of a model across all classifiable thresholds. An AUC of 0.5 indicates that the model has no discriminatory power, equivalent to random guessing, while an AUC of 1.0 indicates perfect discrimination. In risk prediction models in the medical field, an AUC ≥ 0.85 is widely considered a high-performance standard. In most cases, if a diabetic patient who is eventually diagnosed with pancreatic cancer is randomly selected, and a diabetic patient who does not have pancreatic cancer is selected, a model with an AUC ≥ 0.85 can correctly classify the former as having a higher risk, ensuring the overall effectiveness of the model in screening at the population level.

[0038] To ensure the effective use of clinical resources and avoid excessive anxiety among patients, and to guarantee the efficiency and ethical nature of subsequent clinical interventions, an accuracy rate of ≥0.80 is defined as the model evaluation standard. Diagnostic examinations for pancreatic cancer are costly and invasive. If the model accuracy rate is too low, it will lead to unnecessary invasive examinations for those who are not diagnosed with the disease, wasting medical resources and causing unnecessary psychological burden on patients.

[0039] A higher recall rate means fewer missed diagnoses by the model. Pancreatic cancer is highly malignant and has a poor prognosis, making early detection crucial. Therefore, the model must be able to identify true potential patients. A recall rate of ≥0.75 is defined as the model evaluation standard, which can strike a balance between identifying as many patients as possible and controlling the false positive rate within an acceptable range. If the standard is too low, too many missed diagnoses will occur, and the model will lose its predictive significance. If the standard is too high, the model may be forced to be too sensitive, producing a large number of false positives. This threshold range is recognized in clinical practice and can significantly improve the early detection rate of pancreatic cancer.

[0040] In one embodiment of the present invention, the step of fusing the five-dimensional quantization vector of each sample individual in the training feature queue with the quantization value of diabetes course into a training feature vector includes: Multiply the quantized values ​​of two different dimensions in the five-dimensional quantization vector, or the quantized value of the diabetes course, with any quantized value in the five-dimensional quantization vector to generate one or more interactive features. The interaction features are added as a new dimension and concatenated with the five-dimensional quantization vector and the quantization value of the diabetes course to obtain a comprehensive feature vector.

[0041] The interactive features include at least the following: a first interactive feature consisting of the product of the central obesity quantification value and the hyperglycemia quantification value; a second interactive feature consisting of the product of the central obesity quantification value and the hypertriglyceridemia quantification value; a third interactive feature consisting of the product of the hyperglycemia quantification value and the hypertriglyceridemia quantification value; and a fourth interactive feature consisting of the product of the quantification value of the diabetes duration and any quantification value in the five-dimensional quantification vector.

[0042] It should be noted that the construction of interaction features should be based on known or hypothesized biological mechanisms between metabolic syndrome, diabetes, and pancreatic cancer, and highly correlated interactions should be avoided. The following is a medical interpretation of the minimum interaction features that should be included: Central obesity × hyperglycemia: Visceral fat secretes inflammatory factors and free fatty acids. Obesity can lead to a certain degree of insulin resistance, which in turn exacerbates hyperglycemia. Central obesity × hypertriglyceridemia: Visceral adipose tissue is the main site of triglyceride synthesis and storage. Obesity is closely related to hypertriglyceridemia and together they reflect the severity of lipid metabolism disorders. High blood sugar × high triglycerides: Insulin resistance inhibits the activity of lipoprotein lipase, leading to impaired triglyceride clearance; Diabetes duration × any dimension quantification vector: A diabetic patient with long-term obesity, high blood sugar / blood pressure / triglycerides or low high-density lipoprotein cholesterol may have more severe cumulative physical damage compared to other patients.

[0043] In one embodiment of the present invention, the steps of collecting clinical data on five physiological abnormalities related to metabolic syndrome in a target individual and generating a five-dimensional quantitative vector of the target individual based on the diagnostic criteria for metabolic syndrome, using the same method as in the training phase, include: Acquire clinical data on five physiological abnormalities associated with metabolic syndrome in the target individual, including the degree of abdominal fat accumulation, blood glucose level, systolic and diastolic blood pressure, serum triglyceride concentration, and serum high-density lipoprotein cholesterol concentration; Missing and outlier values ​​in the clinical data of the target individuals are handled in the same way as in the training phase. Using the same approach as in the training phase, based on the diagnostic criteria for metabolic syndrome, the clinical data of the target individual after processing missing and outlier values ​​are graded and assigned values ​​to generate a five-dimensional quantization vector for the target individual.

[0044] In one embodiment of the present invention, the steps of fusing the five-dimensional quantized vector of the target individual with diabetes disease course information into a comprehensive feature vector, using the same method as in the training phase, and inputting it into the pre-trained risk prediction model to calculate the comprehensive pancreatic cancer risk assessment result of the target individual include: The duration of diabetes diagnosis in a target individual is obtained as a quantitative value of their diabetes course. Using the same method as in the training phase, the five-dimensional quantization vector of the target individual and the quantization value of the diabetes course are fused into a comprehensive feature vector; In the risk prediction model, each dimension of the comprehensive feature vector is multiplied by its corresponding regression coefficient, and all products are added to the model intercept term to obtain a linear combination value. The linear combination value is non-linearly transformed using the Sigmoid function to output a probability value between 0 and 1, which serves as the comprehensive pancreatic cancer risk score for the target individual. A visualized risk prediction report is generated based on the target individual's comprehensive pancreatic cancer risk score.

[0045] It should be noted that the regression coefficients are weight parameters learned through the maximum likelihood estimation method during the model training phase. They are used to quantify the direction and strength of each feature dimension's contribution to pancreatic cancer risk. By multiplying each value in the feature vector with its corresponding regression coefficient, the influence of key predictive factors is amplified, while the influence of irrelevant factors or noise is weakened. Then, the weighted result of all features is added to the intercept term (i.e., the base risk log odds when all feature values ​​are zero) to finally obtain the linear combination value z, which is a real value of any range that integrates the weighted information of all risk factors.

[0046] The Sigmoid function is a classic S-shaped curve function, and its mathematical expression is: Where S(z) is the probability value of the final output of the Sigmoid function, and z is the linear combination value obtained in the previous steps; The function's purpose is to smoothly compress and map the input z-value, regardless of its magnitude (even positive or negative infinity), into a probability value within the (0,1) interval. In this invention, this probability value is interpreted as the risk probability of the target individual having pancreatic cancer. The larger the z-value, the closer the output of the Sigmoid function is to 1, representing a high risk; conversely, the smaller the z-value, the closer the output is to 0, representing a low risk. This non-linear transformation not only satisfies the mathematical requirements of probability output but also aligns with the intuitive logic of medical decision-making, providing doctors with a quantitative, intuitive, and easily interpretable risk score. As one embodiment of the present invention, the step of generating a visualized risk prediction report based on the comprehensive pancreatic cancer risk score of the target individual includes: The comprehensive pancreatic cancer risk score is compared with a predefined risk threshold to classify target individuals into low-risk, medium-risk, or high-risk levels. Based on the risk level, corresponding follow-up monitoring recommendations and intervention measures are generated for the target individuals. The risk score, risk level, follow-up monitoring recommendations, and intervention recommendations are integrated into a single risk prediction report, which is then output as an electronic document or displayed through a graphical user interface.

[0047] It should be noted that the predefined risk threshold is the core criterion connecting quantitative scoring and clinical classification decisions. It is determined based on statistical analysis of the risk score distribution of the retrospective training dataset and combined with the clinical benefit-risk ratio. Common classification methods include classifying individuals with scores in the bottom 50% of the overall distribution as low-risk, the top 10%-15% as high-risk, and the middle portion as medium-risk. This ensures that the positive and negative predictive values ​​corresponding to different levels have clinical significance, so that the classification results can provide a clear and reasonable basis for subsequent differentiated interventions.

[0048] For low-risk individuals, the main recommendations are health education and annual routine check-ups; For individuals at medium risk, it is recommended to strengthen lifestyle interventions and shorten the follow-up period; For high-risk individuals, in addition to lifestyle and drug interventions, targeted early screening is strongly recommended, including tumor marker testing and abdominal imaging. All recommendations regarding medication must cite authoritative clinical guidelines to ensure their scientific validity and standardization.

[0049] Graphical user interfaces (GUIs) are the key medium for human-computer interaction. Their design should adhere to the principles of clarity, intuitiveness, and user-friendliness. Interface elements typically include: ① Risk dashboard, used to display risk level and specific score; ② Bar charts or radar charts are used to illustrate the composition of the five-dimensional quantized vector; ③ A structured list of recommendations, used to list follow-up monitoring recommendations and intervention recommendations in separate bullet points; ④ Data export module, used to automatically generate electronic reports in PDF format.

[0050] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for optimizing a pancreatic cancer risk prediction model based on metabolic syndrome data, characterized in that, This method includes a model training phase and a model application phase: The model training phase includes: Historical clinical data on five physiological abnormalities associated with metabolic syndrome were obtained from multiple individuals, including central obesity, hyperglycemia, hypertension, hypertriglyceridemia, and low high-density lipoprotein cholesterol. Based on the historical clinical data, five physiological abnormalities of each individual sample were graded and assigned values ​​according to the predefined diagnostic criteria for metabolic syndrome, resulting in a training feature queue. Based on the training feature queue, the duration of diabetes in the corresponding sample individuals, and the label of whether pancreatic cancer has been diagnosed, a risk prediction model is trained using a machine learning algorithm. The model application phase includes: Clinical data on five physiological abnormalities related to metabolic syndrome were collected from the target individual, and a five-dimensional quantitative vector of the target individual was generated based on the diagnostic criteria for metabolic syndrome in the same manner as in the training phase. Using the same approach as in the training phase, the five-dimensional quantized vector of the target individual is fused with the duration of diabetes into a comprehensive feature vector. The comprehensive pancreatic cancer risk assessment result of the target individual is then obtained through a pre-trained risk prediction model.

2. The method for optimizing a pancreatic cancer risk prediction model based on metabolic syndrome data according to claim 1, characterized in that: The step of obtaining historical clinical data on five physiological abnormalities associated with metabolic syndrome, including central obesity, hyperglycemia, hypertension, hypertriglyceridemia, and low high-density lipoprotein cholesterol, from multiple individual samples, includes: The historical clinical data includes: abdominal fat accumulation data of the sample individuals as central obesity data, blood glucose level as hyperglycemia data, systolic and diastolic blood pressure as hypertension data, serum triglyceride concentration as hypertriglyceride data, and serum high-density lipoprotein cholesterol concentration as low high-density lipoprotein cholesterol data.

3. The method for optimizing a pancreatic cancer risk prediction model based on metabolic syndrome data according to claim 1, characterized in that: The steps involved in grading and assigning values ​​to five physiological abnormalities in each individual sample based on the historical clinical data and predefined diagnostic criteria for metabolic syndrome, to obtain the training feature cohort, include: Identify missing values ​​in the historical clinical data and fill in the identified missing values ​​using mean imputation. Outliers in the historical clinical data are identified using an outlier detection method based on reasonable medical ranges, and the outliers are corrected. Multiple grading threshold intervals corresponding to the five physiological abnormalities are extracted from the predefined diagnostic criteria for metabolic syndrome. The five physiological abnormality index values ​​of each individual sample are extracted from historical clinical data after processing missing and outlier values. An ordered numerical score is assigned to each index value according to the grading threshold range it falls into. The numerical scores corresponding to the five physiological abnormalities of each sample individual are concatenated into a five-dimensional quantization vector, and the five-dimensional training feature vectors of all sample individuals are concatenated into a training feature queue.

4. The method for optimizing a pancreatic cancer risk prediction model based on metabolic syndrome data according to claim 1, characterized in that: The steps for training the risk prediction model based on the training feature queue, the diabetes history records of the corresponding sample individuals, and the label of whether pancreatic cancer has been diagnosed include: The time from the diagnosis of diabetes to the diagnosis of pancreatic cancer in the sample individuals was obtained as the quantitative value of the duration of diabetes. The five-dimensional quantization vector of each sample in the training feature queue is fused with the quantization value of the diabetes course to form a training feature vector, and the pancreatic cancer diagnosis result is used as the binary classification label of the data record. Using the training feature vector as input features and the binary classification label as target variable, a logistic regression model is trained using the maximum likelihood estimation algorithm to solve for the regression coefficients and intercept terms of each feature that maximize the likelihood function. The trained logistic regression model is evaluated using a validation dataset that was not involved in the training process. The logistic regression model that meets the preset performance threshold is used as the risk prediction model. At the same time, the final determined regression coefficients, intercept terms and data preprocessing rules are saved as a callable model file.

5. The method for optimizing a pancreatic cancer risk prediction model based on metabolic syndrome data according to claim 4, characterized in that: The step of fusing the five-dimensional quantization vector of each sample individual in the training feature queue with the quantization value of the diabetes course to form a training feature vector includes: Multiply the quantized values ​​of two different dimensions in the five-dimensional quantization vector, or the quantized value of the diabetes course, by any quantized value in the five-dimensional quantization vector to generate one or more interactive features. The interactive features are added as a new dimension and concatenated with the five-dimensional quantization vector and the quantization value of the diabetes course to obtain the comprehensive feature vector.

6. The method for optimizing a pancreatic cancer risk prediction model based on metabolic syndrome data according to claim 5, characterized in that: The interactive features include at least: a first interactive feature consisting of the product of the central obesity quantification value and the hyperglycemia quantification value; a second interactive feature consisting of the product of the central obesity quantification value and the hypertriglyceridemia quantification value; a third interactive feature consisting of the product of the hyperglycemia quantification value and the hypertriglyceridemia quantification value; and a fourth interactive feature consisting of the product of the quantification value of the diabetes duration and any quantification value in the five-dimensional quantification vector.

7. The method for optimizing a pancreatic cancer risk prediction model based on metabolic syndrome data according to claim 1, characterized in that: The steps involved in collecting clinical data on five physiological abnormalities related to metabolic syndrome from the target individual, and generating a five-dimensional quantitative vector for the target individual based on the diagnostic criteria for metabolic syndrome, using the same method as in the training phase, include: Acquire clinical data on five physiological abnormalities associated with metabolic syndrome in the target individual, including the degree of abdominal fat accumulation, blood glucose level, systolic and diastolic blood pressure, serum triglyceride concentration, and serum high-density lipoprotein cholesterol concentration; Missing and outlier values ​​in the clinical data of the target individual are processed in the same manner as in the training phase. Using the same approach as in the training phase, based on the aforementioned diagnostic criteria for metabolic syndrome, the clinical data of the target individual after processing missing and outlier values ​​are graded and assigned values ​​to generate a five-dimensional quantization vector for the target individual.

8. The method for optimizing a pancreatic cancer risk prediction model based on metabolic syndrome data according to claim 1, characterized in that: The steps for calculating the comprehensive pancreatic cancer risk assessment result for the target individual, using the same method as in the training phase, involve fusing the five-dimensional quantized vector of the target individual with diabetes disease duration information into a comprehensive feature vector, and inputting this vector into the pre-trained risk prediction model. The duration of the target individual's diabetes diagnosis to the present is obtained as a quantitative value of their diabetes course. Using the same method as in the training phase, the five-dimensional quantization vector of the target individual is fused with the quantization value of the diabetes course into a comprehensive feature vector; In the risk prediction model, each dimension of the comprehensive feature vector is multiplied by its corresponding regression coefficient, and all products are added to the model intercept term to obtain a linear combination value. The linear combination value is non-linearly transformed using the Sigmoid function to output a probability value between 0 and 1, which serves as the comprehensive pancreatic cancer risk score for the target individual. A visualized risk prediction report is generated based on the target individual's comprehensive pancreatic cancer risk score.

9. The method for optimizing a pancreatic cancer risk prediction model based on metabolic syndrome data according to claim 8, characterized in that: The steps involved in generating a visualized risk prediction report based on the target individual's comprehensive pancreatic cancer risk score include: The comprehensive pancreatic cancer risk score is compared with a predefined risk threshold to classify target individuals into low-risk, medium-risk, or high-risk levels. Based on the risk level, corresponding follow-up monitoring recommendations and intervention recommendations are generated for the target individual; The risk score, risk level, follow-up monitoring recommendations, and intervention recommendations are integrated into a single risk prediction report, which is then output as an electronic document or displayed via a graphical user interface.