Method for constructing nomogram pancreatic cystic disease diagnosis model

By constructing a nomotu diagnostic model for pancreatic cystic lesions and combining MRI radiomics and clinical information, the problem of accuracy in differentiating between benign and malignant pancreatic cystic lesions has been solved, achieving a diagnosis with high accuracy and generalization ability, and supporting personalized medicine.

CN120913865APending Publication Date: 2025-11-07抚顺市中心医院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511015350.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Current technologies struggle to accurately differentiate between benign and malignant pancreatic cystic lesions, and imaging methods have limited diagnostic accuracy, leading to misdiagnosis and missed diagnosis, increasing patient suffering and wasting medical resources.

Method used

A nomotu diagnostic model for pancreatic cystic lesions was constructed. Through feature screening, improved secondary discriminant analysis, and combined logistic regression analysis, along with MRI radiomics and clinical information, a visualized diagnostic model was generated to improve diagnostic accuracy and generalization ability.

Benefits of technology

It enables accurate differentiation between benign and malignant pancreatic cystic lesions, reduces misdiagnosis and missed diagnosis, provides personalized medical support, and improves patient survival rate and quality of life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913865A_ABST
    Figure CN120913865A_ABST
Patent Text Reader

Abstract

The invention discloses a nomogram pancreatic cystic disease diagnosis model construction method, which comprises the following steps of S1, acquiring enhanced MRI (Magnetic Resonance Imaging) images and clinical pathological data, and completing data set division; s2, acquiring a focus region of interest for analyzing region calibration; s3, preprocessing the focus region of interest, extracting radiomics features, completing repeatability evaluation and feature screening, and constructing a candidate feature set; s4, an improved quadratic discriminant analysis algorithm is adopted to construct a radiomics model, hyper-parameters are adjusted, indexes such as AUC are evaluated, and radiomics scores are output; s5, screening key clinical variables based on regression analysis, and generating a clinical input feature set; s6, combining radiomics scores and clinical features to construct a combined diagnosis model; and S7, outputting a Nomoh map, performing three-data-set performance evaluation, and verifying model discrimination capability, calibration consistency and clinical effectiveness. The intelligent, accurate and visual pancreatic cystic lesion diagnosis system realizes intelligent, accurate and visual pancreatic cystic lesion diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of medical imaging technology, and in particular to a nomogram pancreatic cystic lesion diagnosis model construction method. BACKGROUND

[0002] Pancreatic cancer (PC) is one of the most deadly cancers worldwide, with incidence and mortality rates almost equal, and a 5-year survival rate of only 13%. It is expected to become the second leading cause of cancer death by 2030. Pancreatic cystic lesions (PCLs), as a key component of pancreatic precancerous lesions, encompass both mucinous and non-mucinous lesions. Among mucinous tumors, intraductal papillary mucinous neoplasms (IPMNs) are the most common, accounting for about half of all PCLs. Non-mucinous tumors include solid-pseudopapillary neoplasms (SPNs) and other types. Some PCLs, such as mucinous cystic neoplasms (MCNs), IPMNs, and cystic neuroendocrine tumors (NETs), have a high malignant potential and can develop into invasive pancreatic ductal adenocarcinoma (PDAC), accounting for 15-20% of pancreatic ductal carcinoma.

[0003] International Pancreatic Association guidelines and many others have clearly defined high-risk signs of PCLs malignancy, such as pancreatic duct dilation, which is of great significance for preoperative differentiation of benign and malignant PCLs. Accurate preoperative identification of PCLs with malignant potential is crucial for improving patient survival and developing personalized treatment plans. Data shows that the 5-year survival rate of patients with non-invasive main pancreatic duct type IPMN after resection is close to 100%, while that of patients with malignant lesions is only 60%.

[0004] However, there are many challenges in achieving accurate preoperative differential diagnosis in clinical practice. Currently, there are no nucleic acid or protein biomarkers in blood that can accurately distinguish between benign and malignant PCLs. CT and MRI are the preferred imaging methods for detecting PCLs, but the imaging features of benign and malignant PCLs are similar, and even experienced radiologists have limited diagnostic accuracy. When CT or MRI diagnosis is unclear, endoscopic ultrasound (EUS) diagnostic methods such as EUS-guided fine needle aspiration (FNA) and cyst fluid analysis are needed, but the amount of cells is often too low to be used for diagnosis, the sensitivity is greatly reduced, and there is a risk of needle tract tumor implantation and bleeding. The limitations of existing methods not only affect medical decision-making, but also increase patient pain and waste medical resources.

[0005] Imaging genomics provides a new way to solve this problem. Imaging genomics based on computer vision technology can automatically extract imaging features from medical images, use machine learning methods to extract high-dimensional quantifiable features from clinical images and data, and generate stable imaging markers for prediction, which has achieved good application results in tumor screening and diagnosis.

[0006] Currently, there is no study to evaluate the role of radiomics in differentiating PCLs using MRI images. According to the latest European evidence-based guidelines, MRI is the preferred examination method for PCLs. Compared with CT, MRI is more sensitive to small PCLs, has higher soft tissue contrast resolution, can better show the solid component of the lesion, and can avoid ionizing radiation and the use of potentially nephrotoxic iodinated contrast agents.

[0007] Therefore, it is urgent to develop a reliable method to classify preoperative PCLs. The present application can realize the accurate identification of benign and malignant PCLs by constructing an MRI radiomics- clinical nomogram pancreatic cystic lesion diagnosis model, which can provide strong support for clinical diagnosis and treatment and promote the development of personalized medicine. SUMMARY

[0008] One object of the present application is to provide a nomogram pancreatic cystic lesion diagnosis model construction method. The present application uses feature selection, improved secondary discriminant analysis and joint Logistic regression analysis to construct a visual diagnostic model, which has the advantages of strong discriminant performance, good model interpretability and applicability to multi-center data.

[0009] According to the nomogram pancreatic cystic lesion diagnosis model construction method of the present application, the following steps are included:

[0010] S1, collect patient data sets from medical centers whose enhanced MRI image diagnosis reports are pancreatic cystic lesions, divide the data sets, and count the clinical and pathological information of the patients;

[0011] S2, obtain the lesion of interest area based on the data set;

[0012] S3, pre-process the lesion of interest area of each patient, extract radiomics features, perform repeatability evaluation and feature selection, and obtain a candidate feature set;

[0013] S4, based on the candidate feature set, an improved secondary discriminant analysis algorithm is used to construct a radiomics model, the hyperparameters are adjusted, and the AUC, accuracy, sensitivity indicators are compared to evaluate the model performance, the optimal radiomics model is selected, and the radiomics score is output;

[0014] S5, perform single-factor and multi-factor Logistic regression analysis on the clinical data of the training set patients, select the clinical variables related to the lesion benignity and malignancy, and generate a clinical input feature set;

[0015] S6, perform joint Logistic regression analysis on the radiomics score output by the optimal radiomics model and the clinical input feature set to construct a joint diagnosis model;

[0016] S7, visualize the joint diagnostic model as radiomics nomogram, draw the receiver operating characteristic curve in the training set, test set and external validation set respectively, calculate the AUC and 95% confidence interval, use the Hosmer-Lemeshow test to evaluate the calibration consistency of the model, calculate the Brier score to measure the prediction bias, combine the decision curve analysis to evaluate the net benefit of the model under different probability thresholds, and verify the clinical effectiveness and practicality of the joint diagnostic model.

[0017] Optionally, the clinical and pathological information of the patient in step S1 includes gender, age, pancreatitis history, jaundice, carbohydrate antigen 19-9 and carcinoembryonic antigen.

[0018] Optionally, the layer thickness of the enhanced MRI image in step S1 is 1.25 mm.

[0019] Optionally, the preprocessing of the lesion region of interest in step S3 includes performing normalization operation, voxel resampling and gray level discretization processing; the radiomics feature extraction includes extracting statistical features, shape features and texture features; the repeatability evaluation adopts intragroup correlation coefficient calculation, and features with intragroup correlation coefficient values not less than 0.75 are retained; the feature screening includes performing screening operation based on feature reproducibility and correlation, and outputting a candidate feature set.

[0020] Optionally, the structure of the improved quadratic discriminant analysis algorithm in step S4 includes:

[0021] S41, based on the candidate feature set, a feature matrix X ∈ R n×d is constructed, where n is the sample number, d is the feature dimension, and is divided into benign sample subset and malignant sample subset according to the benign and malignant diagnosis results of each patient;

[0022] S42, the feature mean vectors μ0, μ1 and the within-class covariance matrices Σ0, Σ1 of the benign sample subset and the malignant sample subset are calculated respectively, a weight adjustment coefficient λ ∈ [0, 1] is introduced, and a harmonic within-class covariance matrix Σ * is constructed.

[0023] S43, singular value decomposition is performed on the harmonic within-class covariance matrix Σ * , Σ * = UΛU T , where U is a column vector matrix composed of the eigenvectors of Σ * , Λ is a diagonal matrix of the eigenvalues of Σ * , the first k eigenvalues are retained to construct a reduced rank covariance matrix Σ k , which is used as a substitute for the within-class covariance term, and a low-rank discriminant space is constructed.

[0024] S44, define the improved secondary discriminant function output sample classification score value f(x);

[0025] S45, perform sample classification score value calculation on all training set samples, judge the belonging category combined with the set score threshold, record the classification result, calculate the accuracy, sensitivity and AUC value of the model in the training set and the test set;

[0026] S46, repeat the calculation of the improved secondary discriminant function under several groups of parameter combinations, select the optimal combination in the search space of the score threshold, the weight adjustment coefficient λ and the reduced dimension k, determine the final discriminant function, evaluate the model performance to obtain the optimal imageomics model, and apply the optimal imageomics model to each sample to output the imageomics score.

[0027] Optionally, the model performance evaluation process in step S4 includes selecting the model with the smallest relative standard deviation as the optimal imageomics model.

[0028] Optionally, the step S5 of generating a clinical input feature set includes:

[0029] S51, extract the original clinical data matrix of the training set patients, and divide it into two types of continuous variables and classification variables according to the variable field;

[0030] S52, perform Shapiro-Wilk test on the continuous variables to determine whether they conform to the normal distribution;

[0031] S53, perform independent sample t-test on the continuous variables conforming to the normal distribution, and perform Mann-Whitney U test on the continuous variables not conforming to the normal distribution, to calculate the difference significance of the continuous variables in benign and malignant patients;

[0032] S54, construct a contingency table for the classification variables, perform Pearson's chi-square test, and calculate the difference significance of the continuous variables in benign and malignant patients;

[0033] S55, combine the variables with difference significance meeting the set threshold into a preliminary screening variable set, perform single factor Logistic regression analysis, and calculate the regression coefficient, standard error and P value of each variable;

[0034] S56, perform multi-factor Logistic regression analysis on single factor variables with P value less than the preset threshold, introduce variables one by one for regression operation by using forward selection method, and screen out feature items with significant collinearity with other variables;

[0035] S57, construct the clinical input feature set from the variables with non-zero coefficients and P values not exceeding the set threshold in the multi-factor regression.

[0036] Optionally, the step S6 specifically comprises: extracting the radiomics score and the corresponding clinical input feature set of each patient in the training set, constructing a joint input vector, inputting the joint input vector into a Logistic regression model, and setting the regression weight and bias parameter corresponding to the radiomics score and each clinical feature variable; calculating the prediction error of each patient in the training set using the lesion benign and malignant label, constructing a model training target based on the cross-entropy loss function, and obtaining the optimal parameter set of the joint regression model by using the maximum likelihood estimation method to perform iterative update; applying the trained joint regression model to the test set and the external validation set samples, respectively calculating the prediction probability value of each patient, and outputting as the discrimination result of the joint diagnosis model.

[0037] Optionally, the step S7 specifically comprises:

[0038] S71, constructing an imageomics nomogram based on the joint diagnosis model, mapping the radiomics score and each clinical input feature on the coordinate axis as the corresponding score scale, setting the total score axis and the risk probability axis, and completing the visualization mapping of the joint model;

[0039] S72, calculating the prediction probability value and the lesion benign and malignant label of each patient in the training set in turn, drawing a receiver operating characteristic curve, and calculating the AUC and the ninety-five percent position confidence interval;

[0040] S73, repeating the receiver operating characteristic curve drawing and AUC calculation operation in the test set, and recording the performance indicators corresponding to the training set;

[0041] S74, performing joint diagnosis model inference in the external validation set, recording the prediction probability value and the lesion benign and malignant label of each patient, drawing a receiver operating characteristic curve, and calculating the AUC and the ninety-five percent position confidence interval;

[0042] S75, drawing the calibration curve of the joint diagnosis model in the training set, the test set and the external validation set respectively, calculating the statistic and the significance level by using the Hosmer-Lemeshow test, and determining the deviation consistency of the prediction probability value and the true label of the joint diagnosis model;

[0043] S76, calculating the Brier score of the joint diagnosis model in each data set, and recording the mean square error result;

[0044] S77, setting a plurality of probability threshold intervals, constructing the corresponding decision curve, calculating the net benefit curve in the training set, the test set and the external validation set respectively, and drawing the net benefit-probability threshold relationship diagram;

[0045] S78. Based on the area under the receiver operating characteristic curve, the Hosmer-Lemeshow test results, and the net benefit curve of the Brier score and the decision curve, determine the discriminative ability, calibration ability, and clinical applicability of the joint diagnostic model in different datasets.

[0046] Optionally, step S77 specifically includes:

[0047] S771. Within the range of predicted probability values ​​output by the joint diagnostic model, set several continuously increasing probability threshold intervals, denoted as T = {t1, t2, ..., t...} k}, where each t i It belongs to the interval [0,1], with a step size no greater than 0.05, covering all possible predicted probability values;

[0048] S772. In the training set, sequentially apply each probability threshold t i Perform the operation, ensuring the predicted probability is not less than t. i Patient samples were labeled as malignant, and the rest were labeled as benign. By comparing the benign or malignant classification labels of the lesions in each patient, a confusion matrix was constructed at the current threshold.

[0049] S773. Calculate the net profit value based on each confusion matrix. The net profit is defined as the difference between the gain from correct identification and the loss from incorrect identification of the joint diagnostic model at the current threshold. (This is repeated for each t-value.) i The corresponding net profit values ​​form the training set net profit curve;

[0050] S774. Repeat steps S772 and S773 on the test set to form the net profit curve of the test set.

[0051] S775. Perform steps S772 and S773 on the external validation set to construct the net profit curve of the validation set.

[0052] S776. Plot the net profit-probability threshold relationship in each dataset, and plot each t... i The corresponding net income is used as the vertical axis and the probability threshold is used as the horizontal axis to form a continuous line graph.

[0053] S777. Mark the maximum net profit point and the corresponding probability threshold on each net profit curve, and record whether the positions of the maximum net profit in the three datasets are consistent and the range of differences.

[0054] S778. Based on the changes in the shape of the net benefit curves in the three datasets, analyze the benefit distribution characteristics of the joint diagnostic model under different probability decision strategies, and form a basis for evaluating the clinical adaptability of the model based on probability threshold classification.

[0055] The beneficial effects of this invention are:

[0056] 1. Improve diagnostic accuracy: By combining MRI image features and clinical information, the diagnostic model can more accurately distinguish between benign and malignant pancreatic cystic lesions, reducing misdiagnosis and missed diagnosis.

[0057] 2. Enhance model generalization: Use multi-center data set for model training and validation, improve the generalization ability of the model, so that it can maintain good diagnostic performance in different medical environments.

[0058] 3. Provide visualization tools: By constructing an image feature nomogram, the complex model prediction results are displayed in an intuitive way to doctors, which is convenient for doctors to understand and apply.

[0059] 4. Promote personalized medicine: The method of the present application provides strong support for clinical diagnosis and treatment, which helps to develop personalized treatment plans and improve the survival rate and quality of life of patients. BRIEF DESCRIPTION OF DRAWINGS

[0060] The accompanying drawings are used to provide a further understanding of the present application, and constitute a part of the specification, together with embodiments of the present application, to explain the present application, and do not constitute a limitation of the present application. In the drawings:

[0061] Fig. 1 A nomogram pancreatic cystic lesion diagnostic model construction method is proposed for the present application, and the overall flow chart of the method is shown in the figure.

[0062] Fig. 2 The 8 optimal image feature correlation coefficient heat map of the nomogram pancreatic cystic lesion diagnostic model construction method proposed by the present application is shown in the figure.

[0063] Fig. 3 The combined diagnostic model image feature nomogram of the nomogram pancreatic cystic lesion diagnostic model construction method proposed by the present application is shown in the figure. DETAILED DESCRIPTION

[0064] The present application will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, and only illustrate the basic structure of the present application in a schematic manner, and therefore only show the components related to the present application.

[0065] Reference Figs. 1-3 A nomogram pancreatic cystic lesion diagnostic model construction method, comprising the following steps:

[0066] S1, collect the enhanced MRI image diagnostic report of the patient data set from the medical center for pancreatic cystic lesions, divide the data set, and count the clinical and pathological information of the patients;

[0067] S2, obtain the lesion region of interest based on the data set;

[0068] S3, pre-process the lesion region of interest of each patient, extract the imaging features, perform repeatability evaluation and feature screening, and obtain a candidate feature set;

[0069] S4, based on the candidate feature set, construct an imaging feature model by using an improved quadratic discriminant analysis algorithm, adjust the hyperparameters, compare the AUC, accuracy, sensitivity indicators, and evaluate the model performance, select the optimal imaging feature model, and output the imaging feature score;

[0070] S5, perform single-factor and multi-factor Logistic regression analysis on the clinical data of the training set patients, screen out clinical variables related to the benign and malignant lesions, and generate a clinical input feature set;

[0071] S6, perform joint Logistic regression analysis on the imaging feature score output by the optimal imaging feature model and the clinical input feature set, and construct a joint diagnostic model;

[0072] S7, visualize the joint diagnostic model as an imaging feature nomogram, draw the receiver operating characteristic curve in the training set, test set and external validation set respectively, calculate the AUC and 95% confidence interval, use the Hosmer-Lemeshow test to evaluate the calibration consistency of the model, calculate the Brier score to measure the prediction bias, combine the decision curve analysis to evaluate the net benefit of the model under different probability thresholds, and verify the clinical effectiveness and practicality of the joint diagnostic model.

[0073] In the embodiment, the clinical and pathological information of the patient in step S1 includes gender, age, pancreatitis history, jaundice, carbohydrate antigen 19-9 and carcinoembryonic antigen.

[0074] In the embodiment, the enhanced MRI image layer thickness in step S1 is 1.25 mm.

[0075] In the embodiment, the pre-processing of the lesion region of interest in step S3 includes performing normalization operation, voxel resampling and gray level discretization processing; the imaging feature extraction includes extracting statistical features, shape features and texture features; the repeatability evaluation uses intragroup correlation coefficient calculation, and the features with intragroup correlation coefficient value not less than 0.75 are retained; the feature screening includes performing screening operation based on feature reproducibility and correlation, and outputting the candidate feature set.

[0076] In the embodiment, the improved quadratic discriminant analysis algorithm structure in step S4 includes:

[0077] S41, based on the candidate feature set, construct a feature matrix X n×d wherein n is the number of samples, d is the feature dimension, and is divided into benign sample subset and malignant sample subset according to the benign and malignant diagnosis results of each patient.

[0078] S42, respectively, the feature mean vector μ0, μ1 and the within-class covariance matrix Σ0, Σ1 are calculated for the benign sample subset and the malignant sample subset, a weight adjustment coefficient λ ∈ [0, 1] is introduced, and a harmonic within-class covariance matrix Σ is constructed * :

[0079] Σ * = λΣ0+ (1- λ)Σ1;

[0080] S43, the harmonic within-class covariance matrix Σ * is singular value decomposed, Σ * = UΛU T , wherein U is a column vector matrix composed of the eigenvectors of Σ * , Λ is a diagonal matrix of the eigenvalues of Σ * , the first k eigenvalues are retained to construct a reduced-rank covariance matrix Σ k , which is used as a substitute within-class covariance term to construct a low-rank discriminant space;

[0081] S44, define the improved quadratic discriminant function to output sample classification score value f(x):

[0082]

[0083] Wherein x is a single sample feature vector, P0, P1 are the prior probabilities respectively;

[0084] S45, the sample classification score value calculation is performed on all training set samples, the classification category is judged by combining the set score threshold value, the classification result is recorded, and the accuracy, sensitivity and AUC value of the model in the training set and the test set are calculated;

[0085] S46, the improved quadratic discriminant function is repeatedly calculated under several groups of parameter combinations, the optimal combination is selected in the search space of the score threshold value, the weight adjustment coefficient λ and the reduced-rank dimension k, the final discriminant function is determined, the model performance is evaluated to obtain the optimal image genomics model, and the optimal image genomics model is applied to each sample to output the image genomics score.

[0086] In the embodiment, the model performance evaluation process in step S4 includes selecting a model with the smallest relative standard deviation as the optimal image genomics model.

[0087] In the embodiment, the generation step of the clinical input feature set in step S5 includes:

[0088] S51, extract the original clinical data matrix of the training set patients, and divide them into two categories of continuous variables and classification variables according to the variable fields;

[0089] S52, perform Shapiro-Wilk test on continuous variables to determine whether they conform to normal distribution, specifically including: grouping each continuous variable according to the benign and malignant labels of the lesions, calculating the Shapiro-Wilk statistic and its corresponding P value for each group of data, setting a significance threshold, and regarding the variable as conforming to normal distribution when the P value is greater than the set threshold, and determining the variable as a non-normal distribution variable when the P value is not greater than the set threshold;

[0090] S53, perform independent sample t-test on continuous variables conforming to normal distribution, and perform Mann-Whitney U test on continuous variables not conforming to normal distribution, to calculate the significance of the difference of the continuous variables between benign and malignant patients, specifically including: for continuous clinical variables determined by Shapiro-Wilk test to conform to normal distribution, using independent sample t-test to compare the mean difference of the variable between the benign sample group and the malignant sample group, calculating the corresponding t-statistic and two-sided P value; for continuous clinical variables not conforming to normal distribution, using Mann-Whitney U test to compare the rank distribution of the benign group and the malignant group samples, calculating the corresponding U statistic and P value. The difference test of all continuous variables is based on the training set samples, and a unified significance threshold is set to determine whether there is a statistically significant difference in the variable between benign and malignant classification;

[0091] S54, construct a contingency table for categorical variables, perform Pearson's chi-square test, and calculate the difference significance of the continuous variable between benign and malignant patients, specifically including: for all discrete type clinical variables in the training set, construct a contingency table between the benign group and the malignant group, count the frequency distribution of each category in the two groups of samples, perform Pearson's chi-square test, calculate the chi-square statistic and the corresponding P value of the variable, and use the set significance threshold as the basis for discrimination to determine whether the distribution difference of the variable under different classification labels has statistical significance. The classification variable with significant difference is recorded as a candidate variable as a preliminary screening basis for constructing a clinical input feature set;

[0092] S55, combine variables with difference significance meeting the set threshold into a preliminary screening variable set, perform single factor Logistic regression analysis, and calculate the regression coefficient, standard error and P value of each variable; specifically including: combining the continuous variables and categorical variables screened by difference test into a preliminary screening variable set, constructing a single factor Logistic regression model, taking the benign or malignant classification label of the patient's lesion as the dependent variable, and sequentially taking each preliminary screening variable as the independent variable to independently construct the regression equation. Perform parameter fitting process on each single variable regression model to calculate the regression coefficient, standard error and two-sided P value of the single variable. Structure the statistical results of all single factor models and sort the variables according to the P value and standard error for input basis for subsequent multi-factor variable screening and joint regression model construction.

[0093] S56, performing multi-factor logistic regression analysis on single-factor variables with P values less than a preset threshold, introducing variables one by one for regression operation by forward selection method, and screening out characteristic items with significant collinearity with other variables; specifically including: extracting variables with P values less than a set threshold from the single-factor logistic regression analysis results to constitute a multi-factor regression candidate variable pool, taking the benign or malignant classification label of the lesion as the dependent variable, initializing the regression model as an empty variable set, introducing the candidate variables into the model one by one according to the forward selection method, calculating the log-likelihood function value and information criterion score of the current model in each round of introduction, recording the regression coefficient, P value and standard error of the introduced variable, performing variance inflation factor calculation on the variables that have entered the model, identifying characteristic items with significant collinearity with other variables, removing the collinear characteristic items from the variable set, repeating the variable introduction and collinearity removal process until all variables that meet the admission conditions are screened, and outputting the final variable group of the multi-factor logistic regression model for constructing the clinical input feature set;

[0094] S57, constructing the clinical input feature set from variables with non-zero coefficients and P values not exceeding a set threshold in the multi-factor regression.

[0095] In this embodiment, the step S6 specifically comprises: extracting the imageomics score and the corresponding clinical input feature set of each patient in the training set, constructing a joint input vector, inputting the joint input vector into a Logistic regression model, and setting the regression weight and bias parameter corresponding to the imageomics score and each clinical feature variable; calculating the prediction error of each patient in the training set using the lesion benign and malignant label, constructing a model training target based on the cross-entropy loss function, performing iterative update using the maximum likelihood estimation method, and obtaining the optimal parameter set of the joint regression model; applying the trained joint regression model to the test set and the external validation set samples, respectively calculating the prediction probability value of each patient, and outputting as the discrimination result of the joint diagnosis model. Wherein, the Logistic regression model is composed of a feature space formed by an input vector, a weight parameter set, a bias term and a Sigmoid mapping function, the input vector is composed of the imageomics score and the screened clinical input feature variable, and the joint input matrix is formed by column splicing; each input feature corresponds to a trainable regression weight parameter, which is used to measure the contribution degree of the feature to the classification result, the bias parameter is an independent term, which controls the overall output offset of the model, the linear combination part of the model is composed of the inner product of the joint input vector and the regression weight vector and the summation of the bias term, the linear combination result is taken as the input of the Sigmoid function, and the output value after the Sigmoid function transformation is limited between zero and one, which represents the prediction probability of the patient as a malignant lesion, the objective function of the model is the cross-entropy loss function, the maximum likelihood estimation is performed on the model parameters based on the joint input vector and the lesion benign and malignant label of the samples in the training set, and the optimal regression coefficient set and the bias term are obtained. The model structure after training can perform joint prediction on the test set and the external validation set samples, and output the lesion risk probability of each patient.

[0096] In this embodiment, the step S7 specifically comprises:

[0097] S71, constructing an imageomics nomogram based on the joint diagnosis model, mapping the imageomics score and each clinical input feature to the corresponding score scale on the coordinate axis, setting the total score axis and the risk probability axis, and completing the joint model visualization mapping;

[0098] S72, calculating the prediction probability value and the lesion benign and malignant label of each patient in the training set in sequence, drawing a receiver operating characteristic curve, calculating the AUC and the ninety-five percent position confidence interval;

[0099] S73, repeating the receiver operating characteristic curve drawing and AUC calculation operation in the test set, and recording the performance indicators corresponding to the training set;

[0100] S74, perform joint diagnosis model inference in the external validation set, record the prediction probability value and lesion benignity or malignancy label of each patient, draw the receiver operating characteristic curve, calculate the AUC and the ninety-five percent position confidence interval;

[0101] S75, draw the calibration curve of the joint diagnosis model in the training set, test set and external validation set respectively, calculate the statistical quantity and significance level by Hosmer-Lemeshow test to determine the deviation consistency of the prediction probability value and the true label of the joint diagnosis model, wherein the calculation process of Hosmer-Lemeshow test is as follows: the prediction probability value of the joint diagnosis model in the training set, test set and external validation set is arranged in ascending order, and is divided into multiple equal-capacity groups, each group contains samples with similar prediction probability value, the observed benign sample number and the cumulative sum of predicted benign probability are calculated for each group respectively, the actual benign frequency and the expected benign frequency of the group are formed, and the chi-square statistical quantity is calculated according to the observed frequency and the expected frequency of each group, the specific calculation method is that the square of the difference of each group in all groups is divided by the sum of the expected frequency, and the corresponding significance level is calculated according to the chi-square statistical quantity and the degrees of freedom obtained by subtracting the number of model parameters from the number of groups, the statistical quantity and P value of Hosmer-Lemeshow test are used to evaluate the consistency between the model output probability and the true label, and the output is used as the calibration curve evaluation index;

[0102] S76, calculate the Brier score of the joint diagnosis model in each data set, and record the mean square error result; wherein the calculation process of Brier score is as follows: the prediction probability value and the lesion benignity or malignancy classification label of each patient in the training set, test set and external validation set are extracted respectively, the prediction output sequence and the true label sequence at the sample level are constructed, and the square error value between the prediction probability value and the corresponding true label of each patient sample is calculated; based on the square error values of all samples, the arithmetic mean value is calculated as the Brier score result under the data set, and the Brier score of each data set is used to measure the mean square error degree between the prediction probability of the joint diagnosis model and the actual classification label, and the lower the value is, the closer the probability output of the model is to the true state;

[0103] S77, set several probability threshold intervals, construct the corresponding decision curve, calculate the net benefit curve in the training set, test set and external validation set respectively, and draw the net benefit-probability threshold relationship diagram;

[0104] S78, according to the area of the receiver operating characteristic curve, the Hosmer-Lemeshow test result, the Brier score and the net benefit curve of the decision curve, the discrimination ability, the calibration ability and the clinical applicability of the joint diagnosis model in different data sets are judged.

[0105] In this embodiment, step S77 specifically includes:

[0106] S771. Within the range of predicted probability values ​​output by the joint diagnostic model, set several continuously increasing probability threshold intervals, denoted as T = {t1, t2, ..., t...} k}, where each t i It belongs to the interval [0,1], with a step size no greater than 0.05, covering all possible predicted probability values;

[0107] S772. In the training set, sequentially apply each probability threshold t i Perform the operation, ensuring the predicted probability is not less than t. i Patient samples were labeled as malignant, and the rest were labeled as benign. By comparing the benign or malignant classification labels of the lesions in each patient, a confusion matrix was constructed at the current threshold.

[0108] S773. Calculate the net profit value based on each confusion matrix. The net profit is defined as the difference between the gain from correct identification and the loss from incorrect identification of the joint diagnostic model at the current threshold. (This is repeated for each t-value.) i The corresponding net benefit values ​​form the training set net benefit curve. The process of calculating the net benefit value based on each confusion matrix is ​​as follows: Set the current probability threshold in the training set, label the samples with a predicted probability value not lower than the threshold as malignant, construct the binary classification prediction results under the threshold, and count the number of true malignant, false malignant, true benign, and false benign lesions in the confusion matrix according to the prediction results and the classification labels of benign or malignant lesions. Based on the confusion matrix results, calculate the net benefit value under the threshold. Define the discrimination behavior of diagnosing malignant as malignant as intervention behavior, set the unit benefit value brought by the intervention, set the unit loss value caused by misjudging as malignant, and construct the net benefit calculation formula as: Net benefit = (number of true malignant lesions × unit benefit - number of false malignant lesions × unit loss) / total number of samples. The net benefit value corresponding to each set threshold is matched one-to-one with the threshold and recorded as a node in the training set net benefit curve, thus constructing a complete net benefit-probability threshold relationship graph.

[0109] S774. Repeat steps S772 and S773 on the test set to form the net profit curve of the test set.

[0110] S775. Perform steps S772 and S773 on the external validation set to construct the net profit curve of the validation set.

[0111] S776. Plot the net profit-probability threshold relationship in each dataset, and plot each t... i The corresponding net income is used as the vertical axis and the probability threshold is used as the horizontal axis to form a continuous line graph.

[0112] S777, marking the maximum net income point and the corresponding probability threshold on each net income curve, recording whether the maximum net income positions in the three data sets are consistent and the difference range;

[0113] S778, according to the shape change of the net income curve in the three data sets, analyzing the income distribution characteristics of the joint diagnostic model under different probability decision strategies, and forming the model clinical adaptability evaluation basis based on probability threshold classification.

[0114] Embodiment 1:

[0115] In order to verify the feasibility of the present application in implementation, the present application is applied to a multi-center pancreatic cystic lesion clinical auxiliary diagnosis research, and a MRI imaging- clinical combined nomogram pancreatic cystic lesion diagnosis model is constructed. The discrimination performance, stability and clinical applicability are verified.

[0116] The research object is a plurality of pathologically diagnosed pancreatic cystic lesion patients. The enhanced MRI images, complete clinical data and postoperative pathological diagnosis information of the patients have been collected. The image images are reviewed by a senior radiologist, and the layer-by-layer delineation of the lesion region of interest is manually completed by another physician. Through image standardization and region calibration, the imaging features are extracted. The repeatability test is performed on the extracted results, and the stable features with an intra-group correlation coefficient not less than 0.75 are reserved, and a candidate feature set is constructed through the feature screening process.

[0117] In the model construction stage, the improved quadratic discriminant analysis algorithm based on the logistic regression framework is used for image feature score generation. By adjusting the regularization parameter and the covariance constraint coefficient, the performance comparison of five parameter combinations is performed, and finally the parameter combination is selected. The model with AUC of 0.914, accuracy of 0.870, sensitivity of 0.893 and F1 score of 0.881 obtained on the training set is selected as the optimal image feature model.

[0118] In the aspect of clinical feature construction, the normality test and difference analysis are performed on the continuous variables, and the chi-square test is performed on the classification variables. A plurality of variables with statistical significance are screened out. Further through single factor and multi-factor logistic regression analysis, three clinical variables are finally selected as joint input features: CEA elevation, lesion wall nodule existence and pancreatic duct dilatation.

[0119] The optimal image-based score and three clinical input variables are input into a joint Logistic regression model to construct a joint diagnostic model and generate a visual nomogram. The model is evaluated by the receiver operating characteristic curve, the calibration curve, the Brier score and the decision curve in the training set, the test set and the external validation set. In the test set, the AUC of the model reaches 0.940, the accuracy is 0.865, the sensitivity is 0.864, and the specificity is 0.867; the Brier score is 0.089, the Hosmer-Lemeshow calibration P value is 0.637, and the decision curve has obvious positive benefits when the probability threshold is between 0.25 and 0.80.

[0120] Further distribution visualization analysis of the image-based score and the clinical feature total score of each patient shows that the median score of the benign patient group is 32, and the IQR interval is 26 to 39; the median score of the malignant patient group is 67, and the IQR interval is 59 to 74, and the difference between the two is statistically significant. According to the score result, 40 is set as the classification threshold, and the joint model has stable resolution ability at this threshold.

[0121] To verify the stability of the model under different sample distributions, the data in the external validation set is divided into three groups, and subset evaluation is performed according to different lesion sites and different patient age layers. In different subsets, the AUC remains above 0.88, and the Brier score is below 0.10, indicating that the model has good generalization ability and application potential.

[0122] The following table shows the main performance indicators of the model in the training set, the test set and the validation set, as well as the stability performance of the joint model in the three subsets. Through the visualization comparison of the indicators, it can be clearly seen that the joint model has good discrimination ability, calibration ability and benefit performance in different data sets and different subgroups, proving that the method proposed in the present application has significant practicality and accuracy in actual scenarios.

[0123] Table 1: Performance evaluation results of the joint diagnostic model in different data sets and subsets

[0124]

[0125] In the training set, the area of the receiver operating characteristic curve of the joint diagnostic model reaches 0.960, the accuracy is 0.882, the sensitivity is 0.911, the specificity is 0.850, the F1 score is 0.889, the Brier score is 0.081, the Hosmer-Lemeshow calibration P value is 0.691, and the net benefit maximum point corresponds to the threshold of 0.43, indicating that the model has high discrimination ability, balance and calibration consistency in the training sample.

[0126] In the test set, the AUC of the model is 0.940, the accuracy is 0.865, the sensitivity is 0.864, the specificity is 0.867, the F1 score is 0.864, the Brier score is 0.089, the calibration P value is 0.637, and the maximum net benefit point threshold is 0.41, indicating that the model maintains good generalization performance on non-training samples and shows robust recognition ability.

[0127] In the external validation set, the model AUC is 0.886, the accuracy is 0.829, the sensitivity is 0.792, the specificity is 0.867, the F1 score is 0.810, the Brier score is 0.097, the calibration P value is 0.682, and the maximum net benefit point is 0.45, which verifies that the model still has strong practicability and stability in the environment of heterogeneous samples.

[0128] In the feature grouping analysis, the model AUC of the pelvic lesion subset is 0.894, the accuracy is 0.845, the sensitivity is 0.823, the specificity is 0.868, the F1 score is 0.833, the Brier score is 0.091, the calibration P value is 0.659, and the maximum net benefit threshold is 0.42; the model AUC of the pancreatic head lesion subset is 0.912, the accuracy is 0.861, the sensitivity is 0.877, the specificity is 0.844, the F1 score is 0.855, the Brier score is 0.087, the calibration P value is 0.701, and the maximum net benefit threshold is 0.44, reflecting that the performance of the model in different anatomical sites of lesions fluctuates less and has good adaptability.

[0129] In the patient population aged ≥60 years, the model AUC is 0.888, the accuracy is 0.837, the sensitivity is 0.806, the specificity is 0.867, the F1 score is 0.821, the Brier score is 0.093, the calibration P value is 0.668, and the maximum net benefit point is 0.47, indicating that the model is also applicable and reliable for the elderly patient population.

[0130] Overall analysis shows that the combined diagnostic model performs evenly in various data sets and clinically stratified populations, with multiple key performance indicators maintaining high standards, especially in Brier score control, net benefit interval stability, and multi-group AUC consistency, which can provide a quantitative and reliable auxiliary tool for risk stratification of pancreatic cystic lesions.

[0131] The above is only the preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can make equivalent replacement or change according to the technical solution and the inventive concept of the present application within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.

Claims

1. A method for constructing a nomogram pancreatic cystic lesion diagnosis model, characterized in that, The method comprises the following steps: S1, collecting a patient data set from a medical center whose enhanced MRI image diagnosis report is pancreatic cystic lesion, performing data set division, and counting clinical and pathological information of the patient; S2, obtaining a lesion region of interest based on the data set; S3, preprocessing the lesion region of interest of each patient, extracting image features, performing repeatability evaluation and feature screening, and obtaining a candidate feature set; S4, constructing an image feature model based on the candidate feature set combined with an improved quadratic discriminant analysis algorithm, adjusting the hyperparameters, comparing AUC, accuracy, sensitivity indicators, and evaluating the model performance, selecting the optimal image feature model, and outputting an image feature score; S5, performing single-factor and multi-factor Logistic regression analysis on the clinical data of the training set patients, screening out clinical variables related to the benign and malignant lesions of the lesion, and generating a clinical input feature set; S6, performing joint Logistic regression analysis on the image feature score output by the optimal image feature model and the clinical input feature set, and constructing a joint diagnosis model; S7, visualizing the joint diagnosis model as an image feature nomogram, drawing a receiver operating characteristic curve in the training set, test set and external validation set respectively, calculating AUC and 95% confidence interval, using Hosmer-Lemeshow test to evaluate the calibration consistency of the model, calculating Brier score to measure the prediction bias, combining decision curve analysis to evaluate the net benefit of the model under different probability thresholds, and verifying the clinical effectiveness and practicality of the joint diagnosis model.

2. The method according to claim 1, wherein the method is characterized by, The clinical and pathological information of the patient in step S1 includes gender, age, history of pancreatitis, jaundice, carbohydrate antigen 19-9 and carcinoembryonic antigen.

3. The method according to claim 2, wherein the method is characterized by, The layer thickness of the enhanced MRI image in step S1 is 1.25 mm.

4. The nomogram pancreatic cystic lesion diagnosis model construction method according to claim 3, characterized in that, The preprocessing of the lesion region of interest in step S3 includes performing normalization operation, voxel resampling and gray level discretization processing; the image feature extraction includes extracting statistical features, shape features and texture features; the repeatability evaluation adopts intragroup correlation coefficient calculation, and the features with intragroup correlation coefficient value not less than 0.75 are reserved; The feature screening includes performing screening operation based on feature reproducibility and correlation, and outputting a candidate feature set.

5. The method according to claim 4, wherein the method is characterized by, The improved quadratic discriminant analysis algorithm structure in step S4 includes: S41、based on the candidate feature set, constructing a feature matrix X ∈ R n×d wherein n is the number of samples, d is the feature dimension, and is divided into a benign sample subset and a malignant sample subset according to the benign and malignant diagnosis results corresponding to each patient. S42, respectively, the benign sample subset and the malignant sample subset are calculated the feature mean vector μ0, μ1 and the within-class covariance matrix Σ0, Σ1, the weight adjustment coefficient λ ∈ [0, 1] is introduced, and the harmonic within-class covariance matrix Σ is constructed * : S43, the harmonic covariance matrix Σ * performing singular value decomposition, where U is a column vector matrix composed of the eigenvectors of Σ * , Λ is a diagonal matrix of the eigenvalues of Σ * , the first k eigenvalues are reserved to construct a reduced rank covariance matrix Σ k , and a low rank discriminant space is constructed as an alternative intra-class covariance term. S44, defining an improved quadratic discriminant function to output sample classification score value f(x); S45, performing sample classification score value calculation on all training set samples, combining the set score threshold to judge the belonging category, recording the classification result, calculating the accuracy, sensitivity and AUC value of the model in the training set and test set; S46, repeatedly calculating the improved quadratic discriminant function under several groups of parameter combinations, selecting the optimal combination in the search space of the score threshold, the weight adjustment coefficient λ and the reduced dimension k, determining the final discriminant function, evaluating the model performance to obtain the optimal image feature model, and applying the optimal image feature model to each sample to output the image feature score.

6. The nomogram pancreatic cystic lesion diagnosis model construction method according to claim 5, characterized in that, The process of evaluating the model performance in step S4 includes selecting the model with the smallest relative standard deviation as the optimal image feature model.

7. The method according to claim 6, wherein the method is characterized by, The step S5 of generating the clinical input feature set includes: S51, extracting the original clinical data matrix of the training set patients, and dividing into two categories of continuous variables and classification variables according to variable fields; S52, performing Shapiro-Wilk test on the continuous variables to determine whether they conform to normal distribution; S53, performing independent sample t-test on the continuous variables conforming to normal distribution, and performing Mann-Whitney U test on the continuous variables not conforming to normal distribution to calculate the difference significance of the continuous variables in benign and malignant patients; S54, constructing a contingency table for the classification variables, performing Pearson chi-square test, and calculating the difference significance of the continuous variables in benign and malignant patients; S55, combining the variables with difference significance satisfying the set threshold into a preliminary screening variable set, performing single factor Logistic regression analysis, and calculating the regression coefficient, standard error and P value of each variable; S56, performing multi-factor Logistic regression analysis on the single factor variables with P value less than the preset threshold, introducing the variables one by one for regression operation by using forward selection method, and screening out feature items with significant collinearity with other variables; S57, constructing the clinical input feature set from the variables with non-zero coefficient and P value not more than the set threshold in the multi-factor regression.

8. The method according to claim 7, wherein the method is characterized by, The step S6 specifically includes: extracting the imageomics score of each patient in the training set and the corresponding clinical input feature set, constructing a joint input vector, inputting the joint input vector into a Logistic regression model, setting the regression weight and bias parameter corresponding to the imageomics score and each clinical feature variable; calculating the prediction error of each patient in the training set based on the lesion benignity and malignancy label, constructing a model training target based on the cross-entropy loss function, performing iterative update by using the maximum likelihood estimation method, and obtaining the optimal parameter group of the joint regression model; applying the trained joint regression model to the test set and the external validation set samples, respectively calculating the prediction probability value of each patient, and outputting as the discrimination result of the joint diagnosis model.

9. The method according to claim 8, wherein the method is characterized by, The step S7 specifically includes: S71, constructing an imageomics nomogram based on the joint diagnosis model, mapping the imageomics score and each clinical input feature on the coordinate axis to the corresponding score scale, setting the total score axis and the risk probability axis, and completing the visualization mapping of the joint model; S72, calculating the prediction probability value and the lesion benignity and malignancy label of each patient in the training set in turn, drawing a receiver operating characteristic curve, and calculating AUC and ninety-five percent position confidence interval; S73, repeating the receiver operating characteristic curve drawing and AUC calculation operation in the test set, and recording the performance indicators corresponding to the training set; S74, performing joint diagnosis model reasoning in the external validation set, recording the prediction probability value and the lesion benignity and malignancy label of each patient, drawing a receiver operating characteristic curve, and calculating AUC and ninety-five percent position confidence interval; S75, calculating the prediction probability value and the lesion benignity and malignancy label of each patient in the test set in turn, drawing a receiver operating characteristic curve, and calculating AUC and ninety-five percent position confidence interval. S75, draw calibration curves of the joint diagnostic model in the training set, test set and external validation set respectively, calculate the statistical quantity and significance level by Hosmer-Lemeshow test to determine the deviation consistency of the prediction probability value and the true label of the joint diagnostic model; S76, calculate the Brier score of the joint diagnostic model in each data set and record the mean square error result; S77, set several probability threshold intervals, construct the corresponding decision curve, calculate the net income curve in the training set, test set and external validation set respectively, and draw the net income-probability threshold relationship graph; S78, determine the discrimination ability, calibration ability and clinical applicability of the joint diagnostic model in different data sets according to the area under the receiver operating characteristic curve, the Hosmer-Lemeshow test result, the Brier score and the net income curve of the decision curve.

10. The method according to claim 9, wherein the method is characterized by, The step S77 specifically comprises: S771、In the prediction probability value range output by the joint diagnosis model, set several continuous increasing probability threshold intervals, denoted as T={t1, t2, …, t k},wherein each t i belongs to the interval [0, 1], the step is not greater than 0.05, covering all possible prediction probability values; S772、In the training set, each probability threshold t is sequentially processed i Patients with a prediction probability not less than t i are labeled as malignant, and the rest are labeled as benign. By comparing the benign or malignant classification labels of each patient's lesion, a confusion matrix under the current threshold is constructed. S773、based on each confusion matrix, calculate a net benefit value, the net benefit being defined as the difference between the benefit brought by correct joint diagnosis model discrimination and the loss brought by incorrect joint diagnosis model discrimination under the current threshold, and calculate all t i The corresponding net benefit values constitute a training set net benefit curve; S774, repeat the step S772 and the step S773 in the test set to form the test set net income curve; S775, execute the step S772 and the step S773 in the external validation set to construct the validation set net income curve; S776. Plot the net profit-probability threshold relationship in each dataset, and plot each t... i The corresponding net income is used as the vertical axis and the probability threshold is used as the horizontal axis to form a continuous line graph. S777, mark the maximum net income point and the corresponding probability threshold on each net income curve, and record whether the maximum net income positions in the three data sets are consistent and the difference range; S778, analyze the income distribution characteristics of the joint diagnostic model under different probability decision strategies according to the shape change of the net income curve in the three data sets to form the model clinical applicability evaluation basis based on the probability threshold classification.