Patient information data management platform and method for vascular surgery
By integrating multi-source patient data, deep processing and stratified variable training models, the problems of low data utilization and neglect of individual differences in existing technologies are solved, accurate and personalized prediction and report generation of postoperative complications are achieved, and the efficiency and accuracy of clinical decision-making are improved.
Patent Information
- Application Number
- CN202510900146.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-09-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing vascular surgery, existing technologies have difficulty in effectively integrating multi-source data and ignore individual differences and key medical imaging features, resulting in a lack of personalization and accuracy in postoperative complication prediction results.
The data processing module collects and preliminarily organizes patient health data, and uses the feature extraction module for in-depth processing to calculate the degree of vascular stenosis and identify plaque types. Combined with texture analysis technology, a patient data set is constructed, and the probability prediction module sets hierarchical variables to train global and local models to generate a postoperative complication probability assessment report.
It improves the accuracy and personalization of postoperative complication prediction, provides detailed evaluation reports to assist clinical decision-making, and enhances the efficiency and accuracy of clinical decision-making.
Smart Images

Figure CN120708910A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of patient information data management for vascular surgery, and more particularly, to a patient information data management platform and method for vascular surgery. Background Art
[0002] In modern medical practice, especially in vascular surgery, accurate prediction of postoperative complications is crucial to improving treatment outcomes. However, existing technologies often rely on a single data source or simple statistical models, which makes it difficult to fully consider individual patient differences and complex physiological characteristics.
[0003] Existing technologies have significant shortcomings in processing patient information and predicting postoperative complications, primarily in three key areas: First, these systems often struggle to effectively integrate data from diverse sources, resulting in low data utilization; second, most prediction models fail to fully account for individual differences between patients, making predictions lacking in personalization and precision; and finally, traditional analysis methods often overlook key medical imaging features, such as the degree of vascular stenosis and plaque type, which are crucial for accurately predicting postoperative complications. These issues collectively limit the effectiveness and reliability of existing technologies in clinical practice. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art and achieve the above-mentioned objectives, the present invention provides the following technical solutions: a patient information data management platform and method for vascular surgery, comprising:
[0005] Data processing module: used to collect comprehensive health data of patients and perform preliminary sorting to form a multi-source patient data set;
[0006] Feature extraction module: This module is used to perform deep processing on multi-source patient datasets. Based on the deep-processed multi-source patient datasets, the degree of vascular stenosis is calculated, and texture analysis technology is used to identify plaque types and size distributions. The vascular branching pattern and vascular wall thickness characteristics are further analyzed to obtain image features. Correlation analysis is performed on the text data and numerical data to construct the patient dataset.
[0007] Probability prediction module: Based on the patient data set, stratified variables are set to obtain different subsets. Postoperative complication prediction models are trained based on these subsets. The global model is trained using all subsets, and the local model is trained using a single subset. The global model outputs a universal postoperative complication probability result, while the local model outputs a personalized postoperative complication probability result.
[0008] Report generation module: used to generate a postoperative complication probability assessment report based on the postoperative complication probability results, providing a basis for clinical decision-making.
[0009] Furthermore, the comprehensive health data includes electronic medical record data, imaging examination data, vital sign monitoring data, laboratory test result data, and medication reaction record data;
[0010] Electronic medical record data includes patients' personal information, medical history, diagnosis records, and treatment plans; imaging examination data includes vascular structure image data obtained through imaging technology;
[0011] Vital sign monitoring data include blood pressure, heart rate, and blood oxygen saturation; laboratory test result data include test results of blood and urine samples; and medication response record data include the patient's response to different medications;
[0012] Initial cleaning includes using data processing tools to remove duplicates, standardize data formats, and fill in missing values;
[0013] Integrate the initially collated comprehensive health information into a multi-source patient dataset.
[0014] Furthermore, the method of constructing the patient data set includes:
[0015] Step 31: Perform deep processing on the text, numerical, and image data in the multi-source patient dataset, including:
[0016] For text data, use regular expressions to remove special symbols, punctuation marks, and non-alphanumeric characters from the text, and convert all text to lowercase. Use spelling checkers to correct spelling errors, and then apply natural language processing techniques to extract key terms and entities from the text, convert them into embedding vectors, and use the embedding vectors as text features.
[0017] Among them, key terms include disease name, treatment method and drug name, and entities include patient name, hospital affiliation and diagnosis date;
[0018] For numerical data, the Z-score standardization method is used to standardize the numerical data, and the standardized numerical data is used as the numerical feature;
[0019] For image data, Gaussian filter is applied to remove image noise, and histogram equalization is used to enhance image contrast;
[0020] Step 32: Based on all text features, numerical features, and image features, potential connections between features are obtained through association analysis, and features are further screened;
[0021] Step 33: Integrate all types of filtered features and new features into a patient dataset.
[0022] Furthermore, the method for calculating the degree of vascular stenosis includes:
[0023] The internal adjustment function is defined by the length of the contour and curvature Obtained by weighted summation; where LK is the contour, is the initial contour, which is represented by n points ( , ), i=1,2,...,n represents the number of points on the contour;
[0024] The length formula is: ;in, represents the total length of the profile LK, Refers to the loop operation, which means that when i=n, the distance between the last point and the first point is calculated, and mod represents the operator;
[0025] The curvature formula is:
[0026] ;
[0027] in, Represents the curvature of the contour LK. For each point ( , ), i represents the position index of the center point of the current curvature calculation, and It represents the coordinate after the average smoothing of three points. Take one point before and after it and perform local smoothing to get and , , , represents the Gaussian weight function, , represents the preset standard deviation parameter, and j represents the distance of other points in the formula relative to the center point;
[0028] The external adjustment function is defined as ;
[0029] in, Represents the negative gradient magnitude of the image, defined as , Representing an image At the point The gradient at A negative sign indicates that the contour tends to move toward the direction of high gradient;
[0030] According to the adjustment function, the position of each point on the contour is updated using the gradient descent method. When the difference between the updated contour position and the last updated contour position is less than the preset change threshold, the contour is determined to have converged.
[0031] According to the contour position after convergence, the coordinates of the contour points are extracted. For each contour point, along the normal direction Search the edge of the blood vessel and get two points PD1 and PD2, PD1=( , )+d× , PD2=( , ) d× ;
[0032] Among them, d represents the preset search distance, normal Represents a straight line perpendicular to the tangent direction of the curve at the contour point;
[0033] The vascular diameter is obtained by calculating the distance between PD1 and PD2, and the degree of vascular stenosis is obtained by calculating the ratio of the difference between the maximum and minimum vascular diameters in the image to the maximum vascular diameter.
[0034] Furthermore, the method of using texture analysis technology to identify plaque type and size distribution and further analyzing blood vessel branching pattern and blood vessel wall thickness characteristics includes:
[0035] The image data is converted into a grayscale image, and the gray-level co-occurrence matrix method is used to count the number of adjacent occurrences of pixels of different brightness in the grayscale image. Then, the texture features in the image are extracted based on the gray-level co-occurrence matrix, including the intensity of local changes, the correlation of gray levels between adjacent pixels, and the uniformity of image texture.
[0036] The extracted texture features are used as input and a support vector machine is used to train a plaque classification model. Different types of plaques are distinguished and identified based on the plaque classification model, and the area and type of the plaque are output.
[0037] Use image segmentation technology to separate the plaques from the image and calculate the area of each plaque. Then group the plaques according to their area and type, and count the number of plaques in each group. The area and grouping results of the plaques are used as the size and distribution characteristics of the plaques.
[0038] The grayscale image is converted into a binary image using a global threshold segmentation method, and the binary image is processed using a skeletonization algorithm to extract the connectivity path of the vascular structure.
[0039] For each branch in the connected path, measure the distance from one branch point to another branch point or endpoint to obtain the branch length of each branch;
[0040] Among them, a branch point is a node on a connected path where paths in at least three directions intersect, and an endpoint is a point on a connected path where only one direction of the path extends;
[0041] For adjacent branches in a connected path, the angle between each pair of adjacent branches is calculated as the branch angle;
[0042] The branch lengths and branch angles of the blood vessels were horizontally spliced to obtain the vascular branching pattern;
[0043] Based on the image data, the inner and outer boundary points of the blood vessel wall in the image are identified and extracted by edge detection technology. The polynomial surface is used as the fitting model according to the extracted boundary point data to fit the height of the blood vessel wall zb= ;
[0044] Among them, u and v are the coordinates of the boundary point, zb represents the height of the blood vessel wall at the boundary point, c and b are the exponents of u and v respectively, and U and V are the highest powers of u and v respectively obtained by k-fold cross validation. are the fitting parameters to be estimated, Represents the height of the fitting surface at the boundary point. The sum of the squares of the vertical distances between the boundary point and the fitting surface is minimized by the least squares method to estimate the optimal fitting parameters. ;
[0045] According to the boundary point coordinates (u, v , and the height of the fitted surface at the boundary points and The absolute difference between and is used to obtain the thickness of the blood vessel wall at the boundary point (u, v);
[0046] The average thickness of the vascular wall at all boundary points is taken as the vascular wall thickness.
[0047] Furthermore, the method of obtaining potential connections between features through association analysis and further screening features includes:
[0048] Based on all text features, numerical features and image features, a correlation matrix of all features is constructed. Each position in the matrix represents the Spearman correlation coefficient between the corresponding two features.
[0049] Based on the constructed correlation matrix, the feature pairs whose absolute values of the Spearman correlation coefficient are greater than the preset correlation threshold are identified as highly correlated feature pairs;
[0050] Remove one feature from each highly correlated feature pair, and calculate the multicollinearity values of all remaining features until the multicollinearity values of all features are lower than the preset collinearity threshold.
[0051] Methods for calculating multicollinearity values include:
[0052] Select any unselected feature as the dependent variable and all other features as independent variables. Use the multiple linear regression model to perform multiple linear regression on the dependent variable and all independent variables to obtain the linear regression determination coefficient.
[0053] The reciprocal of the difference between 1 and the linear regression coefficient of determination was taken as the multicollinearity value of the dependent variable characteristics;
[0054] According to the remaining features after the removal process, the multicollinearity value of the remaining features is calculated;
[0055] If the multicollinearity value of a feature is greater than the collinearity threshold, it is removed and the collinearity values of all remaining features are recalculated. The calculation is iterated until the multicollinearity values of all remaining features are less than or equal to the collinearity threshold.
[0056] Furthermore, the method of setting stratification variables based on the patient data set to obtain different subsets includes:
[0057] Based on all the characteristics of each patient after screening, stratification variables were defined, including the patient's age, sex, and medical history;
[0058] Among them, the age of the patients is divided into nl groups according to the preset age range, and the patients are divided into male and female groups according to gender. Each group is used as a stratification variable to obtain different age stratification variables and gender stratification variables;
[0059] All types of diseases in the medical history of all patients were counted. Each type of disease was used as an independent medical history stratification variable. For each patient and each medical history stratification variable, if the patient had the corresponding disease, it was recorded as 1, and if the patient did not have the corresponding disease, it was recorded as 0;
[0060] According to the defined stratification variables, each type of stratification variable is combined. For each combination, an independent subset is generated, and all patients are assigned to the corresponding subset. Each subset contains patients with the same stratification type and their screened characteristics.
[0061] For each subset, if the number of patients in the subset is less than the preset sample number threshold, it is marked as a less common subset, and all less common subsets are merged into a new subset;
[0062] For each subset, the polynomial feature generation tool is used to transform the features of each patient in the subset, converting the single features into polynomial features and interaction features of power terms, and using the generated polynomial features and interaction features of all patients as the input of the subset.
[0063] Furthermore, the method of training the postoperative complication prediction model includes:
[0064] Collect the data for training and divide it into different subsets based on the stratification variables, where each subset represents a group of samples with similar characteristics. Merge all subsets into a global dataset and merge the inputs of all subsets into the input of the global dataset.
[0065] The global dataset is divided into training and test sets in proportion, and the global model part of the postoperative complication prediction model is constructed using the random forest model. The input of the global dataset is used as the input of the global model to output the estimated probability of postoperative complications in patients.
[0066] Initialize the model's hyperparameters and use Bayesian optimization to tune the hyperparameters. Use k-fold cross-validation to evaluate the cross-validation scores of the model under different hyperparameter combinations and select the optimal parameter combination.
[0067] Define the cross entropy loss function as the loss function to evaluate the prediction accuracy of the model. In each iteration, the model is forward propagated using the training set to calculate the loss function between the predicted value and the true value, and the model parameters are updated through backpropagation.
[0068] Use the R2 score as the evaluation metric and calculate the R2 score of the current iteration on the validation set. For each iteration, calculate the difference between the R2 score value after the current iteration and the R2 score value of the previous iteration, which is recorded as the iteration difference;
[0069] Set the iteration difference threshold. If the iteration difference is greater than the iteration difference threshold, it is determined that the performance of the model has improved; if the iteration difference is less than or equal to the iteration difference threshold, it is determined that the performance of the model has not improved.
[0070] If the performance of the model on the validation set does not improve in consecutive DD iterations, the training is stopped and the trained global model is obtained;
[0071] The global dataset is redivided into different subsets. For each subset, the global model training process is repeated to obtain a local model of the postoperative complication prediction model for the corresponding subset. The output of the local model is a personalized estimate of the probability of postoperative complications for each patient in the corresponding subset.
[0072] Furthermore, the method of generating the postoperative complication probability assessment report includes:
[0073] Based on the comprehensive health information of the new patient, the global model of the postoperative complication prediction model is used to predict the patient's estimated probability of postoperative complications, and then the local model of the postoperative complication prediction model is used to predict the patient's personalized estimated probability of postoperative complications;
[0074] The average of the prediction results of the global model and the local model is taken as the final estimate of the patient's probability of postoperative complications;
[0075] Set a risk threshold. If the final estimated value is less than or equal to the preset risk threshold, the patient is judged to have a low probability of postoperative complications and a safety signal is generated.
[0076] If the final estimated value is greater than the preset risk threshold, the patient is judged to have a high probability of postoperative complications and a danger signal is generated;
[0077] Generate a visual assessment report on the probability of postoperative complications, including the patient's comprehensive health information, the estimated probability of postoperative complications predicted by the global model, the personalized estimated probability of postoperative complications predicted by the local model, and the final estimated probability of postoperative complications;
[0078] When dangerous signals are detected, warnings will be issued in the postoperative complication probability assessment report to remind doctors.
[0079] Furthermore, a method for managing patient information data for vascular surgery is characterized by comprising:
[0080] S1. Collect comprehensive health data of patients and conduct preliminary collation to form a multi-source patient data set;
[0081] S2. Deeply process the multi-source patient dataset. Based on the deeply processed multi-source patient dataset, calculate the degree of vascular stenosis and use texture analysis technology to identify plaque type and size distribution. Further analyze the vascular branching pattern and vascular wall thickness characteristics to obtain image features. Combined with text data and numerical data for correlation analysis, a patient dataset is constructed.
[0082] S3. Based on the patient data set, set stratification variables to obtain different subsets, and train postoperative complication prediction models based on the different subsets. Specifically, train the global model using all subsets, and then train the local model using a single subset. The global model outputs a universal postoperative complication probability result, and the local model outputs a personalized postoperative complication probability result.
[0083] S4. Generate a postoperative complication probability assessment report based on the postoperative complication probability results to provide a basis for clinical decision-making.
[0084] The technical effects and advantages of the patient information data management platform and method for vascular surgery of the present invention are as follows:
[0085] The present invention aims to deeply explore the characteristics of vascular diseases and predict the probability of postoperative complications by integrating and analyzing the patient's comprehensive health data, and then generate an evaluation report to assist clinical decision-making. First, by collecting and preliminarily organizing the patient's comprehensive health data, a multi-source patient data set containing all relevant information is formed, ensuring the quality and consistency of the data, and providing a reliable foundation for subsequent feature extraction and model training; secondly, by deeply processing multi-source data and performing multiple processing on the image data (including calculating the degree of vascular stenosis, identifying plaque types and their size distribution), features closely related to vascular health are extracted, providing rich input information for subsequent predictions; then, the patient data is divided into different subsets according to the stratification variables. , and construct a global model for predicting postoperative complications for all subsets and a local model for different subsets, combining the advantages of the global and local models to provide a more accurate prediction of the probability of postoperative complications, fully considering the individual differences of patients; finally, a detailed postoperative complication probability assessment report is generated based on the prediction results (including the patient's comprehensive health information, the prediction results of the global model and the local model), providing an intuitive and easy-to-understand result display, helping doctors to quickly formulate treatment plans and issue warnings when necessary, thereby improving the efficiency and accuracy of clinical decision-making. The present invention not only improves the accuracy and personalization level of postoperative complication prediction, but also provides a powerful tool to support the clinical decision-making process, with significant application value and technical advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 This is a schematic diagram of a patient information data management platform for vascular surgery according to the present invention;
[0087] Figure 2 A schematic diagram of a patient data set obtained by constructing a patient information data management platform for vascular surgery according to the present invention;
[0088] Figure 3 This is a schematic diagram of a patient information data management method for vascular surgery according to the present invention. DETAILED DESCRIPTION
[0089] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0090] Example 1
[0091] See also Figure 1 and Figure 2As shown, the patient information data management platform for vascular surgery described in this embodiment includes:
[0092] Data processing module: used to collect comprehensive health data of patients and perform preliminary sorting to form a multi-source patient data set;
[0093] Feature extraction module: This module is used to perform deep processing on multi-source patient datasets. Based on the deep-processed multi-source patient datasets, the degree of vascular stenosis is calculated, and texture analysis technology is used to identify plaque types and size distributions. The vascular branching pattern and vascular wall thickness characteristics are further analyzed to obtain image features. Correlation analysis is performed on the text data and numerical data to construct the patient dataset.
[0094] Probability prediction module: Based on the patient data set, stratified variables are set to obtain different subsets. Postoperative complication prediction models are trained based on these subsets. The global model is trained using all subsets, and the local model is trained using a single subset. The global model outputs a universal postoperative complication probability result, while the local model outputs a personalized postoperative complication probability result.
[0095] Report generation module: used to generate a postoperative complication probability assessment report based on the postoperative complication probability results, providing a basis for clinical decision-making;
[0096] Comprehensive health data includes electronic medical record data, imaging examination data, vital sign monitoring data, laboratory test results data, and medication response record data;
[0097] Electronic medical record data includes patients' personal information, medical history, diagnosis records, and treatment plans. Imaging examination data includes vascular structure image data obtained through imaging technologies (such as CT angiography (CTA) and magnetic resonance angiography (MRA)). These data are obtained using the API interface provided by the hospital information system (such as HIS, LIS, PACS) or standard protocols such as HL7 and FHIR.
[0098] Vital sign monitoring data includes blood pressure, heart rate, and blood oxygen saturation (connected to vital sign monitoring devices (such as smart watches and blood pressure monitors) to upload data in real time via Bluetooth or Wi-Fi technology); laboratory test results include blood and urine sample test results (such as blood lipid levels and blood glucose concentrations); and medication response records include patients' responses to different medications (including but not limited to side effects and efficacy);
[0099] Initial cleaning includes using data processing tools (such as Python's Pandas library and OpenRefine) to remove duplicates, standardize data formats, and fill in missing values;
[0100] Integrate the initially collated comprehensive health information into a multi-source patient dataset;
[0101] It is used to deeply process multi-source patient datasets, apply image analysis technology to identify and quantify features related to vascular health status, and combine text data and numerical data for correlation analysis. The methods for constructing patient datasets include:
[0102] Step 31: Perform deep processing on the text, numerical, and image data in the multi-source patient dataset, including:
[0103] For text data, use regular expressions to remove special symbols, punctuation, and non-alphanumeric characters from the text, and convert all text to lowercase. Use spellcheck tools to correct spelling errors, and then apply natural language processing techniques (such as TF-IDF, Word2Vec, and BERT) to extract key terms and entities from the text, convert them into embedding vectors, and use the embedding vectors as text features.
[0104] Among them, key terms include disease name, treatment method and drug name, and entities include patient name, hospital affiliation and diagnosis date;
[0105] For numerical data, the Z-score standardization method is used to standardize the numerical data, and the standardized numerical data is used as the numerical feature;
[0106] For image data, Gaussian filter is applied to remove image noise, and histogram equalization is used to enhance image contrast;
[0107] Step 32: Based on the deeply processed multi-source patient dataset and image data, image analysis techniques are used to calculate the degree of vascular stenosis, texture analysis techniques are used to identify plaque type and size distribution, and vascular branching patterns and vascular wall thickness characteristics are further analyzed to obtain image features. The image features include vascular geometry features and plaque features. The vascular geometry features include the degree of vascular stenosis, vascular branching patterns, and vascular wall thickness. The plaque features include plaque type, plaque size, and distribution.
[0108] Step 33: Based on all text features, numerical features, and image features, potential connections between features are obtained through association analysis, and features are further screened;
[0109] Step 34: Integrate all types of filtered features and new features into a patient dataset;
[0110] The method for calculating the degree of vascular stenosis includes:
[0111] Based on the multi-source patient dataset after deep processing and the image data, the global threshold segmentation method is applied to binarize the image, extract the vascular area in the image, and use morphological operations (such as dilation and erosion) to refine the vascular contour of the vascular area to obtain the initial contour of the blood vessel. ;
[0112] in, is a series of points ( , ), i=1,2,...,n represents the number of points on the contour;
[0113] According to the initial contour, an adjustment function of the contour is defined, and the position of the contour is adjusted by the adjustment function. The value of the adjustment function is the sum of the internal adjustment function and the external adjustment function;
[0114] The internal adjustment function is defined as the length term of the contour and curvature term Obtained by weighted summation;
[0115] The length term formula is: ;in, represents the total length of the contour LK; is the Euclidean distance between two points, indicating the distance from point ( , ) to point ( , )’s line segment length; Represents a mathematical operation used to loop through the elements of an indexed array or list. Here, i represents the index of the current point, n is the number of points (i.e., the number of points on the contour), and mod is the modulo operator, which returns the remainder after dividing two numbers. Its purpose is to calculate the distance from the last point to the first point. When calculating the total length of a closed curve, the starting and ending points of the curve need to be connected to form a complete closed path. If this is not done, the distance between the last and first points may be ignored in the calculation, resulting in inaccurate results. By minimizing the total length of the contour, the contour can be made smoother, avoiding unnecessary twists and turns and noise.
[0116] By using the length term formula , which ensures that when calculating the total length of a closed curve, the last point can be correctly connected to the first point, thus forming a complete closed path. This is very important for calculating the total length of vascular contours in the vascular surgery patient information data management platform, because it ensures that all points are taken into account and no distance is missed. This adjustment improves the accuracy of the calculation and makes the final result more reliable.
[0117] The formula for the curvature term is:
[0118] ;
[0119] Here, we use the range i=2 to n-2 to avoid boundary effects and ensure that there are enough points to calculate the curvature. Represents the curvature of the contour LK. For each point ( , ), i represents the position index of the center point of the current curvature calculation, and It represents the coordinate after the average smoothing of three points. Take one point before and after it and perform local smoothing to get and , , , represents the Gaussian weight function, , represents the preset standard deviation parameter, and j represents the distance of the point relative to the center point (in the optimized curvature calculation formula, the center point refers to the point around which the curvature is currently calculated); What is calculated is the growth of the cross product between each set of three consecutive points, which reflects the relative position relationship between the three points. The numerator calculates the sum of the cross products between the points in the local area, which reflects the relative position relationship between the vectors formed by these points. The denominator calculates the cube root of the sum of the squares of the distances between the points in the local area, which is used for normalization to ensure that the curvature is not affected by scale changes. The Euclidean distance squared is calculated between two adjacent points by using Gaussian weights The contribution of each point is weighted and summed to ensure that points closer to the center have a greater impact on the result;
[0120] It should be noted that the curvature term formula is obtained by improving and optimizing the original curvature term formula. The original curvature term formula is , the advantages of the optimized formula include:
[0121] The original formula directly uses the coordinates of the original points, which is easily affected by noise. By performing three-point averaging smoothing on each point, the impact of noise on the curvature calculation is reduced and the stability of the result is improved. The original formula only uses three points and may not fully reflect the local geometric features. By expanding the local area to five points (instead of three points), the geometric features of the curve can be better captured, especially in complex shapes or with large noise. The original formula treats the contributions of all points equally and does not consider the relative importance of the points. By introducing the Gaussian weight function, the points close to the center point have a greater impact on the curvature calculation, which enhances the robustness and accuracy of the calculation.
[0122] The external adjustment function is defined as ;
[0123] in, Represents the negative gradient magnitude of the image, defined as , Representing an image At the point The gradient at A negative sign indicates that the contour tends to move toward the direction of high gradient; represents the integral;
[0124] According to the adjustment function, the position of each point on the contour is updated using the gradient descent method. When the difference between the updated contour position and the last updated contour position is less than the preset change threshold, the contour is determined to have converged.
[0125] For each point on the contour ( , ), and its update rules are as follows:
[0126] ;in, represents the learning rate, k represents the number of iterations, and Represents the adjustment function right and The partial derivative of Refers to partial derivative;
[0127] According to the contour position after convergence, the coordinates of the contour points are extracted. For each contour point, along the normal direction Search the edge of the blood vessel and get two points PD1 and PD2, PD1=( , )+d× , PD2=( , ) d× ;
[0128] Among them, d represents the preset search distance, normal Represents a straight line perpendicular to the tangent direction of the curve at the contour point, the contour point ( , ) is in the direction of the tangent line , then the normal direction is , the distance between the vessel walls can be measured using the normal direction to obtain the vessel diameter;
[0129] The vascular diameter is obtained by calculating the distance between PD1 and PD2, and the difference between the maximum and minimum vascular diameters in the image is calculated by the ratio of the maximum vascular diameter to the maximum vascular diameter to obtain the degree of vascular stenosis.
[0130] Methods for using texture analysis technology to identify plaque types and size distribution and further analyze vascular branching patterns and vessel wall thickness characteristics include:
[0131] Convert the image data into a grayscale image and apply the grayscale co-occurrence matrix method to count the number of times pixels of different brightness appear adjacent to each other in the grayscale image (specifically, calculate the probability of the grayscale value combination between each pixel and its neighboring pixels (such as the pixel on the right or below), that is, count how many times in the image, when each pixel has a certain brightness value, the pixel next to it has another brightness value. For example, if a pixel has a brightness value of 100 and an adjacent pixel has a brightness value of 120, record this pair of brightness values (100, 120), and then repeat this process throughout the image). For each pair of such pixel combinations, the number of times they appear together is counted. For example, there are 100 pixels in an image, and there are two pairs of brightness values (100, 120) and (110, 130). After statistics, it is found that there are 12 pairs of (100, 120) and 10 pairs of (110, 130) in the image. Then, the texture features in the image are extracted based on the gray-level co-occurrence matrix, including the intensity of local changes (contrast), the correlation of gray levels between adjacent pixels (correlation), and the uniformity of image texture (energy). These features can effectively describe the patch characteristics in the image;
[0132] The extracted texture features are used as input and a support vector machine is used to train a patch classification model. Based on the patch classification model, different types of patches are distinguished and identified, and the area and type of the patches are output. (This step aims to accurately identify various patches in the image and classify and label them.)
[0133] Use image segmentation techniques (such as threshold segmentation or region growing) to separate the plaques from the image and calculate the area of each plaque. Group the plaques according to their area and type, and count the number of plaques in each group. Use the area and grouping results of the plaques as the size and distribution characteristics of the plaques.
[0134] The grayscale image is converted into a binary image using a global threshold segmentation method, and the binary image is processed using a skeletonization algorithm to extract the connectivity path of the vascular structure.
[0135] For each branch in the connected path, measure the distance from one branch point to another branch point or endpoint to obtain the branch length of each branch;
[0136] Among them, a branch point is a node on a connected path where paths in at least three directions intersect, and an endpoint is a point on a connected path where only one direction of the path extends;
[0137] For adjacent branches in a connected path, the angle between each pair of adjacent branches is calculated as the branch angle;
[0138] The branch lengths and branch angles of the blood vessels were horizontally spliced to obtain the vascular branching pattern;
[0139] Based on the image data, the inner and outer boundary points of the blood vessel wall in the image are identified and extracted by edge detection technology. The polynomial surface is used as the fitting model according to the extracted boundary point data to fit the height of the blood vessel wall zb= ;
[0140] Among them, u and v are the coordinates of the boundary point, zb represents the height of the blood vessel wall at the boundary point, c and b are the exponents of u and v respectively, and U and V are the highest powers of u and v respectively obtained by k-fold cross validation. are the fitting parameters to be estimated, Represents the height of the fitting surface at the boundary point. The sum of the squares of the vertical distances between the boundary point and the fitting surface is minimized by the least squares method to estimate the optimal fitting parameters. (In mathematics and computer science, surface fitting refers to constructing a mathematical model (usually a surface) from a set of known data points so that the model can describe the distribution trend of these data points as accurately as possible. Surface fitting is widely used in various fields, such as engineering design, computer graphics, and medical image processing.)
[0141] According to the boundary point coordinates (u, v) and the actual height zb, as well as the height of the fitting surface at the boundary point , calculate zb and The absolute difference between and is used to obtain the thickness of the blood vessel wall at the boundary point (u, v);
[0142] The average thickness of the vascular wall at all boundary points is taken as the vascular wall thickness;
[0143] The potential connections between features are obtained through association analysis, and the methods for further feature screening include:
[0144] Based on all text features, numerical features and image features, a correlation matrix of all features is constructed. Each position in the matrix represents the Spearman correlation coefficient between the corresponding two features.
[0145] Based on the constructed correlation matrix, identify the feature pairs whose absolute value of the Spearman correlation coefficient is greater than the preset correlation threshold (usually, when the correlation coefficient is greater than 0.8 or less than -0.8, the feature pair can be called highly correlated) as highly correlated feature pairs;
[0146] Gradually remove one feature from the highly correlated feature pairs, calculate the multicollinearity values of all remaining features, and stop when the multicollinearity values of all features are lower than the preset collinearity threshold (such as 5 or 10), and obtain the remaining features after removal processing;
[0147] Methods for calculating multicollinearity values include:
[0148] Select any unselected feature as the dependent variable and all other features as independent variables. Use the multiple linear regression model to perform multiple linear regression on the dependent variable and all independent variables to obtain the linear regression determination coefficient.
[0149] The form of the multiple linear regression model is: , where Y is the dependent variable, is the independent variable, is the intercept term, is the coefficient corresponding to each independent variable;
[0150] The reciprocal of the difference between 1 and the linear regression coefficient of determination was taken as the multicollinearity value of the dependent variable characteristics;
[0151] According to the remaining features after the removal process, the multicollinearity value of the remaining features is calculated;
[0152] If the multicollinearity value of a feature is greater than the collinearity threshold, it is removed and the collinearity values of all remaining features are recalculated. The calculation is iterated until the multicollinearity values of all remaining features are less than or equal to the collinearity threshold.
[0153] Based on the patient data set, stratification variables are set to obtain different subsets in the following ways:
[0154] Based on all the characteristics of each patient after screening, stratification variables were defined, including the patient's age, sex, and medical history;
[0155] Among them, the age of the patients is divided into nl groups according to the preset age range, and the patients are divided into male and female groups according to gender. Each group is used as a stratification variable to obtain different age stratification variables and gender stratification variables;
[0156] All types of diseases in the medical history of all patients were counted. Each type of disease was used as an independent medical history stratification variable. For each patient and each medical history stratification variable, if the patient had the corresponding disease, it was recorded as 1, and if the patient did not have the corresponding disease, it was recorded as 0;
[0157] According to the defined stratification variables, each type of stratification variable is combined and for each combination, an independent subset is generated (for example, if age is divided into 3 groups, sex is divided into 2 groups, and cm types of medical history are considered, theoretically 3×2×2 cm subsets. In fact, because not all combinations will appear, the actual number of subsets will be less than this theoretical value), and all patients are assigned to the corresponding subsets. Each subset contains patients with the same stratification type and their screened characteristics.
[0158] For each subset, if the number of patients in the subset is less than the preset sample size threshold (set by industry experts based on experience), it is marked as a rare subset, and all rare subsets are merged into a new subset (because some subsets may have small data volumes, which may lead to overfitting or poor model performance when training the model. By merging rare subsets, data sparsity problems can be effectively prevented);
[0159] For each subset, a polynomial feature generation tool (such as PolynomialFeatures in Scikit-learn) is used to transform the features of each patient in the subset, converting single features into polynomial features of power terms and interaction terms. The generated polynomial features and interaction terms of all patients are used as input to the subset to train the postoperative complication prediction model, and the output is the probability of postoperative complications. (By defining stratification variables to divide patients into different subsets and creating polynomials and interaction terms between features in each subset to achieve personalized postoperative complication prediction, this not only captures more complex patterns and interactions between features, helping to improve the personalized prediction ability of the model, but also better captures differences between different groups.)
[0160] Methods for training postoperative complication prediction models include:
[0161] Collect the data for training and divide it into different subsets based on the stratification variables, where each subset represents a group of samples with similar characteristics. Merge all subsets into a global dataset and merge the inputs of all subsets into the input of the global dataset.
[0162] The global dataset is divided into training and test sets in proportion, and the global model part of the postoperative complication prediction model is constructed using the random forest model. The input of the global dataset is used as the input of the global model to output the estimated probability of postoperative complications in patients.
[0163] Initialize the model's hyperparameters and use Bayesian optimization to tune the hyperparameters. Use k-fold cross-validation to evaluate the cross-validation scores of the model under different hyperparameter combinations and select the optimal parameter combination.
[0164] Define the cross entropy loss function as the loss function to evaluate the prediction accuracy of the model. In each iteration, the model is forward propagated using the training set to calculate the loss function between the predicted value and the true value, and the model parameters are updated through backpropagation.
[0165] Use the R2 score as the evaluation metric and calculate the R2 score of the current iteration on the validation set. For each iteration, calculate the difference between the R2 score value after the current iteration and the R2 score value of the previous iteration, which is recorded as the iteration difference;
[0166] Set the iteration difference threshold. If the iteration difference is greater than the iteration difference threshold, it is determined that the performance of the model has improved; if the iteration difference is less than or equal to the iteration difference threshold, it is determined that the performance of the model has not improved.
[0167] If the performance of the model on the validation set does not improve in consecutive DD iterations, the training is stopped and the trained global model is obtained;
[0168] The global dataset is re-divided into different subsets. For each subset, the global model training process is repeated to obtain a local model of the postoperative complication prediction model for the corresponding subset. The output of the local model is a personalized estimate of the probability of postoperative complications for each patient in the corresponding subset.
[0169] Methods for generating a postoperative complication probability assessment report include:
[0170] Based on the comprehensive health information of the new patient, the global model of the postoperative complication prediction model is used to predict the patient's estimated probability of postoperative complications, and then the local model of the postoperative complication prediction model is used to predict the patient's personalized estimated probability of postoperative complications;
[0171] The average of the prediction results of the global model and the local model is taken as the final estimate of the patient's probability of postoperative complications;
[0172] A risk threshold is set (set by industry experts or doctors based on experience. To ensure sufficient redundancy, a very small value can be set to prevent unexpected situations). If the final estimated value is less than or equal to the preset risk threshold, the patient is judged to have a low probability of postoperative complications and a safety signal is generated.
[0173] If the final estimated value is greater than the preset risk threshold, the patient is judged to have a high probability of postoperative complications and a danger signal is generated;
[0174] Generate a visual assessment report on the probability of postoperative complications, including the patient's comprehensive health information, the estimated probability of postoperative complications predicted by the global model, the personalized estimated probability of postoperative complications predicted by the local model, and the final estimated probability of postoperative complications;
[0175] When dangerous signals are detected, warnings will be issued in the postoperative complication probability assessment report to remind doctors.
[0176] This embodiment collects and preliminarily organizes comprehensive patient health data to form a multi-source patient dataset containing all relevant information, ensuring data quality and consistency and providing a reliable foundation for subsequent feature extraction and model training. Secondly, by deeply processing multi-source data and performing various image processing operations (including calculating the degree of vascular stenosis and identifying plaque type and size distribution), features closely related to vascular health are extracted, providing rich input information for subsequent predictions. Then, the patient data is divided into different subsets based on stratification variables, and a global model for predicting postoperative complications is constructed for all subsets, as well as local models for different subsets. Combining the advantages of global and local models, this provides a more accurate prediction of the probability of postoperative complication occurrence, fully accounting for individual patient differences. Finally, a detailed postoperative complication probability assessment report is generated based on the prediction results (including the patient's comprehensive health information, the prediction results of the global model, and the local model). This provides an intuitive and easy-to-understand result presentation, helping doctors quickly formulate treatment plans and issuing warnings when necessary, thereby improving the efficiency and accuracy of clinical decision-making. This not only improves the accuracy and personalization of postoperative complication predictions, but also provides a powerful tool to support the clinical decision-making process, demonstrating significant application value and technical advantages.
[0177] Example 2
[0178] See also Figure 3 As shown, for parts not described in detail in this embodiment, please refer to the description of Example 1. A method for managing patient information data for vascular surgery is provided, comprising:
[0179] S1. Collect comprehensive health data of patients and conduct preliminary collation to form a multi-source patient data set;
[0180] S2. Deeply process the multi-source patient dataset. Based on the deeply processed multi-source patient dataset, calculate the degree of vascular stenosis and use texture analysis technology to identify plaque type and size distribution. Further analyze the vascular branching pattern and vascular wall thickness characteristics to obtain image features. Combined with text data and numerical data for correlation analysis, a patient dataset is constructed.
[0181] S3. Based on the patient data set, set stratification variables to obtain different subsets, and train postoperative complication prediction models based on the different subsets. Specifically, train the global model using all subsets, and then train the local model using a single subset. The global model outputs a universal postoperative complication probability result, and the local model outputs a personalized postoperative complication probability result.
[0182] S4. Generate a postoperative complication probability assessment report based on the postoperative complication probability results to provide a basis for clinical decision-making.
[0183] Example 3
[0184] This embodiment discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the operation mode of the above-mentioned method for managing patient information data for vascular surgery is implemented.
[0185] Since the electronic device described in this embodiment is an electronic device used to implement a method for managing patient information data for vascular surgery in the embodiment of this application, those skilled in the art will be able to understand the specific implementation of the electronic device of this embodiment and its various variations based on the method for managing patient information data for vascular surgery in the embodiment of this application. Therefore, how the electronic device implements the method in the embodiment of this application will not be described in detail here. As long as those skilled in the art can implement the electronic device used in the method for managing patient information data for vascular surgery in the embodiment of this application, it falls within the scope of protection to be provided by this application.
[0186] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field according to actual conditions.
[0187] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for users of ordinary skill in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A patient information data management platform for vascular surgery, characterized in that: include: Data processing module: used to collect comprehensive health data of patients and perform preliminary sorting to form a multi-source patient data set; Feature extraction module: This module is used to perform deep processing on multi-source patient datasets. Based on the deeply processed multi-source patient datasets, the degree of vascular stenosis is calculated, and texture analysis technology is used to identify plaque type and size distribution. The vascular branching pattern and vascular wall thickness characteristics are further analyzed to obtain image features. Correlation analysis is performed on the text data and numerical data to construct a patient dataset. The image features include vascular geometry and plaque features. The vascular geometry features include the degree of vascular stenosis, vascular branching pattern, and vascular wall thickness. The plaque features include plaque type, plaque size, and distribution. Among them, based on the image data of the multi-source patient dataset after deep processing, the global threshold segmentation method is applied to binarize the image, the vascular region in the image is extracted, and the vascular contour of the vascular region is refined using morphological operations to obtain the initial contour of the blood vessel; based on the initial contour, a contour adjustment function is defined, and the position of the contour is adjusted by the adjustment function. The value of the adjustment function is the sum of the internal adjustment function and the external adjustment function; according to the adjustment function, the contour position is updated, and the coordinates of the contour points are extracted based on the updated contour position to obtain the degree of vascular stenosis; Probability prediction module: Based on the patient data set, stratified variables are set to obtain different subsets. Postoperative complication prediction models are trained based on these subsets. The global model is trained using all subsets, and the local model is trained using a single subset. The global model outputs a universal postoperative complication probability result, while the local model outputs a personalized postoperative complication probability result. Report generation module: used to generate a postoperative complication probability assessment report based on the postoperative complication probability results, providing a basis for clinical decision-making.
2. A vascular surgery patient information data management platform according to claim 1, characterized in that: The comprehensive health data includes electronic medical record data, imaging examination data, vital sign monitoring data, laboratory test result data, and drug reaction record data; Electronic medical record data includes patients' personal information, medical history, diagnosis records, and treatment plans; imaging examination data includes vascular structure image data obtained through imaging technology; Vital sign monitoring data include blood pressure, heart rate, and blood oxygen saturation; laboratory test result data include test results of blood and urine samples; and medication response record data include the patient's response to different medications; Initial cleaning includes using data processing tools to remove duplicates, standardize data formats, and fill in missing values; Integrate the initially collated comprehensive health data into a multi-source patient dataset.
3. A vascular surgery patient information data management platform according to claim 2, characterized in that: The method of constructing the patient data set includes: Step 31: Perform deep processing on the text, numerical, and image data in the multi-source patient dataset, including: For text data, use regular expressions to remove special symbols, punctuation marks, and non-alphanumeric characters from the text, and convert all text to lowercase. Use spelling checkers to correct spelling errors, and then apply natural language processing techniques to extract key terms and entities from the text, convert them into embedding vectors, and use the embedding vectors as text features. Among them, key terms include disease name, treatment method and drug name, and entities include patient name, hospital affiliation and diagnosis date; For numerical data, the Z-score standardization method is used to standardize the numerical data, and the standardized numerical data is used as the numerical feature; For image data, Gaussian filter is applied to remove image noise, and histogram equalization is used to enhance image contrast; Step 32: Based on all text features, numerical features, and image features, potential connections between features are obtained through association analysis, and features are further screened; Step 33: Integrate all types of filtered features and new features into a patient dataset.
4. A vascular surgery patient information data management platform according to claim 3, characterized in that: The method for calculating the degree of vascular stenosis includes: The internal adjustment function is defined by the length of the contour and curvature Obtained by weighted summation; where LK is the contour, is the initial contour, which is represented by n points ( , ), i=1,2,...,n represents the number of points on the contour; The length formula is: ;in, represents the total length of the profile LK, Refers to the loop operation, which means that when i=n, the distance between the last point and the first point is calculated, and mod represents the operator; The curvature formula is: ; in, Represents the curvature of the contour LK. For each point ( , ), i represents the position index of the center point of the current curvature calculation, and It represents the coordinate after the average smoothing of three points. Take one point before and after it and perform local smoothing to get and , , , represents the Gaussian weight function, , represents the preset standard deviation parameter, and j represents the distance of other points in the formula relative to the center point; The external adjustment function is defined as ; in, Represents the negative gradient magnitude of the image, defined as , Representing an image At the point The gradient at According to the adjustment function, the position of each point on the contour is updated using the gradient descent method. When the difference between the updated contour position and the last updated contour position is less than the preset change threshold, the contour is determined to have converged. According to the contour position after convergence, the coordinates of the contour points are extracted. For each contour point, along the normal direction Search the edge of the blood vessel and get two points PD1 and PD2, PD1=( , )+d× , PD2=( , ) d× ; Among them, d represents the preset search distance, normal Represents a straight line perpendicular to the tangent direction of the curve at the contour point; The vascular diameter is obtained by calculating the distance between PD1 and PD2, and the degree of vascular stenosis is obtained by calculating the ratio of the difference between the maximum and minimum vascular diameters in the image to the maximum vascular diameter.
5. A vascular surgery patient information data management platform according to claim 4, characterized in that: The method of using texture analysis technology to identify plaque type and size distribution and further analyzing blood vessel branching pattern and blood vessel wall thickness characteristics includes: The image data is converted into a grayscale image, and the gray-level co-occurrence matrix method is used to count the number of adjacent occurrences of pixels of different brightness in the grayscale image. Then, the texture features in the image are extracted based on the gray-level co-occurrence matrix, including the intensity of local changes, the correlation of gray levels between adjacent pixels, and the uniformity of image texture. The extracted texture features are used as input and a support vector machine is used to train a plaque classification model. Different types of plaques are distinguished and identified based on the plaque classification model, and the area and type of the plaque are output. Use image segmentation technology to separate the plaques from the image and calculate the area of each plaque. Then group the plaques according to their area and type, and count the number of plaques in each group. The area and grouping results of the plaques are used as the size and distribution characteristics of the plaques. The grayscale image is converted into a binary image using a global threshold segmentation method, and the binary image is processed using a skeletonization algorithm to extract the connectivity path of the vascular structure. For each branch in the connected path, measure the distance from one branch point to another branch point or endpoint to obtain the branch length of each branch; Among them, a branch point is a node on a connected path where paths in at least three directions intersect, and an endpoint is a point on a connected path where only one direction of the path extends; For adjacent branches in a connected path, the angle between each pair of adjacent branches is calculated as the branch angle; The branch lengths and branch angles of the blood vessels were horizontally spliced to obtain the vascular branching pattern; Based on the image data, the inner and outer boundary points of the blood vessel wall in the image are identified and extracted by edge detection technology. The polynomial surface is used as the fitting model according to the extracted boundary point data to fit the height of the blood vessel wall zb= ; Among them, u and v are the coordinates of the boundary point, zb represents the height of the blood vessel wall at the boundary point, c and b are the exponents of u and v respectively, and U and V are the highest powers of u and v respectively obtained by k-fold cross validation. are the fitting parameters to be estimated, Represents the height of the fitting surface at the boundary point. The sum of the squares of the vertical distances between the boundary point and the fitting surface is minimized by the least squares method to estimate the optimal fitting parameters. ; According to the boundary point coordinates (u, v) and the actual height zb, as well as the height of the fitting surface at the boundary point ,calculate and The absolute difference between and is used to obtain the thickness of the blood vessel wall at the boundary point (u, v); The average thickness of the vascular wall at all boundary points is taken as the vascular wall thickness.
6. A vascular surgery patient information data management platform according to claim 5, characterized in that: The method of obtaining potential connections between features through association analysis and further screening features includes: Based on all text features, numerical features and image features, a correlation matrix of all features is constructed. Each position in the matrix represents the Spearman correlation coefficient between the corresponding two features. Based on the constructed correlation matrix, the feature pairs whose absolute values of the Spearman correlation coefficient are greater than the preset correlation threshold are identified as highly correlated feature pairs; Remove one feature from each highly correlated feature pair, and calculate the multicollinearity values of all remaining features until the multicollinearity values of all features are lower than the preset collinearity threshold. Methods for calculating multicollinearity values include: Select any unselected feature as the dependent variable and all other features as independent variables. Use the multiple linear regression model to perform multiple linear regression on the dependent variable and all independent variables to obtain the linear regression determination coefficient. The reciprocal of the difference between 1 and the linear regression coefficient of determination was taken as the multicollinearity value of the dependent variable characteristics; According to the remaining features after the removal process, the multicollinearity value of the remaining features is calculated; If the multicollinearity value of a feature is greater than the collinearity threshold, it is removed and the collinearity values of all remaining features are recalculated. The calculation is iterated until the multicollinearity values of all remaining features are less than or equal to the collinearity threshold.
7. A vascular surgery patient information data management platform according to claim 6, characterized in that: The method of setting stratification variables based on the patient data set to obtain different subsets includes: Based on all the characteristics of each patient after screening, stratification variables were defined, including the patient's age, sex, and medical history; Among them, the age of the patients is divided into nl groups according to the preset age range, and the patients are divided into male and female groups according to gender. Each group is used as a stratification variable to obtain different age stratification variables and gender stratification variables; All types of diseases in the medical history of all patients were counted. Each type of disease was used as an independent medical history stratification variable. For each patient and each medical history stratification variable, if the patient had the corresponding disease, it was recorded as 1, and if the patient did not have the corresponding disease, it was recorded as 0; According to the defined stratification variables, each type of stratification variable is combined. For each combination, an independent subset is generated, and all patients are assigned to the corresponding subset. Each subset contains patients with the same stratification type and their screened characteristics. For each subset, the polynomial feature generation tool is used to transform the features of each patient in the subset, converting the single features into polynomial features and interaction features of power terms, and using the generated polynomial features and interaction features of all patients as the input of the subset.
8. A vascular surgery patient information data management platform according to claim 7, characterized in that: The method of training the postoperative complication prediction model includes: Collect the data for training and divide it into different subsets based on the stratification variables, where each subset represents a group of samples with similar characteristics. Merge all subsets into a global dataset and merge the inputs of all subsets into the input of the global dataset. The global dataset is divided into training and test sets in proportion, and the global model part of the postoperative complication prediction model is constructed using the random forest model. The input of the global dataset is used as the input of the global model to output the estimated probability of postoperative complications in patients. Initialize the model's hyperparameters and use Bayesian optimization to tune the hyperparameters. Use k-fold cross-validation to evaluate the cross-validation scores of the model under different hyperparameter combinations and select the optimal parameter combination. Define the cross entropy loss function as the loss function to evaluate the prediction accuracy of the model. In each iteration, the model is forward propagated using the training set to calculate the loss function between the predicted value and the true value, and the model parameters are updated through backpropagation. Use the R2 score as the evaluation metric and calculate the R2 score of the current iteration on the validation set. For each iteration, calculate the difference between the R2 score value after the current iteration and the R2 score value of the previous iteration, which is recorded as the iteration difference; Set the iteration difference threshold. If the iteration difference is greater than the iteration difference threshold, it is determined that the performance of the model has improved; if the iteration difference is less than or equal to the iteration difference threshold, it is determined that the performance of the model has not improved. If the performance of the model on the validation set does not improve in consecutive DD iterations, the training is stopped and the trained global model is obtained; The global dataset is redivided into different subsets. For each subset, the global model training process is repeated to obtain a local model of the postoperative complication prediction model for the corresponding subset. The output of the local model is a personalized estimate of the probability of postoperative complications for each patient in the corresponding subset.
9. A vascular surgery patient information data management platform according to claim 8, characterized in that: The method of generating the postoperative complication probability assessment report includes: Based on the comprehensive health data of the new patient, the global model of the postoperative complication prediction model is used to predict the patient's estimated probability of postoperative complications, and then the local model of the postoperative complication prediction model is used to predict the patient's personalized estimated probability of postoperative complications; The average of the prediction results of the global model and the local model is taken as the final estimate of the patient's probability of postoperative complications; Set a risk threshold. If the final estimated value is less than or equal to the preset risk threshold, the patient is judged to have a low probability of postoperative complications and a safety signal is generated. If the final estimated value is greater than the preset risk threshold, the patient is judged to have a high probability of postoperative complications and a danger signal is generated; Generate a visual assessment report on the probability of postoperative complications, including the patient's comprehensive health data, the estimated probability of postoperative complications predicted by the global model, the personalized estimated probability of postoperative complications predicted by the local model, and the final estimated probability of postoperative complications; If danger signals are detected, warnings will be issued in the postoperative complication probability assessment report to remind doctors.
10. A method for managing patient information data for vascular surgery, which is implemented based on a patient information data management platform for vascular surgery according to any one of claims 1 to 9, characterized in that: include: S1. Collect comprehensive health data of patients and conduct preliminary collation to form a multi-source patient data set; S2. Deeply process the multi-source patient dataset. Based on the deeply processed multi-source patient dataset, calculate the degree of vascular stenosis and use texture analysis technology to identify plaque type and size distribution. Further analyze the vascular branching pattern and vascular wall thickness characteristics to obtain image features. Combined with text data and numerical data for correlation analysis, a patient dataset is constructed. S3. Based on the patient data set, set stratification variables to obtain different subsets, and train postoperative complication prediction models based on the different subsets. Specifically, train the global model using all subsets, and then train the local model using a single subset. The global model outputs a universal postoperative complication probability result, and the local model outputs a personalized postoperative complication probability result. S4. Generate a postoperative complication probability assessment report based on the postoperative complication probability results to provide a basis for clinical decision-making.