Coronary heart disease risk prediction method based on multi-source data processing
By extracting features from multi-source data using a multilayer perceptron and a one-dimensional convolutional neural network, and combining time decay weights and physiological constraints, an asymmetric decision tree is constructed. This solves the problem of insufficient data utilization in traditional coronary heart disease risk assessment and improves the accuracy and reliability of predictions.
Patent Information
- Application Number
- CN202511062815.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-04
AI Technical Summary
Existing methods for assessing the risk of coronary heart disease mainly rely on limited structured clinical data, ignoring time-series data of physiological signals. Furthermore, traditional models cannot effectively utilize time attributes and complex nonlinear relationships, resulting in limited prediction accuracy and a lack of reliability.
Multilayer perceptron and one-dimensional convolutional neural network are used to extract features from multi-source data. Combined with time decay weights and physiological constraints, an asymmetric decision tree construction strategy is used to build a coronary heart disease risk prediction model, which enhances the model's utilization of time information and physiological knowledge.
It improves the accuracy and reliability of coronary heart disease risk prediction, can better explore the nonlinear correlation between multi-source features, and enhances the model's expressive power.
Smart Images

Figure CN120895246A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of disease prediction, and in particular relates to a method for predicting the risk of coronary heart disease based on multi-source data processing. Background Technology
[0002] Early and accurate prediction and screening of coronary artery disease (CAD) risk play a crucial role in disease prevention, timely intervention, and improved patient prognosis, reducing or minimizing the suffering and losses associated with the disease. Traditional CAD risk assessment primarily relies on clinical experience and classic risk scoring models. While these methods are simple and easy to implement, they are mainly based on limited structured clinical data such as age, gender, blood pressure, and blood lipids, neglecting physiological signals related to pathology, such as time-series data like electrocardiograms (ECGs). Furthermore, these models typically employ linear assumptions, but in reality, the complex nonlinear relationships between various risk factors limit predictive accuracy. In recent years, the rapid development of artificial intelligence has led to the increasing application of machine learning and deep learning models in disease prediction, achieving promising results. Models such as Support Vector Machines (SVMs), Random Forests, and Gradient Boosting Decision Trees (GBDTs) have demonstrated powerful capabilities in processing structured clinical data. However, when processing longitudinal medical data containing specified features, traditional models often ignore the temporal attribute of the data. That is, records at different time points should contribute differently to current risk prediction. Furthermore, most models operate as black boxes, simply learning statistical patterns from the data without integrating with physiological knowledge. This can lead to statistically reasonable but clinically illogical predictions, reducing the reliability of the predictions. Moreover, when constructing tree-based models like GBDT, a fixed, symmetrical tree growth strategy is typically used. This may fail to capture complex interactions between features from different sources and with different properties, limiting the model's expressive power and predictive accuracy. Summary of the Invention
[0003] To improve the accuracy of coronary heart disease prediction, this invention proposes a coronary heart disease risk prediction method based on multi-source data processing, comprising the following steps: Acquire multi-source data containing structured clinical data and physiological signal time-series data; extract the structured clinical data using a multilayer perceptron to obtain a first feature vector; process the physiological signal time-series data using a one-dimensional convolutional neural network to obtain a second feature vector; concatenate the first feature vector and the second feature vector to obtain a fused feature vector; The coronary heart disease risk prediction model is trained based on the fused feature vectors. Each iteration of the training process includes the following steps: Based on the timestamp information of the data samples, a time decay function is used to assign time decay weights to specified features in the fused feature vector, and the specified features are numerically encoded using the time decay weights. Based on the current model state, calculate the prediction uncertainty for each sample and the deviation of the sample from the preset physiological constraint equation; combine the prediction uncertainty and the deviation to perform regularization adjustment on the first and second gradients of the loss function to obtain the adjusted gradient; When constructing the decision tree for this iteration using the adjusted gradient, the interaction gain of the feature pairs is calculated first. If the maximum interaction gain exceeds a preset threshold, an asymmetric splitting strategy is used to construct the decision tree; otherwise, a symmetric splitting strategy is used to construct the decision tree. The decision trees generated in all iterations are combined to obtain the coronary heart disease risk prediction model; the trained coronary heart disease risk prediction model is used to process the fusion feature vector of the target object and output the coronary heart disease risk prediction value.
[0004] Preferably, the step of extracting the structured clinical data using a multilayer perceptron to obtain the first feature vector includes: The structured clinical data is input into a multilayer perceptron containing at least one hidden layer, the hidden layer employing a nonlinear activation function; The output layer of the multilayer perceptron outputs a vector of a preset dimension as the first feature vector.
[0005] Preferably, the step of processing the physiological signal time-series data using a one-dimensional convolutional neural network to obtain the second feature vector includes: The physiological signal time series data is input into a one-dimensional convolutional neural network containing at least one one-dimensional convolutional layer and at least one pooling layer; The last feature map is flattened and passed through at least one fully connected layer to obtain a vector of a preset dimension as the second feature vector.
[0006] Preferably, the step of assigning time decay weights to specified features in the fused feature vector using a time decay function, and then using the time decay weights to numerically encode the specified features, includes: For each data sample, the time decay weight W is calculated using an exponential decay function: ,in This is the difference between the sample timestamp and the current time. The attenuation coefficient; The weighted target sum S and the weighted sample number C for each specified feature value are calculated using the time decay weight. The numerical encoding value of the specified feature is ,in is the smoothing coefficient, and P is the global mean of the target variable.
[0007] Preferably, the step of combining the prediction uncertainty and the bias to regularize the first and second gradients of the loss function to obtain the adjusted gradient includes: By performing multiple random forward propagations on the current coronary heart disease risk prediction model, the statistical dispersion of the prediction results is calculated and used as the prediction uncertainty U; Define at least one physiological constraint equation and calculate the bias D for each sample based on the equation; Calculate the regularization term ,in and These are preset weighting coefficients; Keeping the original first-order gradient g unchanged, and adjusting the original second-order gradient h according to the regularization term R, we obtain the adjusted second-order gradient. The adjustment makes The size of is negatively correlated with the size of R.
[0008] Preferably, when constructing the decision tree for the current iteration using the adjusted gradient, the interaction gain of the feature pairs is first calculated. If the maximum interaction gain exceeds a preset threshold, an asymmetric splitting strategy is used to construct the decision tree; otherwise, a symmetric splitting strategy is used to construct the decision tree, including: Calculate the interaction gain of the preset feature pairs and find the maximum interaction gain; If the maximum interaction gain exceeds a preset threshold, then the feature pair (F1, F2) with the maximum interaction gain is selected, feature F1 is used to split the current node, and feature F2 is used preferentially or forcibly for the next split on the generated child node to achieve asymmetric splitting. Otherwise, calculate the split gain of all individual features and select the feature with the maximum split gain for symmetrical splitting.
[0009] Compared to existing technologies, this invention, during model training, introduces time decay weights to encode specified features, enabling the model to more effectively utilize the timeliness information of medical records and focusing more on the impact of recent data on current risk. Simultaneously, by combining prediction uncertainty with physiological constraint equations to adjust the model gradient, medical knowledge is integrated into model construction, constraining the model's learning direction and avoiding predictions that contradict clinical common sense, thus enhancing the reliability of the results. By selecting different decision tree construction strategies based on feature interaction gains, the complex nonlinear relationships between multi-source features can be more accurately uncovered, improving the model's expressive power. Attached Figure Description
[0010] Figure 1 A schematic diagram illustrating the process of calculating the fused feature vector; Figure 2A schematic diagram illustrating the process of calculating the second eigenvector; Figure 3 This is a schematic diagram of gradient regularization adjustment; Figure 4 Schematic diagrams of symmetrical and asymmetrical splitting. Detailed Implementation
[0011] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0012] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0013] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0014] A specific embodiment is a method for predicting the risk of coronary heart disease based on multi-source data processing, including the following steps: S1, acquire multi-source data containing structured clinical data and physiological signal time-series data; extract the structured clinical data using a multilayer perceptron to obtain a first feature vector; process the physiological signal time-series data using a one-dimensional convolutional neural network to obtain a second feature vector; concatenate the first feature vector and the second feature vector to obtain a fused feature vector; The patient's information includes structured clinical data and physiological signal time-series data. Specifically, the structured clinical data is obtained from an electronic medical record system, and includes, for example, the patient's demographic information, such as age and gender; lifestyle information, such as smoking history; and physical examination indicators, such as blood pressure and blood lipids. The physiological signal time-series data is obtained from an electrocardiogram (ECG) monitoring device, such as a 12-lead ECG signal. In one embodiment, a multilayer perceptron model is constructed, comprising an input layer, three hidden layers, and an output layer. The hidden layers use the ReLU activation function. Preprocessed structured clinical data is input into the multilayer perceptron model, and the output is the first feature vector. A one-dimensional convolutional neural network is constructed, consisting of multiple one-dimensional convolutional layers, batch normalization layers, max pooling layers, and global average pooling layers. ECG time-series data is input into the one-dimensional convolutional neural network, and the output is the second feature vector. The first feature vector and the second feature vector are concatenated dimensionally to form a longer fused feature vector, such as... Figure 1 As shown.
[0015] S2, train the coronary heart disease risk prediction model based on the fused feature vector. Each iteration of the training process includes the following steps: S21, based on the timestamp information of the data samples, a time decay function is used to assign time decay weights to the specified features in the fused feature vector, and the specified features are numerically encoded using the time decay weights. The time decay function preferably uses an exponential decay function, for example, the weight equals... Where Δt is the time difference between the historical time point and the current predicted time point, and α is the decay coefficient. In one embodiment, the specified feature is a predefined feature; in another embodiment, the specified feature is a medication record, infusion record, etc. For a specified feature such as a medication record, the weighted target statistic for each value, such as aspirin or clopidogrel, is the target value of its corresponding sample, i.e., the weighted average of whether or not one has coronary heart disease. The weight is the time decay weight of each sample. The calculated weighted target statistic value replaces the corresponding part of the fused feature vector. The category of the specified feature is encoded, for example, whether or not one smokes. For example, smoking is encoded as 0.121, and non-smoking is encoded as 0.238, etc. Then, the feature for smoking in all samples is 0.121, and the code for non-smoking is 0.238. In the numerical encoding of the specified feature using the time decay weight, the value of the specified feature in all samples and the time decay weight are used.
[0016] In an alternative embodiment, after obtaining the time decay weights, a specified feature in the original structured data is extracted separately and weighted using the time decay method to obtain a weighted encoding vector; then this encoding vector replaces the corresponding feature in the original fused feature vector.
[0017] S22, based on the current model state, calculate the prediction uncertainty of each sample and the deviation of the sample from the preset physiological constraint equation; combine the prediction uncertainty and the deviation to perform regularization adjustment on the first and second gradients of the loss function to obtain the adjusted gradient; The prediction uncertainty for each sample is estimated using the Monte Carlo dropout method, which involves using the dropout layer multiple times during model prediction and calculating the variance of the prediction results. The physiological constraint equation is, for example, that systolic blood pressure is always greater than diastolic blood pressure, and the bias is the difference between diastolic blood pressure and systolic blood pressure, taking a positive value. If the constraint is not violated, the bias is zero. The adjusted gradient is equal to the original gradient multiplied by a regularization term, preferably a regularization term of 1 plus a weighted sum, where the weighted sum is the weighted sum of prediction uncertainty and bias.
[0018] S23, when constructing the decision tree for this iteration using the adjusted gradient, first calculate the interaction gain of the feature pairs. If the maximum interaction gain exceeds a preset threshold, then an asymmetric splitting strategy is used to construct the decision tree; otherwise, a symmetric splitting strategy is used to construct the decision tree. The SHAP framework is used to calculate the interaction effect value of all feature pairs as the interaction gain. When constructing the decision tree, if the calculated maximum interaction gain is greater than a preset threshold, an asymmetric split is performed using a leaf-wise strategy, selecting the leaf with the largest gain from all current leaves for each split. If the maximum interaction gain is not greater than the threshold, a symmetric split is performed using a level-wise strategy, splitting all nodes in the same level indiscriminately.
[0019] S3, combine the decision trees generated in all iterations to obtain the coronary heart disease risk prediction model; use the trained coronary heart disease risk prediction model to process the fusion feature vector of the target object and output the coronary heart disease risk prediction value.
[0020] The predictions from each iteration of the decision tree are summed to obtain a final raw output score, which represents the logarithmic odds of the risk. The raw output score of the target object is then transformed using a logistic function, the Sigmoid function, to map it to a probability value between 0 and 1. This probability value is the predicted risk of the target object developing coronary heart disease in the future.
[0021] In an optional embodiment, the step of extracting the structured clinical data using a multilayer perceptron to obtain a first feature vector includes: The structured clinical data is input into a multilayer perceptron containing at least one hidden layer, the hidden layer employing a nonlinear activation function; The output layer of the multilayer perceptron outputs a vector of a preset dimension as the first feature vector.
[0022] Specifically, structured clinical data includes patient demographic information, laboratory test results, and vital signs. For example, a data sample can be represented as a vector containing information such as age 58 years, sex male, systolic blood pressure 140 mmHg, diastolic blood pressure 90 mmHg, and serum creatinine 1.2 mg / dL. This vector serves as the input to a multilayer perceptron. This multilayer perceptron contains two hidden layers: the first hidden layer has 128 neurons, and the second hidden layer has 64 neurons. Each hidden layer uses the modified linear unit ReLU as a nonlinear activation function to increase the model's nonlinear expressive power.
[0023] During data processing, the input patient data vector is first transformed by a weight matrix and a bias term is added. Then, it undergoes a nonlinear mapping using the ReLU activation function of the first hidden layer. The output of this mapping is then used as the input to the second hidden layer, and this process is repeated. The output layer of the multilayer perceptron is designed to output a 64-dimensional vector. This 64-dimensional vector is the first feature vector, which is a high-dimensional abstract representation learned from the original, dispersed clinical indicators, encoding the patient's static physiological state information.
[0024] In an optional embodiment, the step of processing the physiological signal time-series data using a one-dimensional convolutional neural network to obtain the second feature vector includes: The physiological signal time series data is input into a one-dimensional convolutional neural network containing at least one one-dimensional convolutional layer and at least one pooling layer; The last feature map is flattened and passed through at least one fully connected layer to obtain a vector of a preset dimension as the second feature vector.
[0025] Taking a 10-second ECG signal with a sampling frequency of 256 Hz as an example, the input data is a sequence of 2560 data points. This sequence is fed into a one-dimensional convolutional neural network. This network may contain two one-dimensional convolutional layers. The first convolutional layer uses 32 kernels of size 5 to capture basic waveforms in the ECG signal, such as the P wave or T wave; it is then followed by a max-pooling layer of size 2 to reduce the data dimensionality and retain the most salient features. The second convolutional layer uses 64 kernels of size 3 to learn more complex patterns composed of basic waveforms, such as the QRS complex. After the signal flows through all convolutional and pooling layers, the final feature map is flattened into a one-dimensional long vector. This vector is then fed into a fully connected layer containing 128 neurons for deep feature integration, generating a 64-dimensional vector through an output layer, such as... Figure 2 As shown, this 64-dimensional vector is the second feature vector, which contains the dynamic patterns and rhythmic characteristics of the electrocardiogram signal in the time dimension.
[0026] In an optional embodiment, the step of assigning time decay weights to specified features in the fused feature vector using a time decay function, and then using the time decay weights to numerically encode the specified features, includes: For each data sample, the time decay weight W is calculated using an exponential decay function: ,in This is the difference between the sample timestamp and the current time. The attenuation coefficient; The weighted target sum S and the weighted sample number C for each specified feature value are calculated using the time decay weight. The numerical encoding value of the specified feature is ,in is the smoothing coefficient, and P is the global mean of the target variable.
[0027] Suppose we need to encode a specific feature—the type of disease—in historical diagnostic records, with the target variable being the probability of a patient developing heart failure within the next year. Assuming a decay coefficient α of 0.005, for a diagnostic record from three days ago, with a Δt of 3 days, the time decay weight W is approximately 0.985. However, for a diagnostic record from six months ago (180 days ago), the weight is approximately 0.407. This indicates that recent diagnostic records contribute significantly more to current predictions than older records.
[0028] When calculating the coding value for disease types, all historical records of hypertension are considered. Assuming a global heart failure incidence rate of 0.1 (P=0.1) and a smoothing coefficient λ of 10, the target value of all hypertension records is multiplied by their corresponding time decay weights to obtain the weighted target sum S, and all weights are summed to obtain the weighted sample size C. For example, if S is calculated to be 15.2 and C to be 80.5, the final coding value for the hypertension category is approximately 0.179. This reflects both the historical association between hypertension and heart failure and considers the recentity of the data and global statistical trends, achieving reliable numerical representation of the specified feature. If the disease type is diabetes, its possible coding value is 0.251, thus different diseases have different coding values.
[0029] More specifically, a weight is calculated for each sample using a time decay function, and for a certain category, such as whether or not one smokes, two weighted objective statistics are calculated: S: Weighted target sum, the sum of the weights of all samples in this category multiplied by the target value; C: Weighted sample size, the sum of the weights of all samples in this category; Encoding after smoothing: using the formula This is used to obtain the final numerical code.
[0030] λ and P are used to prevent overfitting. When there is too little data for a certain category, the encoded value of it tends to a global average value P, making the results more stable.
[0031] In an optional embodiment, the step of combining the prediction uncertainty and the bias to regularize the first and second gradients of the loss function to obtain the adjusted gradient includes: By performing multiple random forward propagations on the current coronary heart disease risk prediction model, the statistical dispersion of the prediction results is calculated and used as the prediction uncertainty U; Define at least one physiological constraint equation and calculate the bias D for each sample based on the equation; Calculate the regularization term ,in and These are preset weighting coefficients; Keeping the original first-order gradient g unchanged, and adjusting the original second-order gradient h according to the regularization term R, we obtain the adjusted second-order gradient. The adjustment makes The size of is negatively correlated with the size of R.
[0032] Prediction uncertainty U is quantified by introducing randomness into the model's predictions. For example, using a model with random dropout to make 50 predictions for the same patient sample, if the probabilities of sepsis in the 50 predictions are 0.7, 0.2, 0.8, etc., respectively, and the standard deviation is large, for example, 0.25, it indicates that the model's prediction for that sample is very unstable, and its uncertainty U is high. Bias D is set based on medical knowledge. A physiological constraint is that mean arterial pressure (MAP) should be approximately equal to diastolic pressure (DBP) plus one-third of pulse pressure (SBP) minus DBP. If a sample's recorded values are SBP=100, DBP=90, and MAP=70, the calculated MAP should be 93.3, which deviates significantly from the recorded value. Therefore, the bias D for that sample is high, for example, 23.3.
[0033] During the gradient adjustment phase, weight coefficients are set. It is 1.0. The value is 0.5. For the sample with high uncertainty and high bias, the regularization term R = 11.9. After calculating the original second-order gradient h of the sample, it is passed through an inverse proportional function as follows: Adjustments were made, and the adjusted second gradient was obtained. It will become very small, such as Figure 3As shown, when constructing a gradient boosting decision tree, the second-order gradient determines the weight of the sample in the split gain calculation. Therefore, this sample contributes almost nothing to the structure of the decision tree, effectively preventing the model from being misled by these unreliable or illogical data points.
[0034] In an optional embodiment, when constructing the decision tree for the current iteration using the adjusted gradient, the interaction gain of the feature pairs is first calculated. If the maximum interaction gain exceeds a preset threshold, an asymmetric splitting strategy is used to construct the decision tree; otherwise, a symmetric splitting strategy is used to construct the decision tree, including: Calculate the interaction gain of the preset feature pairs and find the maximum interaction gain; If the maximum interaction gain exceeds a preset threshold, then the feature pair (F1, F2) with the maximum interaction gain is selected, feature F1 is used to split the current node, and feature F2 is used preferentially or forcibly for the next split on the generated child node to achieve asymmetric splitting. Otherwise, calculate the split gain of all individual features and select the feature with the maximum split gain for symmetrical splitting.
[0035] When constructing a node in a decision tree, the model evaluates the interaction effects of feature pairs. For example, it evaluates the interaction gain of the feature pair age and lactate level. It calculates the gain from splitting by age alone, and the gain from splitting by lactate level alone. Then, the model calculates the total gain from splitting by age first, and then splitting by lactate level on its child nodes. The interaction gain is this total gain minus the sum of the individual split gains of the two features. Assuming the interaction gain between age and lactate level is calculated to be 0.8, and the interaction gain between age and white blood cell count is 0.3, then the maximum interaction gain is 0.8.
[0036] If the preset interaction gain threshold is 0.5, and the maximum interaction gain of 0.8 exceeds this threshold, the model will perform an asymmetric split. It will select age as the splitting feature for the current node, for example, dividing it into two branches: one for ages less than 65 and the other for ages greater than or equal to 65. On these two child nodes, when searching for the next optimal splitting feature, the model will force or prioritize lactate level for splitting. This method can explicitly reveal the unique risk pattern of elevated lactate levels in specific patient groups, such as elderly patients. Conversely, if the maximum interaction gain of 0.3 does not exceed the threshold, the model will revert to the standard procedure, i.e., calculating the splitting gain of all individual features such as age, lactate level, and white blood cell count separately, and selecting the feature with the highest gain for a standard symmetric split, such as... Figure 4 As shown.
[0037] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, compact disc read-only memory (CD-ROM), optical storage, etc.) containing computer-usable program code.
[0038] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0039] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for predicting the risk of coronary heart disease based on multi-source data processing, characterized in that, Includes the following steps: Acquire multi-source data containing structured clinical data and physiological signal time-series data; extract the structured clinical data using a multilayer perceptron to obtain a first feature vector; process the physiological signal time-series data using a one-dimensional convolutional neural network to obtain a second feature vector; concatenate the first feature vector and the second feature vector to obtain a fused feature vector; The coronary heart disease risk prediction model is trained based on the fused feature vectors. Each iteration of the training process includes the following steps: Based on the timestamp information of the data samples, a time decay function is used to assign time decay weights to specified features in the fused feature vector, and the specified features are numerically encoded using the time decay weights. Based on the current model state, calculate the prediction uncertainty for each sample and the deviation of the sample from the preset physiological constraint equation; combine the prediction uncertainty and the deviation to perform regularization adjustment on the first and second gradients of the loss function to obtain the adjusted gradient; When constructing the decision tree for this iteration using the adjusted gradient, the interaction gain of the feature pairs is calculated first. If the maximum interaction gain exceeds a preset threshold, an asymmetric splitting strategy is used to construct the decision tree. Otherwise, a symmetric splitting strategy is used to construct the decision tree; The decision trees generated in all iterations are combined to obtain the coronary heart disease risk prediction model; the trained coronary heart disease risk prediction model is used to process the fusion feature vector of the target object and output the coronary heart disease risk prediction value.
2. The method according to claim 1, characterized in that, The step of extracting the first feature vector from the structured clinical data using a multilayer perceptron includes: The structured clinical data is input into a multilayer perceptron containing at least one hidden layer, the hidden layer employing a nonlinear activation function; The output layer of the multilayer perceptron outputs a vector of a preset dimension as the first feature vector.
3. The method according to claim 1, characterized in that, The process of using a one-dimensional convolutional neural network to process the time-series physiological signal data to obtain the second feature vector includes: The physiological signal time series data is input into a one-dimensional convolutional neural network containing at least one one-dimensional convolutional layer and at least one pooling layer; The last feature map is flattened and passed through at least one fully connected layer to obtain a vector of a preset dimension as the second feature vector.
4. The method according to claim 1, characterized in that, The step of assigning time decay weights to specified features in the fused feature vector using a time decay function, and then using the time decay weights to numerically encode the specified features, includes: For each data sample, the time decay weight W is calculated using an exponential decay function: ,in This is the difference between the sample timestamp and the current time. The attenuation coefficient; The weighted target sum S and the weighted sample number C for each specified feature value are calculated using the time decay weight. The numerical encoding value of the specified feature is ,in is the smoothing coefficient, and P is the global mean of the target variable.
5. The method according to claim 1, characterized in that, The first and second gradients of the loss function are regularized by combining the prediction uncertainty and the bias to obtain the adjusted gradients, including: By performing multiple random forward propagations on the current coronary heart disease risk prediction model, the statistical dispersion of the prediction results is calculated and used as the prediction uncertainty U; Define at least one physiological constraint equation and calculate the bias D for each sample based on the equation; Calculate the regularization term ,in and These are preset weighting coefficients; Keeping the original first-order gradient g unchanged, and adjusting the original second-order gradient h according to the regularization term R, we obtain the adjusted second-order gradient. The adjustment makes The size of is negatively correlated with the size of R.
6. The method according to claim 1, characterized in that, When constructing the decision tree for the current iteration using the adjusted gradient, the interaction gain of the feature pairs is calculated first. If the maximum interaction gain exceeds a preset threshold, an asymmetric splitting strategy is used to construct the decision tree. Otherwise, a symmetric splitting strategy is used to construct the decision tree, including: Calculate the interaction gain of the preset feature pairs and find the maximum interaction gain; If the maximum interaction gain exceeds a preset threshold, then the feature pair (F1, F2) with the maximum interaction gain is selected, feature F1 is used to split the current node, and feature F2 is used preferentially or forcibly for the next split on the generated child node to achieve asymmetric splitting. Otherwise, calculate the split gain of all individual features and select the feature with the maximum split gain for symmetrical splitting.
Citation Information
Cited By
Chemical process safety risk early warning system based on edge calculation
CN121599478A
A chemical process safety risk early warning system based on edge computing
CN121599478B