Medical index model construction method and device and electronic equipment
By constructing a medical index model, using loss function and performance index values to screen the training feature subset, and combining normal distribution transformation and sample equality analysis, the poor model accuracy problem caused by medical data imbalance is solved, and the prediction ability and stability of the model are improved.
Patent Information
- Application Number
- CN202410018308.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, imbalance in the distribution of medical data leads to poor prediction accuracy of model prediction, especially in the prediction of a few samples.
By obtaining the original feature data set, using the loss function to calculate the model loss value and performance index value, select patient data samples with performance index values greater than the set threshold to form a training feature subset, and train the specified medical index model, combining normal distribution transformation, sample equality analysis and synthetic sample generation to improve the generalization ability and prediction accuracy of the model.
It effectively solves the problem of medical data imbalance, improves the prediction accuracy and generalization ability of the model, especially the prediction ability of a few types of samples, and reduces the risk of overfitting.
Smart Images

Figure CN120280133A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of medical technologies, and in particular, to a method, an apparatus, and an electronic device for constructing a medical index model. Background Art
[0002] In terms of disease prediction, doctors usually predict the development trend of patients' conditions in advance based on their manual experience, so as to formulate more targeted treatment plans. For example, for patients with cardiovascular diseases, doctors can predict the risks of adverse events such as myocardial infarction and stroke in the future of the patients, and thus take intervention measures in advance to reduce the mortality rate of the patients.
[0003] Therefore, in the prior art, there have been solutions to predict diseases based on medical data by introducing models. However, the medical data often has the problem of unbalanced data distribution. For example, when predicting whether a patient will have breast cancer, the number of patients with tumors is often much less than the number of non-patient populations. As a result, the prediction results of the model are poor in accuracy. Summary of the Invention
[0004] The purpose of the present application is to provide a method, an apparatus, and an electronic device for constructing a medical index model to solve or overcome the above-mentioned technical problems existing in the prior art.
[0005] According to the first aspect of the embodiments of the present application, a method for constructing a medical index model is provided, which includes:
[0006] Obtain an original feature data set, where the original feature data set includes a plurality of patient data samples, and each patient data sample includes a plurality of patient features corresponding to the same patient;
[0007] Input the patient data samples in the original feature set into the constructed first medical index model in sequence for identification to predict the medical index of the patient;
[0008] Based on the constructed loss function, calculate the loss value of the first medical index model according to the predicted medical index of the patient;
[0009] Calculate the performance index value of the first medical index model according to the loss value;
[0010] Select some patient data samples corresponding to when the performance index value is greater than the set index threshold and add them to the pre-constructed feature subset to obtain a training feature subset;
[0011] Train the specified second medical index model based on the training feature subset.
[0012] Optionally, it further includes:
[0013] Perform a normal distribution transformation on the original data feature set to obtain data points;
[0014] Based on the data points, draw a quantity - quantile plot;
[0015] Judge whether the data points satisfy the normal distribution according to the quantity - quantile plot;
[0016] If it is satisfied, perform the step of sequentially inputting the patient data samples in the original feature set into the constructed first medical index model for recognition to predict the medical index of the patient.
[0017] Optionally, it further includes: performing sample balance analysis on the training feature subset to obtain a balance estimate. If the balance estimate is less than the set balance threshold, determine the minority sample region in the training feature subset, and generate synthetic samples based on the minority samples in the minority sample region to fill into the training feature subset.
[0018] Optionally, the generating synthetic samples based on the minority samples in the minority sample region to fill into the training feature subset includes:
[0019] For any minority sample in the minority sample region, determine at least two adjacent minority samples;
[0020] Perform a difference operation between the at least two samples to generate the synthetic samples and fill them into the minority sample region.
[0021] Optionally, the method further includes:
[0022] Divide the training feature subset into several folds;
[0023] Take any selected fold as the test set, and the remaining other folds as the training set, to train the disease recognition model based on the training set, and test the second medical index model based on the test set to judge whether the second medical index model trained based on the training set meets the set model performance indicators.
[0024] Optionally, the judging whether the second medical index model trained based on the training set meets the set model performance indicators includes:
[0025] Calculate the model performance indicators of the second medical index model when each fold is used as the test set and the other folds are used as the training set;
[0026] Based on the mean value of all model performance indicators, evaluate whether the second medical index model trained based on the training set meets the set model performance indicators.
[0027] Optionally, the constructed loss function calculates the loss value of the first medical index model according to the predicted medical indexes of the patient, including:
[0028] Based on the constructed loss function, calculate the standard loss value of the first medical index model according to the predicted medical indexes of the patient;
[0029] Calculate the regularization loss value of the first medical index model according to the parameters of the medical indexes when predicting the medical indexes of the patient and the set regularization hyperparameters;
[0030] Calculate the regression loss value of the first medical index model according to the standard loss value and the regularization loss value.
[0031] Optionally, the method further includes:
[0032] Calculate the data quality of the model training subset;
[0033] Based on the data quality, match with the medical index model library to match out a second medical index model from multiple second medical index models in the medical index model library as the specified second medical index model.
[0034] A device for constructing a medical index model, which includes:
[0035] A data acquisition unit for acquiring an original feature data set, where the original feature data set includes multiple patient data samples, and each patient data sample includes multiple patient features corresponding to the same patient;
[0036] A pre-training unit for performing the following steps:
[0037] Input the patient data samples in the original feature set into the constructed first medical index model in sequence for identification to predict the medical indexes of the patient;
[0038] Based on the constructed loss function, calculate the loss value of the first medical index model according to the predicted medical indexes of the patient;
[0039] Calculate the performance index value of the first medical index model according to the loss value;
[0040] A feature subset construction unit for selecting some patient data samples corresponding to when the performance index value is greater than the set index threshold and adding them to the pre-constructed feature subset to obtain a training feature subset;
[0041] A model training unit for training the specified second medical index model based on the training feature subset.
[0042] An electronic device includes a memory and a processor. A computer-executable program is stored on the memory, and the processor is configured to run the computer-executable program to execute the method described in any one of the present application.
[0043] In the solution provided by the embodiments of the present application, by obtaining an original feature dataset, the original feature dataset includes multiple patient data samples, and each patient data sample includes multiple patient features corresponding to the same patient; sequentially inputting the patient data samples in the original feature set into a constructed first medical index model for recognition to predict the medical index of a patient; based on a constructed loss function, calculating the loss value of the first medical index model according to the predicted medical index of the patient; calculating the performance index value of the first medical index model according to the loss value; selecting some of the patient data samples corresponding to when the performance index value is greater than a set index threshold and adding them to a pre-constructed feature subset to obtain a training feature subset; training a specified second medical index model based on the training feature subset, thereby avoiding poor model prediction accuracy caused by unbalanced medical data. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Some specific embodiments of the embodiments of the present application will be described in detail hereinafter with reference to the drawings in an exemplary but not restrictive manner. The same reference numerals in the drawings denote the same or similar components or parts. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:
[0045] Figure 1 It is a schematic flowchart of a method for constructing a medical index model according to an embodiment of the present application.
[0046] Figure 2 It is a schematic structural diagram of a device for constructing a medical index model according to an embodiment of the present application.
[0047] Figure 3 It is a schematic flowchart of a method for preprocessing patient data according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application shall fall within the protection scope of the embodiments of the present application.
[0049] Figure 1 It is a schematic flowchart of a method for constructing a medical index model according to an embodiment of the present application. As Figure 1 shown, it includes:
[0050] S101. Obtain an original feature dataset, where the original feature dataset includes multiple patient data samples, and each patient data sample includes multiple patient features corresponding to the same patient;
[0051] S102. Input the patient data samples in the original feature set into the constructed first medical index model in sequence for recognition to predict the medical index of the patient;
[0052] S103. Based on the constructed loss function, calculate the loss value of the first medical index model according to the predicted medical index of the patient;
[0053] S104. Calculate the performance index value of the first medical index model according to the loss value;
[0054] S105. Select some patient data samples corresponding to when the performance index value is greater than the set index threshold and add them to the pre-constructed feature subset to obtain a training feature subset;
[0055] S106. Train the specified second medical index model based on the training feature subset.
[0056] By inputting the patient data samples into the first medical index model in sequence for recognition and prediction, combining the calculation of the loss function and the performance index value, selecting some patient data samples corresponding to when the performance index value is greater than the set index threshold and adding them to the feature subset, so as to eliminate the features with less influence on the model prediction, thereby improving the training effect and generalization ability of the model and ensuring the prediction ability and accuracy of the second medical index model.
[0057] Optionally, the method further includes:
[0058] Initialize an empty feature subset: p selected = 0, where 0 indicates that the number of patient features in the training feature subset at initialization is 0;
[0059] Correspondingly, inputting the patient data samples in the original feature set into the constructed first medical index model in sequence for recognition to predict the medical index of the patient includes: for the j-th patient feature x in the patient data sample jInput it into the constructed first medical index model for identification to predict the medical index of the patient. j is a positive integer and does not exceed the total number of patient characteristics. After that, perform the above steps S103 and S104 again to obtain the performance index value of the first medical index model. Traverse all patient characteristics in all patient data samples corresponding to all patients to obtain multiple performance index values of the first medical index model. Then, after performing step S105, select the partial patient characteristics (such as denoted as x best ) included in the partial patient data samples corresponding to the performance index values greater than the set index threshold and add them to the pre-constructed feature subset to obtain the training feature subset p selected .
[0060] Alternatively, in other embodiments, step S105 may be replaced by: determining whether the improvement of the performance index value of the first medical index model is not significant, or the number of patient characteristics added to the feature subset reaches the set number, and then ending the above steps S102 - S105.
[0061] Optionally, the method further includes:
[0062] Perform normal distribution transformation on the original data feature set to obtain data points;
[0063] Based on the data node positions, draw a quantity - quantile plot;
[0064] Judge whether the data points satisfy the normal distribution according to the quantity - quantile plot;
[0065] If it is satisfied, perform the step of inputting the patient data samples in the original feature set into the constructed first medical index model for identification to predict the medical index of the patient.
[0066] Optionally, performing normal distribution transformation on the original data feature set to obtain data points includes: performing normal distribution transformation on the same - type patient characteristics of different patients to obtain data points. The normal distribution transformation is, for example, by performing logarithmic transformation, exponential transformation, square - root transformation, etc. on the patient characteristics to obtain the corresponding data points.
[0067] In this embodiment, through normal distribution transformation, the original data feature set can be converted into data points that satisfy the normal distribution, thereby eliminating skewness and heteroscedasticity in the data and making the data more in line with statistical assumptions. In addition, by drawing a quantity - quantile plot and judging whether the data points satisfy the normal distribution, an objective evaluation of the data distribution can be performed, thereby discovering the distribution characteristics in the data and providing a reference for the specification of the subsequent second medical index model.
[0068] Optionally, it further includes: performing sample balance analysis on the training feature subset to obtain a balance estimate. If the balance estimate is less than a set balance threshold, determining the minority sample region in the training feature subset, and generating synthetic samples based on the minority samples in the minority sample region to fill into the training feature subset.
[0069] In this embodiment, by performing sample balance analysis on the training feature subset and synthesizing minority samples when the balance estimate is less than the set threshold, the problem of class imbalance in the data set is effectively improved, thereby improving the model's learning ability for minority class samples, reducing the overfitting risk of the model for majority class samples, and thus improving the generalization ability and prediction performance of the model. In addition, by generating synthetic samples and filling them into the training feature subset, the number of minority class samples can be increased, the diversity of the data can be enriched, enabling the model to learn and understand the features of the data more comprehensively, improving the robustness and generalization ability of the model, and enhancing the prediction accuracy of the model for new samples. Finally, by filling synthetic samples into the training feature subset, the overfitting risk caused by too few minority class samples is reduced, the prediction ability of the model for minority class samples is improved, and the model can more comprehensively cover all aspects of the data.
[0070] Optionally, the generating synthetic samples based on the minority samples in the minority sample region to fill into the training feature subset includes:
[0071] For any minority sample in the minority sample region, determining at least two adjacent minority samples thereto;
[0072] Performing a difference operation between the at least two samples to generate the synthetic sample and filling it into the minority sample region.
[0073] In this embodiment, by determining at least two adjacent minority samples in the minority sample region and performing a difference operation between them to generate synthetic samples, the characteristic information of the original data can be retained, ensuring that the synthetic samples better represent the distribution characteristics of the original data and improving the model's understanding and learning ability of the data. In addition, by performing a difference operation between adjacent minority samples to generate synthetic samples, the diversity of the synthetic samples can be controlled, ensuring that the synthetic samples maintain a certain similarity with the original data, avoiding the generation of overly discrete or abnormal synthetic samples, and improving the effectiveness and credibility of the synthetic samples. Finally, by generating synthetic samples and filling them into the minority sample region, the sample size of the data set can be increased, and the diversity of the data can be enriched, improving the generalization ability of the model, enhancing the prediction accuracy of the model for minority class samples, and reducing the overfitting risk.
[0074] Optionally, the method further includes:
[0075] Dividing the training feature subset into several folds;
[0076] Take any one of the selected folds as the test set, and the remaining other folds as the training set. Train the disease recognition model based on the training set, and test the second medical index model based on the test set to determine whether the second medical index model trained based on the training set meets the set model performance indicators.
[0077] In this embodiment, by dividing the training feature subset into multiple folds and alternately selecting one fold as the test set, the disease recognition model can be trained and tested multiple times, so as to more comprehensively evaluate the performance of the model, improve the generalization ability and stability of the model, and improve the prediction accuracy of the model. In addition, the dependence of the model on a specific training set is reduced, and the risk of overfitting is reduced. Moreover, through multiple trainings and tests, the performance of the model on different data sets can be better grasped, and the robustness and generalization ability of the model can be improved. During the training process of each cross-validation, the parameters of the model can be adjusted according to the performance of the validation set, so as to optimize the performance of the model, select the best model parameter configuration, improve the prediction ability of the model, effectively avoid overfitting at the same time, and ensure the generalization ability of the model.
[0078] Optionally, the determination of whether the second medical index model trained based on the training set meets the set model performance indicators includes:
[0079] Calculate the model performance indicators of the second medical index model when each fold is used as the test set and the other folds are used as the training set;
[0080] Based on the mean of all model performance indicators, evaluate whether the second medical index model trained based on the training set meets the set model performance indicators.
[0081] For example, based on the following formula, the performance evaluation value of the second medical index model trained based on the training set can be obtained based on the mean of all model performance indicators, and then compared with the set model performance indicators to determine whether it is not less than the set model performance indicators:
[0082]
[0083] D represents the training feature subset, which is divided into k folds, D i represents the i-th fold, CV k represents the performance evaluation value.
[0084] In this embodiment, by testing each fold and calculating the model performance metrics, and then taking the average of the performance metrics of all folds, the performance of the model on different data subsets can be comprehensively evaluated, so as to more comprehensively understand the generalization ability and stability of the model, and improve the accuracy of evaluating the model performance. In addition, since the division of each cross-validation is random, by calculating the average performance metrics of multiple cross-validations, the influence brought by randomness can be reduced, and the performance of the model can be evaluated more accurately. Finally, evaluating the model performance based on the average of all model performance metrics can reduce the dependence on a specific data set, improve the robustness and generalization ability of the model, and ensure that the model can achieve good performance under different data distributions.
[0085] Optionally, based on the constructed loss function, according to the predicted medical metrics of the patient, calculating the loss value of the first medical metric model includes:
[0086] Based on the constructed loss function, according to the predicted medical metrics of the patient, calculating the standard loss value of the first medical metric model;
[0087] Predicting the parameters of the medical metrics when predicting the medical metrics of the patient and the set regularization hyperparameters, and calculating the regularization loss value of the first medical metric model;
[0088] According to the standard loss value and the regularization loss value, calculating the regression loss value of the first medical metric model.
[0089] For example, the regression loss value of the first medical metric model can be calculated based on the following formula:
[0090]
[0091] where y i is the target value of the medical metrics of the i-th patient, x ij is the j-th patient feature of the i-th patient, β j is the model parameter of the first medical metric model corresponding to the j-th patient feature of all patients, n is the total number of patients, p is the total number of patient features in each patient's data, and λ is the regularization hyperparameter.
[0092] represents the calculated standard loss value, represents the calculated regression loss value.
[0093] In this embodiment, by calculating the standard loss value and the regularization loss value, the model parameters of the first medical index model are adjusted, for example, making their values tend to zero, so as to eliminate the patient features that have little influence on the training of the second medical index model, and retain the patient features with small values but actually having a greater impact, so as to reduce the dimension of the data and improve the performance of the training of the second medical index model.
[0094] Optionally, the method further includes:
[0095] Calculating the data quality of the model training subset;
[0096] Based on the data quality, it is matched with the medical index model library to match a second medical index model from multiple second medical index models in the medical index model library as the specified second medical index model.
[0097] In this embodiment, multiple second medical index models in the medical index model library include, for example:
[0098] 1. Support Vector Machine (SVM): Suitable for efficient data classification, especially excellent in complex problems and non-linear data.
[0099] 2. K-Nearest Neighbors: Effective for data with uneven distribution or local patterns.
[0100] 3. Neural Networks: Include Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), and Long Short-Term Memory Networks (LSTM) to handle sequence data or time series prediction tasks.
[0101] 4. Decision Trees and Their Ensemble Methods: Decision trees, Random Forests, AdaBoost, Gradient Boosting, and Bagging and other ensemble methods to improve prediction performance.
[0102] 5. Bayesian Classifiers: Include variants such as Naive Bayes and Gaussian Naive Bayes, suitable for various data types.
[0103] 6. Deep Learning Models: In addition to traditional convolutional neural networks, complex deep learning models such as Generative Adversarial Networks (GAN) and Transformers are also considered to adapt to more complex data features.
[0104] In a specific application scenario, taking XGBoost (eXtreme Gradient Boosting) in Gradient Boosting as an example, the training dataset T = {(x1, y1), (x2, y2), …, (x n , y n )}, where x1 …… x n represent the first to the nth patients, and y1 …… y n represent the medical indicators predicted corresponding to these n patients respectively. The loss function is used to calculate the loss value between the predicted medical indicator and the corresponding actual medical indicator y i , where the value of i ranges from 1 to n; the regularization term is denoted as Ω(f k ) to calculate the regularization loss value, and the overall objective function is:
[0105]
[0106] where i is the i-th patient, k is the k-th tree in XGBoost,
[0107] By performing a second-order Taylor expansion on the above overall objective function, expanding the regularization term, removing the constant term, and combining the coefficients of the linear term and the quadratic term, the final objective function is obtained as follows:
[0108]
[0109] where, G is the sum of the cumulative first-order partial derivatives of the fitting relationship of the characteristics of the same type of patients, and H is the sum of the cumulative second-order partial derivatives of the fitting relationship of the characteristics of the same type of patients.
[0110] where j represents the leaf node, T represents the total number of leaf nodes, and both γ and λ are regularization hyperparameters. For the j-th leaf node:
[0111] w j is the weight of this leaf node.
[0112] G j is the cumulative sum of the first-order derivatives of this leaf node, corresponding to the first-order derivative of the loss function.
[0113] H j is the cumulative sum of the second-order derivatives of this leaf node, corresponding to the second-order derivative of the loss function.
[0114] λ is the coefficient of the regularization term, used to control the complexity of the model.
[0115] T is the number of leaf nodes of the tree.
[0116] γ is the coefficient of the regularization term, which is used to further control the complexity of the tree.
[0117] The objective function consists of three parts:
[0118] Loss term: This part represents the degree of fit of the model to the training data. The first term G j w j corresponds to the accumulation of the first-order derivatives, and the second term corresponds to the accumulation of the second-order derivatives and the regularization term.
[0119] Regularization term: γ T . This part corresponds to the complexity of the tree, where γ is the coefficient of the regularization term and T is the number of leaf nodes of the tree. This helps prevent the model from overfitting.
[0120] Finally, the weight of each leaf node is obtained, and the optimal Obj objective value at this time is: By minimizing the above objective function, the optimal leaf node scores are found, and the optimal leaf node weights w j * .
[0121] In this embodiment, by calculating the data quality of the model training subset, the integrity, accuracy, and reliability of the training data can be evaluated, so as to select a matching medical index model, ensuring that the selected model can effectively utilize the training data and achieve good performance in actual applications. Furthermore, by matching the data quality with the medical index model library, the most suitable second medical index model can be selected according to the actual data situation, thereby realizing personalized model selection and improving the fit between the model and the actual data. In addition, selecting a matching second medical index model can improve the adaptability and prediction accuracy of the model, enabling the model to better adapt to specific training data, and thus improving the effect of the model in actual applications.
[0122] In the above embodiment, whether for the first medical index model or the second medical index model, for example, when we evaluate the performance of a classification model, the following performance metric values can be used: Precision (precision rate), Recall (recall rate), F1 Score, Accuracy (accuracy rate), and AUC-ROC.
[0123] The precision rate refers to the proportion of patient data samples actually belonging to the positive class among those predicted as the positive class by the model, that is, the accuracy of the model when predicting the positive class.
[0124]
[0125] Recall rate refers to the proportion of patient data samples that are actually positive classes and are correctly predicted as positive classes by the model. That is, the model's recognition ability for positive class samples.
[0126]
[0127] F1 Score is the harmonic mean of precision and recall rate.
[0128]
[0129] Accuracy rate refers to the proportion of the number of patient data samples correctly classified by the model to the total number of patient data samples.
[0130]
[0131] The ROC curve is a curve plotted with the false positive rate at different thresholds as the abscissa and the true positive rate (i.e., recall rate) as the ordinate.
[0132] Figure 2 It is a schematic structural diagram of a construction device for a medical index model in an embodiment of the present application. As Figure 2 shown, it includes:
[0133] A data acquisition unit 201, configured to acquire an original feature data set, where the original feature data set includes a plurality of patient data samples, and each patient data sample includes a plurality of patient features corresponding to the same patient;
[0134] A pre-training unit 202, configured to perform the following steps:
[0135] Sequentially input the patient data samples in the original feature set into the constructed first medical index model for recognition to predict the medical index of the patient;
[0136] Based on the constructed loss function, calculate the loss value of the first medical index model according to the predicted medical index of the patient;
[0137] Calculate the performance index value of the first medical index model according to the loss value;
[0138] A feature subset construction unit 203, configured to select some patient data samples corresponding to when the performance index value is greater than a set index threshold and add them to a pre-constructed feature subset to obtain a training feature subset;
[0139] A model training unit 204, configured to train a specified second medical index model based on the training feature subset.
[0140] The original feature dataset used in the above embodiments of the present application can be obtained by using the patient data preprocessing method provided in the following embodiments.
[0141] Figure 3 It is a schematic flowchart of a patient data preprocessing method according to an embodiment of the present application. As Figure 3 shown, it includes:
[0142] S301. Obtain the patient data collected by the patient data collection module;
[0143] S302. Classify the patient data to obtain first-type data and second-type data. The first-type data is collected by wearable sensors configured on the patient, and the second-type data is collected by an application;
[0144] S303. Perform recursive estimation processing on the first-type data based on a set first preprocessing model to generate corresponding predicted state data;
[0145] S304. Fill in the missing values of the second-type data based on a set second preprocessing model to generate corresponding estimated missing values.
[0146] Take the patient data that has completed the preprocessing process shown in steps S301 - S304 as the patient data sample in the above Figure 1 - Figure 2 embodiment.
[0147] In the above embodiment, after obtaining the patient data collected by the patient data collection module, the patient data is classified to obtain first-type data and second-type data, thereby distinguishing which are collected by wearable sensors and which are collected by the application, realizing the preliminary classification of patient data. Further, targeted processing is performed on these two types of data based on different preprocessing models respectively, realizing the in-depth preprocessing of data, improving the quality of patient data, and thus ensuring that the subsequent treatment plan generated has high accuracy and can achieve targeted treatment of patients.
[0148] Optionally, in step S303, based on the set first preprocessing model, the following steps are performed to perform recursive estimation processing on the first-type data to generate corresponding predicted state data:
[0149] S313. Obtain the historical estimated value of the first-type data;
[0150] S323. Based on the historical estimated value, perform recursive estimation processing on the first-type data to generate corresponding predicted state data.
[0151] Optionally, in step S323, based on the current measurement value and the historical estimated value, performing a recursive estimation process on the first type of data to generate corresponding predicted state data, including:
[0152] S3231. Determining a current estimated value based on a set state transition matrix and the historical estimated value;
[0153] S3232. Performing a recursive estimation process on the first type of data based on the current estimated value and the Kalman gain calculated according to the state transition matrix to generate corresponding predicted state data.
[0154] Optionally, the method further includes:
[0155] Determining a predicted error covariance at the current moment based on the state transition matrix and the historical predicted error covariance;
[0156] Calculating the Kalman gain according to the predicted error covariance at the current moment and the state transition matrix.
[0157] In a specific application scenario, the recursive estimation process on the first type of data based on a set first preprocessing model to generate corresponding predicted state data can be implemented according to the following formula (1):
[0158]
[0159] Wherein, is the preliminary predicted state data at time k, i.e., the current estimated value, A is the state transition matrix, is the estimated state at time k - 1, i.e., the historical estimated value, and k is a positive integer.
[0160] Thus, through the above formula (1), determining the current estimated value based on the set state transition matrix and the historical estimated value is achieved.
[0161]
[0162] Wherein, is the predicted error covariance at time k, i.e., the predicted error covariance at the current moment, P k-1 is the error covariance at time k - 1, i.e., the historical predicted error covariance, and Q is the process noise covariance matrix. That is, through the above formula (2), determining the predicted error covariance at the current moment based on the state transition matrix and the historical predicted error covariance is achieved.
[0163]
[0164] K kis the Kalman gain at time k, H is the observation matrix, and R is the measurement noise covariance matrix.
[0165] That is, through the above formula (3), the Kalman gain is calculated based on the predicted error covariance at the current time and the state transition matrix.
[0166]
[0167] is the final state estimate at time k, that is, the predicted state data corresponding to the first type of data at time k, z k is the current measurement value of the first type of data.
[0168] On the basis of the above processing, in order to facilitate the calculation of the predicted state data corresponding to the first type of data at time k + 1, it further includes: updating the predicted error covariance at the current time according to the predicted error covariance at the current time and the observation matrix. For example, it is implemented with reference to the following formula (5).
[0169]
[0170] P k is the final error covariance at time k, and I is the identity matrix.
[0171] In the above embodiments, each matrix can be determined according to the actual application scenario. The process noise covariance Q, the measurement noise covariance matrix R, and the predicted error covariance matrix P are set parameters. First, an initial value is given by experience, and then it is continuously optimized according to the results. The observation matrix H is directly obtained by observation. The Kalman gain K and the state transition matrix A can both be obtained through formula derivation.
[0172] In this embodiment, based on the above first preprocessing model, it is possible to accurately predict the future state of the data on the basis of considering the current measurement value and the historical estimated value, so it can be applied to the non-steady first type of data and can also adapt to the dynamically changing first type of data, thereby providing a more accurate state estimate and providing more stable data for subsequent analysis.
[0173] The wearable sensors that generate the above first type of data include but are not limited to temperature sensors, humidity sensors, etc.
[0174] On the basis of the above embodiments, optionally, the method further includes:
[0175] Dividing the predicted state data into several data samples according to a set time period;
[0176] Performing short-time Fourier transform on each data sample to generate a short-time spectrum;
[0177] Generated according to the short video spectra corresponding to all samples.
[0178] Optionally, in this embodiment, for example, the predicted state data can be divided according to a set time window function so as to divide the predicted state data according to a set time period, and then a number of data samples are obtained.
[0179] Optionally, according to the short video spectra corresponding to all samples, a micro-Doppler time-frequency map is generated. For example, the short-time spectra within all time window functions can be stacked to obtain a micro-Doppler time-frequency map, so as to visually observe the changes of the predicted state data in terms of time and frequency, thereby providing high-quality basic data for subsequent training and testing.
[0180] In addition, since the micro-Doppler time-frequency map is formed by the above stacking, the matrixization of the predicted state data is realized, and the predicted state data is embodied in the form of a matrix, thereby providing high-quality basic data for subsequent training and testing.
[0181] Optionally, in step S304, based on the set second preprocessing model, missing value filling is performed on the second type of data to generate corresponding estimated missing values, including:
[0182] S314. Assign independent categories to the second type of data with missing values;
[0183] S324. Based on the set second preprocessing model, perform missing value prediction on the category to generate corresponding estimated missing values.
[0184] The above second type of data is, for example, the basic attribute data of a patient, which may include gender, region, education level, age, etc.
[0185] In this embodiment, by assigning the above opposing categories, the information of the missing values in the second type of data can be retained, and no additional bias or distortion is introduced. Further, excessive processing of the data can be avoided, thereby maintaining the integrity and consistency of the data.
[0186] Optionally, the method further includes: determining whether the second type of data with missing values is a categorical variable or a continuous variable. For example, the gender is a categorical variable, while the age is a continuous variable.
[0187] If the second type of data with missing values is a categorical variable, then based on the set second preprocessing model, approximate sample prediction is performed on the category to generate corresponding estimated missing values;
[0188] If the second type of data with missing values is a continuous variable, then based on the set second preprocessing model, approximate sample prediction or a statistic filling algorithm is performed on the category to generate corresponding estimated missing values.
[0189] Optionally, the approximate sample prediction for the category based on the set second preprocessing model in the above steps can be implemented based on the following steps:
[0190]
[0191] Where: is the predicted category of the missing value x, c j is a category in the set of optional categories, K is the number of nearest neighbors selected, y i is the category of the i-th training sample closest to the missing value x, I(·) is the indicator function, and if the condition in the parentheses is true, the result is 1, otherwise it is 0.
[0192] Furthermore, assuming that the estimated missing value is x and the training data set is D, It can be expressed by the following formula:
[0193]
[0194] Where, is the estimated missing value of the missing value x, K is the number of nearest neighbors selected, y i is the value of the i-th training sample closest to the missing value x.
[0195] In the above embodiment, the training samples are constructed according to the usage scenario to implement the estimation of the above missing value.
[0196] The above formulas (6) and (7) can be applied to both the case where the second type of data with missing values is a categorical variable and the case of a continuous variable.
[0197] Optionally, the method further includes:
[0198] Performing one-hot encoding on the categorical variable to convert it into a binary variable; and / or
[0199] Performing a linearization transformation on the continuous variable to obtain linear data.
[0200] In this embodiment, through the above one-hot encoding, the reasonable expression of data in the model can be ensured, thereby improving the expression ability of data in the model, reducing the scale difference between features, improving the performance of the model when applied to model training, and improving the stability of the model.
[0201] In the binary variable, some variables take the value of 1 and some variables take the value of 0. This encoding method can effectively handle the magnitude relationship between categorical variables.
[0202] Taking the one-hot encoding of the color column that can be used as a categorical variable as an example:
[0203] Before encoding:
[0204] … Color … Red Green Red
[0205] After encoding:
[0206] … Red Green … 1 0 0 1 1 0
[0207] In this embodiment, when linearizing the continuous variable, for example, it can be achieved through techniques such as logarithmic transformation and normalization to ensure the reasonable expression of data in the model, thereby improving the expression ability of data in the model, reducing the scale difference between features, and improving the performance and stability of the model when applied to model training.
[0208] Optionally, considering that the continuous variable may exhibit a right-skewed (long-tailed) or left-skewed (short-tailed) distribution. Therefore, when linearizing the continuous variable to obtain linear data through logarithmic transformation, specifically, the power relationship of the original data of the continuous variable is transformed into a linear relationship to better display the distribution characteristics of the data, making the transformed data closer to the normal distribution and reducing the impact of skewness such as right-skewed (long-tailed) or left-skewed (short-tailed) on model training.
[0209] In this embodiment, if the linearization of the continuous variable is achieved through normalization, data with different features or different data ranges can be transformed into a unified scale, so that the model can better process and learn, improving the performance and convergence speed of the model.
[0210] Exemplarily, mean normalization can be any one of the following three.
[0211] 1. Min-Max normalization (Min-Max Scaling):
[0212]
[0213] Where X normalized is the normalized value, X is the original data of the continuous variable, X min and X max are the minimum and maximum values of the continuous variable respectively.
[0214] 2. Z-Score normalization (Standardization)
[0215] By zero-centering and unit-variance normalizing the continuous variable, the distribution of the data is made close to the standard normal distribution (mean is 0, standard deviation is 1).
[0216]
[0217] where X normalized is the normalized value, X is the original data of the continuous variable, μ is the mean of the data, and σ is the standard deviation of the data.
[0218] 3. Range Scaling Method (Scaling to Unit Length):
[0219] Scale the data to unit length in the form of a vector, considering the direction of the eigenvector rather than just the magnitude, which is especially applicable to the text data of patients.
[0220]
[0221] where X normalized is the normalized vector, X is the original data vector of the continuous variable, and ‖X‖ represents the L2 norm of the original data vector of the continuous variable.
[0222] Alternatively, if the second type of data with missing values is a continuous variable, in addition to performing approximate sample prediction or statistical filling algorithms on the category based on the set second preprocessing model to generate corresponding estimated missing values, it is also possible to use a classification machine learning algorithm such as a decision tree model. In this case, the continuous variable needs to be converted into a categorical variable, and then the missing value estimation method for categorical variables provided above is used to generate an estimated loss value.
[0223] When converting a continuous variable into a categorical variable, the method of information gain entropy is used to calculate an information gain entropy as the initial splitting point, dividing the continuous variable into two subsets. Then, the information entropy of each subset is calculated, and finally, the best threshold is selected through information gain. Based on this best threshold, patients are classified, such as into a high-risk group and a low-risk group, thus implementing a dichotomy to further judge the logical relationship between the high-risk group and the low-risk group. For example, the case where the risk of getting sick is significantly higher when cholesterol is greater than a certain value than when it is less than a certain value (i.e., high cholesterol accelerates the deterioration of the disease).
[0224] Entropy measures the degree of disorder of a data set, and its calculation formula is as follows:
[0225]
[0226] where S is the data set composed of continuous variables, c is the number of categories of continuous variables, and p i is the proportion of category i in the data set.
[0227] Information Gain is the criterion for selecting the best splitting point in a decision tree, and its calculation formula is as follows:
[0228]
[0229] Among them, S is the parent node data set, A is a continuous variable, Values(A) is the subset divided according to the value of the continuous variable A, and S v is the subset where the value of the continuous variable A is v, |S v | and |S| are the sample numbers of the subset S v and the parent node data set S respectively.
[0230] The steps are as follows:
[0231] 1. Sort the continuous variable A from small to large.
[0232] 2. Select the average of two adjacent values as the threshold and calculate the information gain.
[0233] 3. Repeat step 2 until the information gains of all possible thresholds are calculated.
[0234] 4. Select the threshold with the largest information gain as the classification criterion and divide the continuous variable into two classifications.
[0235] According to the third aspect of the embodiments of the present application, an electronic device is provided, which includes a memory and a processor. A computer-executable program is stored on the memory, and the processor is used to run the computer-executable program to execute the method described in any one of the embodiments of the present application.
[0236] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principle. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present application.
Claims
1. A method for constructing a medical index model, characterized in that Including: Obtain an original feature dataset, where the original feature dataset includes multiple patient data samples, and each patient data sample includes multiple patient features corresponding to the same patient; Sequentially input the patient data samples in the original feature set into a constructed first medical index model for identification to predict the medical index of a patient; Based on the constructed loss function, calculate the loss value of the first medical index model according to the predicted medical index of the patient; Calculate the performance index value of the first medical index model according to the loss value; Select some patient data samples corresponding to when the performance index value is greater than a set index threshold and add them to a pre-constructed feature subset to obtain a training feature subset; Train a specified second medical index model based on the training feature subset.
2. The method according to claim 1, characterized in that, Also including: Perform a normal distribution change process on the original data feature set to obtain data points; Draw a quantity-quantile plot based on the data node positions; Judge whether the data points satisfy the normal distribution according to the quantity-quantile plot; If satisfied, execute the step of sequentially inputting the patient data samples in the original feature set into a constructed first medical index model for identification to predict the medical index of a patient.
3. The method according to claim 1, wherein Also including: Perform sample balance analysis on the training feature subset to obtain a balance estimate. If the balance estimate is less than a set balance threshold, determine the minority sample region in the training feature subset, and generate synthetic samples based on the minority samples in the minority sample region to fill into the training feature subset.
4. The method according to claim 3, characterized in that The generating synthetic samples based on the minority samples in the minority sample region to fill into the training feature subset includes: For any minority sample in the minority sample region, determine at least two adjacent minority samples; Perform a difference operation between the at least two samples to generate the synthetic samples and fill them into the minority sample region.
5. The method according to claim 1, wherein The method further includes: Divide the training feature subset into several folds; Take any selected fold as the test set, and the remaining other folds as the training set, to train a disease identification model based on the training set, and test the second medical index model based on the test set to determine whether the second medical index model trained based on the training set meets the set model performance index.
6. The method according to claim 1, characterized in that The judging whether the second medical index model trained based on the training set meets the set model performance index includes: Calculate the model performance index of the second medical index model when each fold is used as the test set and the other folds are used as the training set; Evaluate whether the second medical index model trained based on the training set meets the set model performance index based on the mean of all model performance indexes.
7. The method according to claim 1, wherein The calculating the loss value of the first medical index model based on the constructed loss function according to the predicted medical index of the patient includes: Based on the constructed loss function, calculate the standard loss value of the first medical index model according to the predicted medical index of the patient; When predicting the medical indicators of the patient, the parameters of the medical indicators and the set regularization hyperparameters are used to calculate the regularization loss value of the first medical indicator model; According to the standard loss value and the regularization loss value, calculate the regression loss value of the first medical indicator model.
8. The method according to claim 1, wherein The method further includes: Calculate the data quality of the model training subset; Based on the data quality, match with the medical indicator model library to match a second medical indicator model from multiple second medical indicator models in the medical indicator model library as the specified second medical indicator model.
9. A method for constructing a medical index model, characterized in that, It includes: A data acquisition unit for acquiring an original feature dataset, where the original feature dataset includes multiple patient data samples, and each patient data sample includes multiple patient features corresponding to the same patient; A pre-training unit for performing the following steps: Sequentially input the patient data samples in the original feature set into the constructed first medical indicator model for identification to predict the medical indicators of the patient; Based on the constructed loss function, according to the predicted medical indicators of the patient, calculate the loss value of the first medical indicator model; According to the loss value, calculate the performance index value of the first medical indicator model; A feature subset construction unit for selecting some patient data samples corresponding to when the performance index value is greater than the set index threshold and adding them to the pre-constructed feature subset to obtain a training feature subset; A model training unit for training the specified second medical indicator model based on the training feature subset.
10. An electronic device, characterized in that, It includes a memory and a processor, and a computer-executable program is stored on the memory, and the processor is used to run the computer-executable program to execute the method according to any one of claims 1-8.