Dengue fever risk prediction method and device based on multi-dimensional index fusion
By integrating multi-dimensional indicators and processing missing data, combined with feature engineering driven by pathological mechanisms, an accurate dengue fever risk prediction system was constructed, which solves the problem of incomplete indicators in existing technologies and achieves more accurate risk prediction and clinical guidance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHANGSHA FIRST HOSPITAL
- Filing Date
- 2026-02-12
- Publication Date
- 2026-06-23
Smart Images

Figure CN122266747A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method and apparatus for predicting dengue fever risk based on the fusion of multi-dimensional indicators. Background Technology
[0002] Dengue fever is an acute vector-borne infectious disease caused by the dengue virus, primarily transmitted by Aedes mosquitoes. Typical clinical symptoms include fever, muscle and joint pain, and rash. In severe cases, it can progress to shock, organ failure, and even death, posing a significant threat to global public health. The World Health Organization estimates that between 50 and 100 million dengue fever cases are contracted globally each year, putting nearly half the population at risk, with particularly high incidence rates in Asia and the Western Pacific.
[0003] Currently, some AI-based dengue fever risk prediction technologies attempt to monitor symptom descriptions and travel history information in electronic medical records. Although these studies have introduced clinical data, they mostly target single features and lack refined strategies for different proportions of missing features. This can easily lead to distortion of feature information or bias in model input, resulting in inaccurate risk prediction. Summary of the Invention
[0004] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.
[0005] The main objective of this invention is to propose a dengue fever risk prediction method and device based on multi-dimensional indicator fusion. Through the fusion of multi-dimensional indicators, refined missing data processing, and feature engineering driven by pathological mechanisms, a robust and accurate dengue fever risk prediction system is constructed. This system can solve the problems of incomplete indicators, improper data processing, and insufficient model performance in existing technologies, and provides doctors with helpful information for guidance.
[0006] To achieve the above objectives, a first aspect of the present invention provides a dengue fever risk prediction method based on multi-dimensional index fusion, the method comprising: In response to a risk prediction command, the system acquires various feature indicators uploaded by the client and determines the missing proportion of the feature indicators; the feature indicators include at least the subject's in vitro blood routine indicators and blood coagulation indicators. When the missing proportion of the feature indicator is greater than a first threshold, the missing state of the feature indicator is encoded as a binary feature, and the indicator distribution of the feature indicator is mapped to a probability feature based on kernel density estimation to obtain the basic feature of the feature indicator; when the missing proportion of the feature indicator is less than or equal to the first threshold and greater than or equal to the second threshold, the feature indicator is imputed by the median, and a deviation coefficient is calculated based on the normal reference interval of the feature indicator. The deviation coefficient is used as a supplementary feature of the feature indicator to obtain the basic feature of the feature indicator; when the missing proportion of the feature indicator is less than the second threshold, the feature indicator is imputed by the median to obtain the basic feature of the feature indicator; the value of the first threshold is greater than the value of the second threshold. Based on the pathological mechanisms associated with dengue fever, determine the pathological association characteristics; The characteristic indicators and the pathological correlation features are input into a preset risk prediction model to obtain the dengue fever risk prediction results output by the risk prediction model.
[0007] The dengue fever risk prediction method based on multi-dimensional index fusion provided in this application has at least the following beneficial effects: This method constructs a robust and accurate dengue fever risk prediction system by fusing multi-dimensional indicators, processing missing data meticulously, and using feature engineering driven by pathological mechanisms. The system can solve the problems of incomplete indicators, improper data processing, and insufficient model performance in existing technologies, and provides doctors with helpful information for guidance.
[0008] To achieve the above objectives, a second aspect of the present invention provides a dengue fever risk prediction device based on multi-dimensional index fusion, the device comprising: The feature indicator acquisition module is used to respond to risk prediction instructions, acquire multiple feature indicators uploaded by the client, and determine the missing proportion of the feature indicators; the feature indicators include at least the subject's in vitro blood routine indicators and blood coagulation indicators. The missing data completion module is used to: encode the missing state of the feature indicator as a binary feature when the missing proportion of the feature indicator is greater than a first threshold, and map the indicator distribution of the feature indicator as a probability feature based on kernel density estimation to obtain the basic feature of the feature indicator; when the missing proportion of the feature indicator is less than or equal to the first threshold and greater than or equal to a second threshold, perform median imputation on the feature indicator, calculate a deviation coefficient based on the normal reference interval of the feature indicator, and use the deviation coefficient as a supplementary feature of the feature indicator to obtain the basic feature of the feature indicator; when the missing proportion of the feature indicator is less than the second threshold, perform median imputation on the feature indicator to obtain the basic feature of the feature indicator; the value of the first threshold is greater than the value of the second threshold. The pathological association feature determination module is used to determine pathological association features based on the pathological mechanisms associated with dengue fever. The risk prediction module is used to input the feature indicators and the pathological correlation features into a preset risk prediction model to obtain the dengue fever risk prediction results output by the risk prediction model.
[0009] To achieve the above objectives, a third aspect of the present invention provides an electronic device, comprising: at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, the instructions being executed by the at least one control processor to enable the at least one control processor to execute the above-described dengue fever risk prediction method based on multi-dimensional index fusion.
[0010] To achieve the above objectives, a fourth aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-described dengue fever risk prediction method based on multi-dimensional indicator fusion.
[0011] It is understood that the beneficial effects of the second to fourth aspects compared with the related technologies are the same as the beneficial effects of the first aspect compared with the related technologies. Please refer to the relevant description in the first aspect above, which will not be repeated here. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram of a dengue fever risk prediction method based on multi-dimensional index fusion provided in an embodiment of this application; Figure 2 This is a schematic diagram of a bar chart showing the number of dengue fever and non-dengue fever samples provided in an embodiment of this application; Figure 3 This is a schematic diagram of a feature correlation heatmap provided in an embodiment of this application; Figure 4 This is a schematic diagram of a cross-validation ROC curve provided in an embodiment of this application; Figure 5 This is a schematic diagram of a test set confusion matrix provided in an embodiment of this application; Figure 6 This is a schematic diagram of a dengue fever risk prediction device based on multi-dimensional index fusion provided in an embodiment of this application; Figure 7 This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0015] like Figure 1 As shown in one embodiment of this application, a method for predicting dengue fever risk based on multi-dimensional index fusion is provided. The method includes: Step S110: Respond to the risk prediction instruction, obtain various feature indicators uploaded by the client, and determine the missing proportion of feature indicators.
[0016] Step S120: When the missing proportion of a feature indicator is greater than a first threshold, the missing state of the feature indicator is encoded as a binary feature, and the indicator distribution of the feature indicator is mapped to a probability feature based on kernel density estimation to obtain the basic feature of the feature indicator; when the missing proportion of a feature indicator is less than or equal to the first threshold and greater than or equal to the second threshold, the feature indicator is imputed by the median, and the deviation coefficient is calculated based on the normal reference interval of the feature indicator. The deviation coefficient is used as a supplementary feature of the feature indicator to obtain the basic feature of the feature indicator; when the missing proportion of a feature indicator is less than the second threshold, the feature indicator is imputed by the median to obtain the basic feature of the feature indicator; the value of the first threshold is greater than the value of the second threshold.
[0017] Step S130: Determine pathological association features based on the pathological mechanisms associated with dengue fever.
[0018] Step S140: Input the characteristic indicators and pathological correlation features into the preset risk prediction model to obtain the dengue fever risk prediction results output by the risk prediction model.
[0019] For ease of understanding, the following explains some key terms in this embodiment: In this embodiment, the characteristic indicators include at least in vitro complete blood count indicators and blood coagulation indicators. Complete blood count indicators typically include white blood cell count, red blood cell count, hemoglobin, platelet count, etc., reflecting the composition and function of blood cells. Blood coagulation indicators include prothrombin time, activated partial thromboplastin time, fibrinogen, etc., used to assess blood coagulation function.
[0020] Binary features refer to features that encode a certain state in the original data into a feature with only two possible values. This encoding method can transform non-numerical or state-based information into numerical inputs that the model can process.
[0021] Kernel density estimation (KDE) is a nonparametric statistical method used to estimate the probability density function of a random variable. KDE allows us to infer the overall distribution of a variable from finite sample data, mapping discrete or continuous index distributions to smooth probabilistic features.
[0022] Median imputation is a common method for handling missing data, which replaces missing values with the median value of all non-missing values of an indicator. Compared to mean imputation, median imputation is less sensitive to outliers and can preserve the original distribution characteristics of the data.
[0023] The normal reference range refers to the statistical range within which a measurement of a biological indicator falls in a healthy population. This range is typically determined by medical laboratories based on the test results of healthy individuals and is used to determine whether an individual's indicator is abnormal.
[0024] A risk prediction model is a trained mathematical model or set of algorithms that can receive input features and output a prediction of dengue fever risk. The model can be a single machine learning model or an ensemble of multiple models.
[0025] In step S110, the method responds to an instruction and acquires various feature indicators uploaded by the client. In practice, the instruction can be issued by the user through a graphical user interface (GUI), such as clicking the "Predict Treatment" button and manually selecting the feature indicator file to be processed. In this method, the feature indicators include at least in vitro blood routine indicators and coagulation indicators. In practice, medical staff can scan paper reports or extract data from unstructured electronic medical records and then upload them. For example, in primary healthcare institutions, doctors may record patients' platelet counts, white blood cell counts, and other blood routine indicators, as well as prothrombin time and other coagulation indicators, in paper medical records. When data is entered into the computer system, if some test results have not yet been returned or have not been performed for some reason, the indicator data will be empty. In this case, it is necessary to count the number of these empty values and calculate their percentage of the total sample, which is the missing value ratio.
[0026] In step S120, this method employs a hierarchical processing strategy to address different levels of missing feature indicators. Specifically, when the missing feature indicator ratio exceeds a first threshold, the missing status of the feature indicator is encoded as a binary feature, and the indicator distribution is mapped to a probability feature based on kernel density estimation to obtain the basic features of the feature indicator. For example, for indicators with a high missing ratio, such as the results of less commonly used cytokine detection, the missing status can be directly encoded as "0" indicating absence and "1" indicating presence.
[0027] When the missing proportion of a feature indicator is less than or equal to the first threshold and greater than or equal to the second threshold, the feature indicator is filled with the median, and the deviation coefficient is calculated based on the normal reference range of the feature indicator. The deviation coefficient is used as a supplementary feature of the feature indicator to obtain the basic feature of the feature indicator.
[0028] Furthermore, when the missing proportion of a feature indicator is less than the second threshold, the feature indicator is imputed using the median to obtain its basic characteristics. It should be noted that the first threshold is greater than the second threshold to ensure the hierarchical relationship of the missing data processing logic.
[0029] In step S130, this method includes determining the pathological association characteristics of the subject based on the pathological mechanisms associated with dengue fever. In practice, combinations of potentially associated indicators can be manually selected based on known knowledge of dengue fever pathophysiology. For example, the ratio of platelet count to white blood cell count, or the ratio of aspartate aminotransferase to alanine aminotransferase, can be calculated as preliminary pathological association characteristics; these ratios can be selected based on experience or literature review.
[0030] In step S140, the feature indicators and pathological correlation features are input into a preset risk prediction model to obtain the dengue fever risk prediction result output by the risk prediction model. The risk prediction model can be a logistic regression model or a decision tree model. For example, the processed blood routine indicators, blood coagulation indicators, and calculated pathological correlation features are used as input, and the model calculates and finally outputs a predicted value for dengue fever risk, such as a probability value between 0 and 1. During model training, cross-validation can be used for evaluation to ensure its generalization ability.
[0031] This embodiment significantly broadens the input dimensions for prediction by identifying multiple characteristic indicators of the subject and explicitly specifying that at least in vitro blood routine indicators and blood coagulation indicators are included. This fusion of multi-dimensional indicators enables the model to comprehensively assess the physiological state of the subject, thus overcoming the limitations of incomplete prediction basis in existing technologies.
[0032] Furthermore, this embodiment proposes a processing strategy based on the hierarchical missing data ratio: for indicators with a high missing data ratio, binary feature encoding and kernel density estimation are used to map them as probabilistic features; for indicators with a medium missing data ratio, median imputation and deviation coefficient calculation are used as supplementary features; and for indicators with a low missing data ratio, median imputation is performed. This sophisticated missing data processing mechanism can retain the effective information in the original data and extract features from different perspectives (missing status, distribution characteristics, and abnormality degree) to effectively address the missing data problem in clinical data, thereby improving the input quality and predictive ability of the model.
[0033] Furthermore, by identifying pathological association features based on the pathological mechanisms associated with dengue fever, the predictive power of the model is further enhanced. This pathology-driven feature engineering enables the model not only to identify statistical associations but also to capture intrinsic biological laws, thereby improving the clinical interpretability and reliability of the prediction results and providing doctors with guiding auxiliary diagnostic information.
[0034] In summary, this method constructs a robust and accurate dengue fever risk prediction system through the fusion of multi-dimensional indicators, refined missing data processing, and pathological mechanism-driven feature engineering. The system can solve the problems of incomplete indicators, improper data processing, and insufficient model performance in existing technologies, and provides doctors with helpful information for guidance.
[0035] In some embodiments of this application, the risk prediction model includes: at least one base classifier and one meta classifier; Before collecting various characteristic indicators of the subjects, the method also includes: Step S210: Train at least one base classifier based on the training set to obtain the predicted probability of dengue fever risk output by each of the at least one base classifier.
[0036] Step S220: Based on the predicted dengue fever risk output by each of the at least one base classifier, train a meta-classifier to obtain the weight values of each of the at least one base classifier output by the meta-classifier.
[0037] Step S230: Based on the predicted probability of dengue fever risk output by at least one basic classifier and its corresponding weight value, the output result of the risk prediction model is obtained.
[0038] Step S240: Backpropagate based on the output results until the trained risk prediction model is obtained.
[0039] Among them, the risk prediction model is an ensemble learning model, the core of which is to improve the overall prediction performance by combining the prediction results of multiple learners.
[0040] This model can employ ensemble strategies such as stacking or weighted voting. The base classifier is the fundamental building block in ensemble learning, its role being to learn from the raw data and generate preliminary predictions. Base classifiers can utilize various machine learning algorithms, such as decision trees, logistic regression, neural networks, and support vector machines. Each base classifier can independently learn different aspects or features of the data. The meta-classifier, also known as a secondary learner, learns how to optimally combine the predictions from the base classifiers. The meta-classifier does not directly operate on the raw data; instead, it takes the predictions from the base classifiers (e.g., predicted probabilities) as input and learns the relationship between these predictions and the true labels to generate the final predictions.
[0041] A training set is a dataset used to train a machine learning model. It contains known input features and corresponding output labels. The model adjusts its internal parameters by learning patterns from the training set.
[0042] The predicted probability of dengue fever risk refers to the model's quantitative assessment of the risk of a subject contracting dengue fever. It is usually a value between 0 and 1, representing the likelihood of contracting the disease.
[0043] Weights are coefficients assigned to the prediction results of each base classifier, indicating the importance or reliability of the base classifier in the final prediction. Weights can be dynamically adjusted based on the performance, diversity, or other criteria of the base classifiers.
[0044] Backpropagation is a commonly used optimization algorithm, especially when training neural networks or ensemble models with differentiable structures. It minimizes prediction error by calculating the gradient of the loss function with respect to the model parameters and updating the parameters in the opposite direction of the gradient.
[0045] This method effectively integrates the predictive advantages of multiple base classifiers and uses a meta-classifier for intelligent weighting and combination, thereby significantly improving the accuracy and generalization ability of dengue fever risk prediction. This ensemble learning approach can better handle the complexity and nonlinear relationships of multi-dimensional indicators, reduce the risk of overfitting that may exist in a single model, and make the prediction results more stable and reliable, providing a more accurate basis for clinical diagnosis and early intervention.
[0046] In some embodiments of this application, the base classifier is a random forest, a lightweight gradient booster, or a support vector machine.
[0047] The base classifier is the basic building block in an ensemble learning model. Its role is to perform preliminary classification or prediction of the input data. Each base classifier independently learns the patterns and rules in the data and outputs its prediction results for the target variable. These prediction results are then integrated and further combined and decided by the meta-classifier to form the final risk prediction. The choice of base classifier directly affects the performance of the ensemble model. Algorithms with different learning mechanisms and advantages are usually selected to improve the robustness and diversity of the model.
[0048] Random forest is an ensemble learning method that improves the accuracy and stability of a model by constructing multiple decision trees and combining their predictions. Lightweight gradient boosting machine is an optimization algorithm based on gradient boosting decision trees (GBDT), which significantly improves training speed and reduces memory consumption while maintaining high accuracy. Support vector machine is a binary classification model whose basic idea is to find an optimal hyperplane in the feature space such that sample points of different classes are separated by the hyperplane with maximum margin.
[0049] This method effectively improves the model's ability to predict dengue fever risk by using random forests, lightweight gradient boosters, or support vector machines as the base classifiers in the risk prediction model. During the training phase, the base classifiers are trained separately, each learning different patterns and feature combinations from the training data. For example, random forests, by integrating multiple decision trees, leverage their inherent randomness and parallelism to capture nonlinear relationships in the data and reduce the risk of overfitting; lightweight gradient boosters, through an iterative optimization process of gradient boosting, gradually correct prediction errors, are particularly adept at handling large-scale and high-dimensional data, and can quickly converge to a better solution; support vector machines, by finding the optimal classification hyperplane, achieve maximum margin classification in the feature space, exhibiting strong generalization ability for well-defined but complex classification problems. These base classifiers, with their different learning mechanisms and advantages, each output a predicted probability for dengue fever risk. Subsequently, the meta-classifier learns how to optimally combine the outputs of these base classifiers based on these predicted probabilities and assigns corresponding weight values to each base classifier. Through this stacked ensemble architecture, the model can fully utilize the advantages of different base classifiers and make up for the shortcomings of a single model, thereby forming a more powerful and robust risk prediction model. This combination method enables the model to understand and analyze the characteristic indicators and pathological correlation features of subjects from multiple perspectives and levels, effectively addressing the challenges of high data complexity and strong feature correlation in dengue risk prediction, and ultimately obtaining more accurate and reliable dengue risk prediction results.
[0050] This method uses random forest, lightweight gradient booster, or support vector machine as the base classifier for the risk prediction model. This method can significantly improve the accuracy and stability of dengue fever risk prediction. Each of these specific base classifiers has a unique learning mechanism and advantages. For example, random forest has the ability to resist overfitting and handle high-dimensional data, lightweight gradient booster has the efficiency and advantages in handling large-scale data, and support vector machine has the strong generalization ability and robustness to small sample data. When these high-performance and diverse base classifiers are integrated into a stacked model, they can capture complex patterns and potential correlations in the data from different perspectives, effectively making up for the limitations that may exist in a single model. This multi-model fusion strategy enables the model to show stronger adaptability and robustness when facing challenges such as nonlinearity, high dimensionality, and imbalance common in dengue fever clinical data, providing clinicians with more reliable early warnings.
[0051] In some embodiments of this application, after obtaining the output of the risk prediction model based on the predicted probabilities of dengue fever risk and their corresponding weight values output by at least one base classifier, the method further includes: By combining hierarchical 5-fold cross-validation with Bayesian optimization, the hyperparameters of at least one base classifier and one meta-classifier are optimized.
[0052] Hierarchical 5-fold cross-validation is a commonly used model evaluation technique. It divides the dataset into 5 subsets and ensures that the sample proportion of each category in each subset is consistent with that of the original dataset. In each iteration, one subset is selected as the validation set and the other 4 subsets are used as the training set. This process is repeated 5 times to evaluate the model's performance multiple times and take the average to obtain a more stable and reliable performance estimate. This method can effectively reduce the randomness of model evaluation results and is especially suitable for imbalanced datasets, such as when there may be few positive cases in dengue fever risk prediction.
[0053] Bayesian optimization is a highly efficient global optimization algorithm, particularly suitable for optimizing computationally expensive black-box functions. It finds the optimal hyperparameter combination in fewer iterations by constructing a probabilistic surrogate model of the objective function and using a sampling function to guide the selection of the next evaluation point. Compared to traditional grid search or random search, Bayesian optimization explores the hyperparameter space more intelligently, balancing exploration and utilization, and significantly improving the efficiency of hyperparameter tuning.
[0054] Hyperparameters are parameters that need to be pre-set before training a machine learning model. They are not learned during the training process but directly affect the model's learning ability and final performance. For the basic classifiers and meta-classifiers mentioned above, hyperparameters may include, but are not limited to, the maximum depth of the decision tree model, the number of weak learners in the ensemble learning model, and regularization parameters. The appropriate selection of these hyperparameters is crucial for improving the model's generalization ability and prediction accuracy.
[0055] This method effectively addresses the problems of low optimization efficiency, unstable model performance, and insufficient ability to handle imbalanced data in dengue fever risk prediction models within complex hyperparameter spaces. Bayesian optimization significantly improves the efficiency of hyperparameter search and reduces computational resources and time consumption, while hierarchical 5-fold cross-validation ensures the accuracy and robustness of model performance evaluation. Especially in medical datasets like dengue fever datasets where sample class imbalance may exist, this optimization strategy enables the trained risk prediction model to have stronger generalization ability and higher prediction accuracy, thus providing more reliable technical support for early warning and clinical decision-making in dengue fever.
[0056] In some embodiments of this application, when training at least one base classifier, a hybrid sampling strategy of SMOTE oversampling and random undersampling of the majority class is used on the training set.
[0057] Training at least one base classifier is the process of adjusting the internal parameters of the classifier using existing labeled data (training set) so that it can learn patterns and rules in the data, thereby accurately classifying or predicting new and unseen data. In the dengue fever risk prediction scenario, this means teaching the classifier how to judge the risk of a subject contracting dengue fever based on the subject's characteristic indicators.
[0058] The training process can iteratively update the model weights using the gradient descent algorithm to minimize the prediction error; or recursively partition the data space using the decision tree construction algorithm to generate a series of decision rules; for support vector machines, the optimal hyperplane is found to maximize the margin between samples of different classes.
[0059] SMOTE oversampling is a strategy for synthesizing minority class samples to address class imbalance in datasets. It increases the number of minority class samples by interpolating between them to generate new synthetic samples, allowing the model to better learn the features of the minority class during training and avoid bias towards the majority class. The SMOTE algorithm finds the K nearest neighbors for each minority class sample, then randomly selects a sample from these neighbors and randomly generates a new synthetic sample between the original sample and the selected nearest neighbor. In addition, variants such as Borderline-SMOTE can be used to oversample only those minority class samples close to the decision boundary.
[0060] Random undersampling of the majority class is a strategy to reduce the number of majority class samples in a dataset, also used to solve the class imbalance problem. It balances the dataset by randomly removing some majority class samples, making the number of majority and minority class samples closer, thus preventing the model from overlearning the features of the majority class. This strategy can directly select a portion of the majority class samples and remove them from the training set until the number of majority class samples reaches a preset proportion. It can also be combined with other strategies, such as clustering the majority class samples before undersampling and then selecting representative samples from each cluster for undersampling. Hybrid sampling strategy refers to combining oversampling and undersampling techniques to deal with the class imbalance problem. This strategy aims to make full use of the advantages of oversampling to increase minority class information and undersampling to reduce majority class noise, so as to more effectively balance the dataset, thereby improving the model's ability to recognize all classes and the overall performance.
[0061] This method effectively solves the model training bias problem caused by class imbalance in dengue risk prediction. By using SMOTE oversampling to increase the representativeness of minority class samples, the base classifier can fully learn the characteristics of dengue-positive cases. At the same time, random undersampling of the majority class reduces redundant information in the majority class samples, avoiding overfitting of the model to the majority class. This hybrid sampling strategy significantly enhances the dengue risk identification ability of the trained base classifier, especially in identifying key positive cases. This improves the accuracy and robustness of the entire multi-dimensional index fusion dengue risk prediction model, providing a more reliable basis for clinical diagnosis.
[0062] In some embodiments of this application, the calculation process of the deviation coefficient includes: Step S310: Calculate the absolute value of the difference between the filled value and the normal reference interval; Step S320: Determine the deviation coefficient based on the absolute value and the normal reference range.
[0063] The calculation of the absolute value of the difference between the filled value and the normal reference interval aims to quantify the deviation of the feature index value after median filling from the preset normal reference interval. By calculating the absolute value, it can be ensured that regardless of whether the filled value is higher or lower than the normal reference interval, the deviation is uniformly represented as a positive value, thus avoiding the influence of the positive and negative signs on the subsequent calculation of the deviation coefficient, making the measurement of the deviation more objective. As one possible implementation method, this can be achieved by comparing the filled value with the center value of the normal reference interval and calculating the absolute value of the difference.
[0064] Based on the absolute value and the normal reference range, the deviation coefficient is determined. This step, based on the absolute difference and normal reference range obtained in the previous step, further quantifies the degree of abnormality of the characteristic indicator and standardizes it into a deviation coefficient. The deviation coefficient can intuitively reflect the severity of the characteristic indicator's deviation from the normal range, providing a more interpretable and discriminative input for subsequent risk prediction models. As one possible implementation, the deviation coefficient can be defined as the ratio of the absolute difference to the width of the normal reference range. Alternatively, a piecewise function or a nonlinear function can be used to determine the deviation coefficient. For example, when the absolute difference is small, the deviation coefficient increases linearly; when the absolute difference is large, the deviation coefficient increases at a faster rate to highlight extreme anomalies.
[0065] This embodiment addresses the problem of effectively quantifying the deviation of a feature indicator from the normal reference range when the proportion of missing features falls within a specific range by clarifying the calculation process of the deviation coefficient. Specifically, when the proportion of missing features is less than or equal to a first threshold and greater than or equal to a second threshold, the missing data is first imputed using the median to ensure data integrity. Based on this, to capture the degree of deviation between the imputed value and the normal physiological range, this method further calculates the absolute value of the difference between the imputed value and the normal reference range. This absolute difference directly reflects the distance of the feature indicator's value from the normal range. Subsequently, the deviation coefficient is determined based on this absolute value and the normal reference range. By correlating the absolute difference with the normal reference range, for example through ratio or normalization, a standardized and comparable deviation coefficient can be obtained, which can more accurately characterize the degree of abnormality of the feature indicator. This calculation method ensures that the deviation coefficient not only reflects the magnitude of the deviation but also considers the background of the normal reference range, thus providing the risk prediction model with more biologically meaningful and discriminative supplementary features, effectively improving the model's ability to identify dengue fever risk.
[0066] In some embodiments of this application, pathological associated features include: Platelets / WBC ratio, AST / ALT ratio, and LDH / Albumin ratio.
[0067] The Platelets / WBC ratio refers to the ratio of platelet count to white blood cell count in a subject's blood. Thrombocytopenia and leukopenia are common clinical manifestations of dengue infection. This ratio can comprehensively reflect changes in immune cells and coagulation function during infection, providing potential indicators of disease progression and severity. The AST / ALT ratio refers to the ratio of aspartate aminotransferase (AST) to alanine aminotransferase (ALT) in a subject's blood. AST and ALT are important indicators of liver function. Dengue infection may cause liver damage, and changes in these ratios are significant in the diagnosis and assessment of liver disease. They can be used to assess liver involvement in dengue risk prediction. The LDH / Albumin ratio refers to the ratio of lactate dehydrogenase (LDH) to albumin (Albumin) in a subject's blood. LDH is a non-specific indicator reflecting cell damage and tissue necrosis, while Albumin reflects liver synthetic function and nutritional status. Severe dengue fever patients often have elevated LDH and decreased Albumin. This ratio can comprehensively reflect the body's inflammatory response, the degree of tissue damage, and pathophysiological changes such as fluid leakage.
[0068] This method incorporates these clinically significant composite features into the risk prediction model, which not only enhances the model's biological interpretability but also significantly improves the accuracy and reliability of dengue fever risk prediction, helping clinicians to make early diagnoses, risk stratify, and develop more effective interventions.
[0069] For ease of understanding, such as Figures 2 to 5 One embodiment of this application provides a dengue fever risk prediction method based on multi-dimensional indicator fusion, the method comprising the following steps: Step S910: Data acquisition and preprocessing; Step S9110, Data Acquisition; The experimental data contains 928 records, involving 13 features and 1 diagnostic label, as detailed below: (1) Characteristic indicators: Complete blood count indicators: WBC (white blood cell count), RBC (red blood cell count), Hemoglobin, Platelets; Blood coagulation parameters: PT (prothrombin time), APTT (activated partial thromboplastin time), TT (thrombin time), D-Dimer (D-dimer); Biochemical / inflammatory / liver function indicators: LDH (lactate dehydrogenase), AST (aspartate aminotransferase), ALT (alanine aminotransferase), Total Protein, Albumin.
[0070] (2) Diagnostic label: Dengue (1=dengue fever, 0=non-dengue fever).
[0071] (3) Sample distribution: 171 cases of dengue fever (18.4%) and 757 cases of non-dengue fever (81.6%). The data dimension is (928,14), and there is a certain imbalance of categories.
[0072] Step S9120, preprocessing; First, a differentiated processing strategy is adopted based on the missing percentage of each indicator: Table 1
[0073] High missing rate indicators (missing ratio > 85%): For PT, APTT, TT, and D-Dimer, the "missing state encoding + indicator distribution mapping" method is used. "Whether it is missing" is used as a binary feature. At the same time, the original indicator distribution is mapped to a probability feature through kernel density estimation (KDE) to retain the potential diagnostic information of missing indicators. Missing rate indicator processing (30%≤missing ratio≤65%): For LDH, Total-Protein, and Albumin, “median imputation + missing bias correction” is adopted. After imputation, the bias coefficient is calculated through the normal reference interval of the indicator (bias coefficient = |imputed value - median of reference interval| / reference interval range) as a supplementary feature; Handling of low missing rate indicators (missing rate <10%): Median imputation was used for WBC, RBC, Hemoglobin, Platelets, AST, and ALT; Secondly, data normalization: MinMaxScaler is used to normalize all indicators and derived features to the [0,1] interval to eliminate dimensional differences.
[0074] Finally, a feature set was constructed based on the pathological mechanisms of dengue fever (thrombocytopenia, hepatocellular damage, and inflammatory response): Basic features: preprocessed original indicators and missing related derived features; Pathological correlation features: Calculate the Platelets / WBC ratio (reflecting the characteristic relative decrease in platelets in dengue fever), the AST / ALT ratio (reflecting the type of hepatocellular damage), and the LDH / Albumin ratio (reflecting the correlation between inflammation and nutritional status).
[0075] Step S920, Exploratory data analysis; Step S9210: Visualize the sample distribution; (1) Bar chart: There is a significant difference between the number of dengue fever patients (171 cases) and non-dengue fever patients (757 cases), with non-dengue fever samples accounting for more than 80%, which confirms the problem of class imbalance.
[0076] (2) Pie chart: The ratio of dengue fever to non-dengue fever samples is approximately 1:4.4, suggesting that the model needs to focus on handling imbalanced data to avoid bias towards the majority class.
[0077] Step S9220, Feature Correlation Analysis; (3) Heat map display: In routine blood tests, RBC and Hemoglobin showed a strong positive correlation (correlation coefficient > 0.8), reflecting a synergistic change in the oxygen-carrying capacity of red blood cells; Among the biochemical indicators, AST and ALT (liver function indicators) showed a moderate positive correlation (correlation coefficient ≈ 0.6), suggesting that hepatocellular damage may affect the levels of both simultaneously. Platelets (platelet count) showed a negative correlation with the Dengue label (correlation coefficient ≈ -0.5), consistent with the clinical characteristics of thrombocytopenia in dengue fever patients, and is a key diagnostic indicator.
[0078] Step S930: Construct an integrated architecture of "basic classifier layer + meta classifier layer" to solve the class imbalance problem; Base classifier layer: Three heterogeneous classifiers are trained in parallel, including, Random Forest (RF, n_estimators=300, class_weight='balanced'); Lightweight gradient booster (LightGBM, num_leaves=31, learning_rate=0.05); Support Vector Machine (SVM, kernel function RBF, class_weight='balanced'); Each classifier outputs a predicted probability (probability of association with dengue fever). During the training phase of the basic classifier, a hybrid sampling strategy of "SMOTE oversampling (minority class) + random undersampling (majority class)" is adopted for the training set to balance the sample distribution. Meta-classifier layer: Taking the predicted probabilities of the three base classifiers as input, logistic regression is trained as the meta-classifier. The results of the base classifiers are fused through weighted voting to output the final association probability. Model optimization: A combination of hierarchical 5-fold cross-validation and Bayesian optimization was used to optimize the hyperparameters of each classifier and ensure model stability.
[0079] Step S940, Decision Support Output Module; Output risk levels (low risk: probability < 0.3; medium risk: 0.3 ≤ probability < 0.7; high risk: probability ≥ 0.7) to assist clinicians in making decisions regarding dengue fever screening and further testing.
[0080] Step S950, Model Validation Module; The performance was validated using five metrics: accuracy, precision, recall, F1-score, and AUC. The test set was required to have an accuracy of ≥97%, a recall of ≥92%, and an AUC of ≥0.99 to ensure the clinical applicability of the model.
[0081] Step S960, Explanation of technical advantages; Compared with the prior art, the advantages of this embodiment are: (1) Breaking through the limitations of the traditional "direct deletion of missing indicators", the potential information of blood coagulation indicators with high missing rate is retained by missing state coding and distribution mapping; and the associated features are constructed in combination with the pathological mechanism of dengue fever, making the features more diagnostically targeted and solving the problem of insufficient utilization of existing model features. (2) A hierarchical ensemble learning architecture is adopted, which integrates the advantages of heterogeneous classifiers. The class imbalance problem is solved by hybrid sampling and weighted voting. Compared with a single classification model, the dengue fever recall rate is improved by 5%-8%, and the false negative rate is significantly reduced. The combination of Bayesian optimization and hierarchical cross-validation ensures the stability of the model on different sample subsets. (3) Based on routine clinical test indicators, no additional test costs are required, and the operation is simple; the model has a fast calculation speed (prediction time for a single sample < 1 second), and the output risk level is intuitive, making it easy for primary care physicians to use quickly.
[0082] like Figure 6 As shown in one embodiment of this application, a dengue fever risk prediction device based on multi-dimensional index fusion is provided. The device includes: The feature indicator acquisition module 1100 is used to respond to risk prediction instructions, acquire multiple feature indicators uploaded by the client, and determine the missing proportion of the feature indicators; the feature indicators include at least the subject's in vitro blood routine indicators and blood coagulation indicators. The missing data completion module 1200 is used to: encode the missing state of the feature indicator as a binary feature when the missing proportion of the feature indicator is greater than a first threshold, and map the indicator distribution of the feature indicator as a probability feature based on kernel density estimation to obtain the basic feature of the feature indicator; when the missing proportion of the feature indicator is less than or equal to the first threshold and greater than or equal to a second threshold, perform median imputation on the feature indicator, calculate a deviation coefficient based on the normal reference interval of the feature indicator, and use the deviation coefficient as a supplementary feature of the feature indicator to obtain the basic feature of the feature indicator; when the missing proportion of the feature indicator is less than the second threshold, perform median imputation on the feature indicator to obtain the basic feature of the feature indicator; the value of the first threshold is greater than the value of the second threshold. The pathological association feature determination module 1300 is used to determine pathological association features based on the pathological mechanisms associated with dengue fever; The risk prediction module 1400 is used to input the feature indicators and the pathological correlation features into a preset risk prediction model to obtain the dengue fever risk prediction result output by the risk prediction model.
[0083] It should be noted that the dengue fever risk prediction device based on multi-dimensional index fusion provided in this embodiment and the dengue fever risk prediction method based on multi-dimensional index fusion described above are based on the same inventive concept. Therefore, the relevant content of the dengue fever risk prediction method based on multi-dimensional index fusion described above also applies to the content of the dengue fever risk prediction device based on multi-dimensional index fusion. Therefore, it will not be repeated here.
[0084] likeFigure 7 This application also provides an electronic device, which includes: At least one memory; At least one processor; At least one program; The program is stored in memory, and the processor executes at least one program to implement the dengue fever risk prediction method based on multi-dimensional index fusion described above in this disclosure.
[0085] Electronic devices can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.
[0086] The electronic devices according to embodiments of this application will now be described in detail.
[0087] The processor 1600 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention. The memory 1700 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1700 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1700 and is called and executed by the processor 1600 to implement the dengue fever risk prediction method based on multi-dimensional indicator fusion of the embodiments of the present invention.
[0088] The input / output interface 1800 is used to implement information input and output. The communication interface 1900 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 2000 transmits information between various components of the device (e.g., processor 1600, memory 1700, input / output interface 1800, and communication interface 1900); The processor 1600, memory 1700, input / output interface 1800 and communication interface 1900 are connected to each other within the device via bus 2000.
[0089] This invention also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the aforementioned dengue fever risk prediction method based on multi-dimensional index fusion.
[0090] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0091] The embodiments described in this invention are intended to more clearly illustrate the technical solutions of the embodiments of this invention, and do not constitute a limitation on the technical solutions provided by the embodiments of this invention. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this invention are also applicable to similar technical problems.
[0092] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0093] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0094] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0095] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0096] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0097] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0098] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0099] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0100] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0101] The above is a detailed description of the preferred embodiments of this application. However, the embodiments of this application are not limited to the above-described implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the embodiments of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of the embodiments of this application.
Claims
1. A dengue fever risk prediction method based on multi-dimensional indicator fusion, characterized in that, The method includes: In response to a risk prediction command, the system acquires various feature indicators uploaded by the client and determines the missing proportion of the feature indicators; the feature indicators include at least the subject's in vitro blood routine indicators and blood coagulation indicators. When the missing proportion of the feature indicator is greater than a first threshold, the missing state of the feature indicator is encoded as a binary feature, and the indicator distribution of the feature indicator is mapped to a probability feature based on kernel density estimation to obtain the basic feature of the feature indicator; when the missing proportion of the feature indicator is less than or equal to the first threshold and greater than or equal to the second threshold, the feature indicator is imputed by the median, and a deviation coefficient is calculated based on the normal reference interval of the feature indicator. The deviation coefficient is used as a supplementary feature of the feature indicator to obtain the basic feature of the feature indicator; when the missing proportion of the feature indicator is less than the second threshold, the feature indicator is imputed by the median to obtain the basic feature of the feature indicator; the value of the first threshold is greater than the value of the second threshold. Based on the pathological mechanisms associated with dengue fever, determine the pathological association characteristics; The characteristic indicators and the pathological correlation features are input into a preset risk prediction model to obtain the dengue fever risk prediction results output by the risk prediction model.
2. The dengue fever risk prediction method based on multi-dimensional index fusion according to claim 1, characterized in that, The risk prediction model includes: at least one base classifier and one meta-classifier; Before obtaining the various feature indicators uploaded by the client, the method further includes: Train the at least one base classifier based on the training set to obtain the predicted probability of dengue fever risk output by each of the at least one base classifier. The meta-classifier is trained based on the predicted dengue fever risk output by each of the at least one base classifier, so as to obtain the weight values of each of the at least one base classifier output by the meta-classifier. The output of the risk prediction model is obtained based on the predicted probability of dengue fever risk output by each of the at least one basic classifier and its corresponding weight value. The process is repeated backpropagating based on the output until the trained risk prediction model is obtained.
3. The dengue fever risk prediction method based on multi-dimensional index fusion according to claim 2, characterized in that, The base classifier is a random forest, a lightweight gradient booster, or a support vector machine.
4. The dengue fever risk prediction method based on multi-dimensional index fusion according to claim 2, characterized in that, After obtaining the output of the risk prediction model based on the predicted dengue fever risk output by each of the at least one base classifier and its corresponding weight value, the method further includes: The hyperparameters of at least one base classifier and one meta classifier are optimized by combining hierarchical 5-fold cross-validation with Bayesian optimization.
5. The dengue fever risk prediction method based on multi-dimensional index fusion according to claim 3, characterized in that, When training the at least one base classifier, a hybrid sampling strategy of SMOTE oversampling and random undersampling of the majority class is used on the training set.
6. The dengue fever risk prediction method based on multi-dimensional index fusion according to claim 1, characterized in that, The calculation process of the deviation coefficient includes: Calculate the absolute value of the difference between the filled value and the normal reference interval; The deviation coefficient is determined based on the absolute value and the normal reference range.
7. The dengue fever risk prediction method based on multi-dimensional index fusion according to claim 1, characterized in that, The pathological association features include: Platelets / WBC ratio, AST / ALT ratio, and LDH / Albumin ratio.
8. A dengue fever risk prediction device based on multi-dimensional index fusion, characterized in that, The device includes: The feature indicator acquisition module is used to respond to risk prediction instructions, acquire multiple feature indicators uploaded by the client, and determine the missing proportion of the feature indicators; the feature indicators include at least the subject's in vitro blood routine indicators and blood coagulation indicators. The missing data completion module is used to: encode the missing state of the feature indicator as a binary feature when the missing proportion of the feature indicator is greater than a first threshold, and map the indicator distribution of the feature indicator as a probability feature based on kernel density estimation to obtain the basic feature of the feature indicator; when the missing proportion of the feature indicator is less than or equal to the first threshold and greater than or equal to a second threshold, perform median imputation on the feature indicator, calculate a deviation coefficient based on the normal reference interval of the feature indicator, and use the deviation coefficient as a supplementary feature of the feature indicator to obtain the basic feature of the feature indicator; when the missing proportion of the feature indicator is less than the second threshold, perform median imputation on the feature indicator to obtain the basic feature of the feature indicator; the value of the first threshold is greater than the value of the second threshold. The pathological association feature determination module is used to determine pathological association features based on the pathological mechanisms associated with dengue fever. The risk prediction module is used to input the feature indicators and the pathological correlation features into a preset risk prediction model to obtain the dengue fever risk prediction results output by the risk prediction model.
9. An electronic device, characterized in that, include: At least one control processor and a memory for communicatively connecting to the at least one control processor; The memory stores instructions that can be executed by the at least one control processor, which, when executed by the at least one control processor, enables the at least one control processor to perform a dengue fever risk prediction method based on multi-dimensional index fusion as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform a dengue fever risk prediction method based on multi-dimensional index fusion as described in any one of claims 1 to 7.