A system and method for drug risk assessment based on genetic polymorphisms
By analyzing CYP2D6 and ADRB1 gene polymorphisms, collecting and processing metoprolol-related medical data, screening features, and establishing a risk prediction model, the problem of individual differences in the response of metoprolol to treatment in elderly patients was solved, and more reliable risk assessment and individualized medication regimens were achieved.
Patent Information
- Application Number
- CN202411302945.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-09-19
AI Technical Summary
The existing technology for metoprolol shows inter-individual variability in treatment response in elderly patients with multiple coexisting diseases, multiple medications, and impaired liver and kidney function. Furthermore, medical data processing is cumbersome and makes effective risk assessment difficult.
By analyzing the polymorphisms of CYP2D6 and ADRB1 genes, relevant medical data were collected, data preprocessing and feature screening were performed, a training dataset was established, and a risk prediction model was used for evaluation, including univariate analysis and Spellman correlation method for feature screening, oversampling to balance the sample, and establishing a risk prediction model.
A more reliable and interpretable risk prediction model was constructed. The model was simplified, redundant features were removed, and the interpretability of features was improved, providing a theoretical basis for personalized metoprolol dosing regimens.
Smart Images

Figure CN119252508B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical data processing technology, and in particular to a drug risk assessment system and method based on gene polymorphism. Background Technology
[0002] In existing technologies, the selective β1 receptor blocker metoprolol is a drug for cardiovascular diseases in the elderly. However, in elderly patients with multiple coexisting diseases, multiple medications, or impaired liver and kidney function, the response to metoprolol treatment exhibits significant inter-individual variability, posing a challenge to the clinical treatment of elderly patients with cardiovascular diseases. Cytochrome P450 2D6 (CYP2D6) gene polymorphism is an important reason for the inter-individual variability in metoprolol response.
[0003] However, existing methods for processing medical data are overly cumbersome. Furthermore, there is still significant room for improvement in areas such as how to effectively assess the risks of metoprolol and how to efficiently process medical data.
[0004] Therefore, it is necessary to provide a drug evaluation method to address the above-mentioned problems. Summary of the Invention
[0005] The present invention aims to provide a drug risk assessment system and method based on gene polymorphism to solve the technical problems of how to effectively assess the risk of metoprolol and effectively process medical data in the prior art. The technical problems to be solved by the present invention are achieved through the following technical solutions.
[0006] The first aspect of this invention proposes a drug risk assessment method based on gene polymorphism, comprising: collecting historical medical data; processing the collected historical medical data, including disease data, drug dosage, physiological indicator data, and test data; performing data preprocessing on the collected historical medical data; and performing feature screening according to specified screening rules, wherein the feature screening includes: calculating the correlation between each feature using univariate analysis for the first screening, and calculating the correlation between features using the Spellman correlation method for the second screening; defining positive and negative samples based on the selected features, and establishing a training dataset, specifically including: oversampling a small number of samples before establishing the training dataset to ensure that the number of positive and negative samples is equal; establishing a risk prediction model; training the risk prediction model using the training dataset to obtain a trained risk prediction model; and inputting the drug dosage and test data of the object to be assessed into the trained risk prediction model to obtain the risk assessment result.
[0007] The second aspect of this invention proposes a drug risk assessment system based on gene polymorphism, employing the drug risk assessment method based on gene polymorphism described in the first aspect of this invention. The system includes: a data collection module for collecting historical medical data, including disease data, drug dosage, physiological indicator data, and test data; a feature screening module for preprocessing the collected historical medical data and screening features according to specified screening rules, including: calculating the correlation between each feature using univariate analysis for the first screening and calculating the correlation between features using the Spellman correlation method for the second screening; a first establishment module for defining positive and negative samples based on the selected features and establishing a training dataset, specifically including: oversampling a small number of samples before establishing the training dataset to ensure that the number of positive and negative samples is equal; a second establishment module for establishing a risk prediction model, training the risk prediction model using the training dataset to obtain a trained risk prediction model; and a result output module for inputting the drug dosage and test data of the object to be assessed into the trained risk prediction model to obtain the risk assessment result.
[0008] A third aspect of the present invention provides an electronic device comprising: one or more processors; a storage device for storing one or more programs; wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in the first aspect of the present invention.
[0009] A fourth aspect of the present invention provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect of the present invention.
[0010] The embodiments of the present invention have the following advantages:
[0011] Compared with existing technologies, this invention analyzes the polymorphisms of the CYP2D6 and ADRB1 genes, records and collects relevant medical data on metoprolol tolerance, adverse reactions, blood pressure and heart rate responses, and long-term prognosis. The collected medical data is then processed to select multiple features to establish a training dataset and a risk prediction model. This model is trained using the training dataset to obtain a well-trained risk prediction model. The drug dosage and test data of the subjects to be evaluated are input into the trained risk prediction model to obtain risk assessment results. This provides a theoretical basis for establishing gene polymorphism-guided personalized metoprolol medication regimens.
[0012] In addition, univariate analysis was used to calculate the features for the first screening, and the Spellman correlation method was used to calculate the feature correlation for the second screening. The two screening processes can help build a more reliable and interpretable model, simplify the model, remove redundant features, and improve the interpretability of features. Attached Figure Description
[0013] Figure 1 This is a flowchart illustrating an example of the drug evaluation method of the present invention;
[0014] Figure 2 This is a partial flowchart illustrating the drug evaluation method of the present invention;
[0015] Figure 3 This is another partial flowchart example of the drug evaluation method of the present invention;
[0016] Figure 4 This is a structural block diagram of the drug evaluation system of the present invention;
[0017] Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention;
[0018] Figure 6 This is a schematic diagram of a computer-readable medium embodiment according to the present invention. Detailed Implementation
[0019] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0020] In view of the above problems, this invention proposes a drug evaluation method. By analyzing the polymorphisms of CYP2D6 and ADRB1 genes, relevant medical data on metoprolol tolerance, adverse reactions, blood pressure and heart rate responses, and long-term prognosis are recorded and collected. The collected medical data is then processed to select multiple features to establish a training dataset and a risk prediction model. The risk prediction model is trained using the training dataset to obtain a well-trained risk prediction model. The drug dosage and test data of the subject to be evaluated are input into the well-trained risk prediction model to obtain the risk assessment results. This provides a theoretical basis for establishing a gene polymorphism-guided personalized metoprolol medication regimen.
[0021] Example 1: Figure 1 This is a flowchart illustrating an example of the drug evaluation method of the present invention.
[0022] The following reference Figure 1 , Figure 2 and Figure 3 The present invention will be described in detail below.
[0023] First, in step S101, historical medical data is collected, including disease data, drug dosage, physiological indicator data, and test data.
[0024] In one specific implementation, for example, historical medical data on the use of metoprolol in a hospital during a specified historical period (e.g., from April 2017 to July 2019) is collected.
[0025] For example, the preferred users are elderly people aged 65 and above. Specifically, there are 1039 male users.
[0026] Specifically, historical medical data includes disease history, medication dosage, physiological indicators, and laboratory data. Disease history includes cardiovascular disease. Medication dosage includes the initial metoprolol dose (e.g., converting all metoprolol doses to metoprolol tartrate doses, specifically metoprolol tartrate 50mg). Physiological indicators include the age and weight of historical users, blood pressure, heart rate, and genetic polymorphism data before and after medication use. The genetic polymorphism data is obtained, for example, by using Sanger sequencing to determine polymorphisms at the CYP2D6*2 (rs16947), *10 (rs1065852), *14 (rs5030865), and ADRB1 (rs1801253) sites.
[0027] In addition, the test data included whether a clinical endpoint event occurred and the date of occurrence. Furthermore, the metoprolol dosage and adverse reactions of the aforementioned users were also collected. Specifically, data was collected at equal intervals for a specific duration (e.g., 12 weeks, 24 weeks, or 36 weeks). For example, blood pressure, heart rate, metoprolol maintenance dose, and related adverse reactions were collected every two weeks. These related adverse reactions specifically included cardiovascular and non-cardiovascular adverse reactions. Cardiovascular adverse reactions were defined as orthostatic hypotension, bradycardia (heart rate ≤55 bpm), cardiac arrest, second-degree or higher atrioventricular block, syncope, and cold extremities. Non-cardiovascular adverse reactions were defined as dyspnea, sleep disturbances and fatigue, headache or dizziness, and depression.
[0028] Specifically, if metoprolol was not discontinued during the last data collection, the metoprolol dose is defined as the metoprolol maintenance dose.
[0029] In another implementation, for example, publicly available medical record data from an Electronic Medical Record (EMR) system can be obtained. Privacy information in the obtained data is anonymized before use. Furthermore, because real-world data suffers from issues such as missing patient information, incomplete medical records, and significant discrepancies between records, censoring and outlier handling are also necessary.
[0030] It should be noted that the above-mentioned methods for collecting medical data can also be achieved through a combination of outpatient inquiries, access to electronic medical record systems, and telephone follow-ups. Furthermore, data can be collected every six months. Clinical endpoint events include recorded all-cause mortality, cardiovascular mortality, myocardial infarction, heart failure, and acute stroke, etc. Moreover, the method of this invention is particularly suitable for elderly individuals aged 65 to 85 years. The above are merely optional examples and should not be construed as limiting the invention.
[0031] Next, in step S102, the collected historical medical data is preprocessed, and feature screening is performed according to the specified screening rules. The feature screening includes: using univariate analysis to calculate each feature for the first screening, and using the Spellman correlation method to calculate the feature correlation for the second screening.
[0032] Specifically, the collected historical medical data underwent data preprocessing. CYP2D6 activity scores were calculated according to the latest consensus and converted to CYP2D6 metabolic phenotypes. For example, Sanger sequencing was used to determine the site polymorphisms of CYP2D6*2 (rs16947), *10 (rs1065852), *14 (rs5030865), and ADRB1 (rs1801253), and optimized large-fragment polymerase chain reaction (PCR) was used to detect the site polymorphism of CYP2D6*5 (gene deletion). CYP2D6 activity scores (AS) were calculated according to the latest consensus, and CYP2D6 genotypes were converted to metabolic phenotypes. In this example, the data included 651 cases (62.7%) of normal metabolites (NMs), 385 cases (37.1%) of intermediate metabolites (IMs), and 3 cases (0.3%) of poorly metabolites (PMs). PMs were excluded in subsequent data analysis due to their small number. All allele frequencies were in Hardy-Weinberg equilibrium. For example, high-performance liquid chromatography-mass spectrometry (HPLC-MS / MS) was used to detect the stable plasma trough concentration of metoprolol.
[0033] Furthermore, the "activity score" is transformed into a categorical variable. Specifically, an activity score ≤ 1 corresponds to a category value of 0, indicating intermediate metabolites (IMs). An activity score > 1 corresponds to a category value of 1, indicating normal metabolites (NMs). This category will be named "Activity Score Classification".
[0034] Convert "metoprolol maintenance dose" into a categorical variable. For example, use one-hot encoding to convert the feature. Specifically, if the metoprolol maintenance dose is ≤25mg / d, the corresponding category value is 1, named "Maintenance Dose_Category 1"; if the metoprolol maintenance dose is >25mg / d, the corresponding category value is 2, named "Maintenance Dose_Category 2" (which is classified according to the metoprolol dose, and can be considered as low dose or normal dose), or use 50mg / d to define the category label.
[0035] The following expressions are used to exponentially convert “brain natriuretic peptide (BNP), cardiac troponin I (cTnI), and cardiac troponin T (cTnT)”.
[0036]
[0037] in, This indicates that the data for "brain natriuretic peptide, troponin I, and troponin T" are converted into exponential form. i It is a positive integer.
[0038] Next, feature filtering is performed on all preprocessed medical data according to the specified filtering rules.
[0039] Univariate analysis was used to calculate the p-value of the significance between each feature and adverse reaction. The calculated p-value was compared with the first specific value for the first screening.
[0040] After the first screening, the Spellman correlation method is used to calculate the feature correlation. The absolute value of the correlation between the calculated features is compared with a second specific value for the second screening.
[0041] Specifically, the specified rules include: deleting features whose calculated P-value is greater than a first specific value; and deleting selected features when the absolute value of the correlation between the selected features is greater than a second specific value.
[0042] Optionally, the first specific value is 0.05 to 0.2, preferably 0.1. The second specific value is 0.6 to 0.8, preferably 0.7.
[0043] In this example, 50% of the features passed the above screening process. The specific P-values are shown in Table 1 below.
[0044]
[0045] Specifically, in the second screening, the Spellman correlation is used to calculate the correlation between the features selected in the first screening and "whether there is an adverse reaction". Features with an absolute correlation value greater than or equal to 0.1 with "whether there is an adverse reaction" are selected as candidate features. At the same time, it is ensured that there is no high correlation between the selected features. Features with a correlation coefficient greater than 0.7 (e.g., approximately 30% of the features are retained) are excluded to avoid redundant information from multiple highly correlated features. Therefore, the above screening process helps to build a more reliable and interpretable model, simplifies the model, removes redundant features, and improves interpretability.
[0046] Next, the candidate features obtained after the first and second screenings are input into the random forest model to determine the feature combination with the optimal AUC value, which yields the following nine features: activity score, age, brain natriuretic peptide, cerebrovascular disease, liver disease, troponin I, heart rate, diastolic blood pressure, and atrial fibrillation.
[0047] Specifically, the random forest model is trained using Gini parameters, the AUC value is calculated, and the importance of different features is evaluated (specifically, the features are sorted in descending order).
[0048] In one specific implementation, the hyperparameters of the random forest model are "n_estimators=100, max_depth=None, random_state=42, criterion='gini'". Subsequently, different numbers of features are progressively selected and input into the random forest model. The model's AUC value is recalculated. By finding the feature combination with the optimal AUC value, the feature combination with the best predictive performance in the feature selection process is determined, which can reduce feature dimensionality and improve model efficiency. Through the above selection process, nine features were selected for modeling: "activity score", "age", "brain natriuretic peptide", "cerebrovascular disease", "liver disease", "troponin I", "heart rate", "diastolic blood pressure", and "atrial fibrillation".
[0049] It should be noted that the above is only an optional example and should not be construed as a limitation of the present invention.
[0050] Next, in step S103, based on the selected features, positive samples and negative samples are defined, and a training dataset is established. Specifically, before establishing the training dataset, a small number of samples are oversampled to make the number of positive samples and negative samples equal.
[0051] Specifically, the nine identified characteristics were divided into continuous and discrete variables. The continuous variables included "age," "brain natriuretic peptide," "troponin I," "heart rate," and "diastolic blood pressure." The discrete variables included "cerebrovascular disease," "liver disease," and "atrial fibrillation."
[0052] Oversampling of a small number of samples, such as Figure 3 As shown, the specific steps include:
[0053] Step S301: Divide the nine identified features into continuous variables and discrete variables.
[0054] Step S302: Use the SMOTE method to synthesize negative samples of all continuous variables.
[0055] Step S303: For each synthesized negative sample, find its k nearest neighbor negative samples for each discrete variable, and calculate the Hamming distance of these negative samples on the class features. Select the nearest neighbor sample with the smallest Hamming distance as a reference, and synthesize the class features of the new sample by randomly selecting a class value different from the reference sample.
[0056] Step S304: Combine the type characteristics of the synthesized discrete variables with the values of the continuous variables to obtain a complete synthesized sample.
[0057] For example, the SMOTE method can be used to synthesize negative samples of all continuous variables.
[0058] Use the K-nearest neighbor algorithm (e.g., k=5) to compute each minority class sample The Euclidean distance between the sample point and its k neighbors is calculated, and one neighbor (i.e., the nearest neighbor sample) is randomly selected. A random interpolation is then performed on the line segment between this sample point and its neighbor, with the interpolation position determined by a random number between 0 and 1. The calculation formula is as follows:
[0059]
[0060] Where, x new These are newly synthesized sample points, x i It is the original minority class sample, x j It is x i The neighbors. rand(0,1) represents a random number between 0 and 1, such as 0.6, 0.7, 0.8, preferably 0.6.
[0061] For each synthesized negative sample x new For each discrete variable, find its k nearest neighbor negative samples 𝑁 new (x newAnd calculate the Hamming distance H of these negative samples on the class features. d Select the nearest neighbor sample X with the minimum Hamming distance. ref As a reference, the class features of the new sample are synthesized by randomly selecting a class value different from that of the reference sample. The specific formula is as follows:
[0062]
[0063] in, and These are samples and The value of the j-th discrete variable; It is the number of possible categories for the d-th discrete variable; This indicates a bitwise XOR operation.
[0064] By combining the type characteristics of the synthesized discrete variables with the values of the continuous variables, a complete synthetic sample is obtained. This synthetic sample consists of nine features used in the modeling.
[0065] Okay, let's use a concrete example to illustrate how to combine the type features of the synthesized discrete variables with the values of continuous variables to obtain a complete synthetic sample. This synthetic sample will contain nine features for modeling.
[0066] For example, there are 9 features, of which the first 5 are continuous variables and the last 4 are discrete variables: 1. Activity score (continuous), 2. Age (continuous), 3. Brain natriuretic peptide (continuous), 4. Troponin I (continuous), 5. Heart rate (continuous), 6. Cerebrovascular disease (discrete), 7. Liver disease (discrete), 8. Troponin T (discrete), 9. Atrial fibrillation (discrete).
[0067] A new sample is synthesized, where the values of continuous variables have already been generated using the SMOTE algorithm, and the values of discrete variables need to be synthesized through the following steps:
[0068] Step 1: Synthesis of continuous variable values
[0069] Suppose that the following continuous variable values were synthesized for a negative sample (a patient who did not experience an adverse reaction) using the SMOTE algorithm:
[0070] - Activity score: 0.5
[0071] - Age: 68
[0072] - Brain natriuretic peptide: 300
[0073] - Troponin I: 0.2
[0074] Heart rate: 72
[0075] Step 2: Synthesis of discrete variable values
[0076] For discrete variables, we find the positive samples (patients experiencing adverse reactions) that are most similar to the original negative samples, and synthesize new category values according to certain rules. Suppose we have the following discrete variable values for positive samples:
[0077] - Cerebrovascular disease: Yes
[0078] - Liver disease: No
[0079] - Troponin T: is
[0080] - Atrial fibrillation: No
[0081] The K-Nearest Neighbors algorithm was used to find the positive samples most similar to the synthesized continuous variable values, and the Hamming distance was calculated. The following positive samples were selected as references:
[0082] - Cerebrovascular disease: No
[0083] - Liver disease: Yes
[0084] - Troponin T: No
[0085] - Atrial fibrillation: Yes
[0086] Step 3: Combine continuous and discrete variables
[0087] By combining the synthesized continuous variable values with the discrete variable values, a complete synthetic sample is obtained. In this example, it might be decided to change the states of "cerebrovascular disease" and "atrial fibrillation" to increase sample diversity. Therefore, the discrete variable values of the synthetic sample might be as follows:
[0088] - Cerebrovascular disease: Yes (change from "No" to "Yes")
[0089] - Liver disease: Yes (remains unchanged)
[0090] - Troponin T: Yes (changes from "No" to "Yes")
[0091] - Atrial fibrillation: Yes (changes from "No" to "Yes")
[0092] Complete synthetic sample
[0093] By combining continuous and discrete variables, the following complete synthetic sample is obtained:
[0094] - Activity score: 0.5
[0095] - Age: 68
[0096] - Brain natriuretic peptide: 300
[0097] - Troponin I: 0.2
[0098] Heart rate: 72
[0099] - Cerebrovascular disease: Yes
[0100] - Liver disease: Yes
[0101] - Troponin T: is
[0102] - Atrial fibrillation: Yes
[0103] This synthetic sample can now be added to the training dataset to train the risk prediction model. In this way, not only is the number of positive samples increased, but the model's ability to identify different patient characteristics is also improved by introducing different combinations of discrete variables.
[0104] Based on the selected features, the above oversampling steps (specifically, oversampling a small number of samples to make the number of positive and negative samples equal) are used to sample the ratio of positive to negative samples to 1:1 to establish a training dataset.
[0105] Therefore, by oversampling a small number of samples as described above, the imbalance of data can be effectively improved, the model performance can be enhanced, and the generalization ability of the model can be strengthened.
[0106] Next, in step S104, a risk prediction model is established. The risk prediction model is trained using the training dataset to obtain the trained risk prediction model.
[0107] For example, a risk prediction model can be established using the Naive Bayes method. The risk prediction model can then be trained using the training set established in step S103.
[0108] Optionally, the dataset is randomly split into training and test sets in a ratio of 8:2 (training:test). After sampling the training set using S103, risk prediction models are built using eight different machine learning models and nodal plots. The test set is then used to evaluate the model performance to determine which Naive Bayes method to use for building the risk prediction model.
[0109] The steps above mentioned using eight different machine learning models to build a risk prediction model and evaluating its performance using a test set. To enhance training, this process can be expanded to include adding more algorithmic models and conducting more detailed model evaluation and selection. The following are the expanded and supplementary steps:
[0110] 1. Dataset Splitting: First, randomly split the dataset into 80% training set and 20% test set. This ensures sufficient data for training the model while providing a separate dataset for evaluating its generalization ability.
[0111] 2. Feature Engineering: Before training the model, feature engineering is performed, including feature selection, feature extraction, and feature transformation. This may include standardizing continuous variables, encoding discrete variables, handling missing values, and transforming nonlinear features.
[0112] 3. Model Training: Train various machine learning models using the training set data. Besides Naive Bayes, the following models can also be considered:
[0113] - Logistic Regression
[0114] Support Vector Machine (SVM)
[0115] - Decision Tree
[0116] Random Forest
[0117] Gradient Boosting Machine (GBM)
[0118] - Extra Trees
[0119] - Artificial Neural Networks
[0120] 4. Model Tuning: Tune each model, including adjusting hyperparameters and selecting the optimal model configuration. This can be achieved using methods such as grid search or random search.
[0121] 5. Cross-validation: To more accurately evaluate model performance, cross-validation (such as k-fold cross-validation) is used to assess the stability and accuracy of each model.
[0122] 6. Model Evaluation: Each model is evaluated using a test set. Evaluation metrics may include accuracy, sensitivity, specificity, precision, F1 score, AUC, etc.
[0123] 7. Model Comparison: Compare the performance of different models and select the best-performing model. If a model performs well across multiple evaluation metrics, it is likely the optimal choice.
[0124] 8. Model Fusion: Consider using model fusion techniques, such as stacking or voting, to combine the predictions of multiple models to improve overall performance.
[0125] 9. Model Interpretation: For the selected best model, use model interpretation tools (such as SHAP, LIME) to interpret the model's decision-making process, which is especially important for models in the medical field.
[0126] 10. Model Deployment: Deploy the trained model to the production environment to perform risk assessments on new patient data.
[0127] 11. Continuous monitoring: After the model is deployed, continuously monitor its performance to ensure that it remains accurate and reliable in practical applications.
[0128] These extended and supplementary steps enable a more comprehensive evaluation and selection of the best machine learning models, and enhance model performance and robustness through augmented training.
[0129] In one alternative implementation, the above nine features are divided into two or more categories of labeled data, and the risk prediction model is enhanced and trained to obtain a well-trained risk prediction model.
[0130] In another alternative implementation, the above nine features are divided into data with multiple labels, and the risk prediction model is trained or enhanced to obtain a trained risk prediction model.
[0131] Next, in step S105, the drug dosage and test data of the object to be evaluated are input into the trained risk prediction model to obtain the risk assessment results.
[0132] When applying a risk prediction model, the drug dosage and test data of the object to be evaluated are input into the trained risk prediction model, which outputs a risk assessment value to determine the likelihood of adverse reactions or other risks. The risk assessment value is, for example, the probability of a risk occurring, and further, based on this probability, a judgment is made regarding the occurrence of adverse reactions or other risks.
[0133] It should be noted that the above is only an optional example and should not be construed as a limitation of the present invention.
[0134] To verify the effectiveness of the present invention, a comparative experiment was conducted, and the comparative results are shown in Table 1 below.
[0135]
[0136] Through the above comparative experiments, the method of the present invention was compared with the existing methods such as 'K-nearest neighbor', 'support vector machine', 'logistic regression', 'decision tree', 'AdaBoost', 'XGBoost', and 'random forest' to calculate the risk assessment, and the evaluation indicators in Table 1 above were obtained.
[0137] Specifically, a confusion matrix can be calculated based on the actual results and the predicted results, and then various evaluation indicators can be calculated, as follows:
[0138] Confusion Matrix
[0139]
[0140]
[0141] Here, AUROC refers to the area under the curve between sensitivity and (1-specificity). AUPRC refers to the area under the curve between precision and sensitivity.
[0142] As can be seen from Table 1, the method of the present invention is superior to other existing methods in all evaluation indicators.
[0143] Furthermore, the accompanying drawings are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes shown in the drawings do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0144] Compared with existing technologies, this invention analyzes the polymorphisms of the CYP2D6 and ADRB1 genes, records and collects relevant medical data on metoprolol tolerance, adverse reactions, blood pressure and heart rate responses, and long-term prognosis. The collected medical data is then processed to select multiple features to establish a training dataset and a risk prediction model. This model is trained using the training dataset to obtain a well-trained risk prediction model. The drug dosage and test data of the subjects to be evaluated are input into the trained risk prediction model to obtain risk assessment results. This provides a theoretical basis for establishing gene polymorphism-guided personalized metoprolol medication regimens.
[0145] In addition, univariate analysis was used to calculate the features for the first screening, and the Spellman correlation method was used to calculate the feature correlation for the second screening. The two screening processes can help build a more reliable and interpretable model, simplify the model, remove redundant features, and improve the interpretability of features.
[0146] The following are system embodiments of the present invention, which can be used to execute the method embodiments of the present invention. For details not disclosed in the system embodiments of the present invention, please refer to the method embodiments of the present invention.
[0147] Figure 4 This is a schematic diagram of an example of a drug evaluation system according to the present invention.
[0148] Reference Figure 4 The second aspect of this disclosure provides a drug evaluation system 400, which employs the drug evaluation method described in the first aspect of the invention. The drug evaluation system includes a data collection module 410, a feature screening module 420, a first establishment module 430, a second establishment module 440, and a result output module 450.
[0149] Specifically, the data collection module 410 collects historical medical data, including disease data, drug dosage, physiological indicator data, and test data. The feature selection module 420 preprocesses the collected historical medical data and performs feature selection according to specified rules. The feature selection includes: calculating the features using univariate analysis for the first selection and calculating the feature correlation using the Spellman correlation method for the second selection. The first establishment module 430 defines positive and negative samples based on the selected features and establishes a training dataset. Specifically, before establishing the training dataset, a small number of samples are oversampled to ensure that the number of positive and negative samples is equal. The second establishment module 440 establishes a risk prediction model by training the risk prediction model using the training dataset. The result output module 450 inputs the drug dosage and test data of the object to be evaluated into the trained risk prediction model to obtain the risk assessment result.
[0150] Compared with existing technologies, this invention analyzes the polymorphisms of the CYP2D6 and ADRB1 genes, records and collects relevant medical data on metoprolol tolerance, adverse reactions, blood pressure and heart rate responses, and long-term prognosis. The collected medical data is then processed to select multiple features to establish a training dataset and a risk prediction model. This model is trained using the training dataset to obtain a well-trained risk prediction model. The drug dosage and test data of the subjects to be evaluated are input into the trained risk prediction model to obtain risk assessment results. This provides a theoretical basis for establishing gene polymorphism-guided personalized metoprolol medication regimens.
[0151] Figure 5 This is a schematic diagram of an embodiment of an electronic device according to the present invention.
[0152] like Figure 5 As shown, the electronic device is embodied in the form of a general-purpose computing device. There can be one or more processors working collaboratively. This invention also does not preclude distributed processing, meaning that processors can be distributed across different physical devices. The electronic device of this invention is not limited to a single entity, but can also be the sum of multiple physical devices.
[0153] The memory stores a computer-executable program, typically machine-readable code. The computer-readable program can be executed by the processor to enable the electronic device to perform the method of the present invention, or at least some steps of the method.
[0154] The memory includes volatile memory, such as random access memory (RAM) and / or cache memory, and may also be non-volatile memory, such as read-only memory (ROM).
[0155] Optionally, in this embodiment, the electronic device further includes an I / O interface for exchanging data with external devices. The I / O interface can represent one or more of several bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0156] It should be understood that Figure 5 The electronic device shown is merely one example of the present invention, and the electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include display units such as displays, and some electronic devices also include human-computer interaction elements such as buttons and keyboards. Any electronic device capable of executing a computer-readable program in memory to implement the method of the present invention or at least some steps of the method can be considered as an electronic device covered by the present invention.
[0157] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software, or by combining software with necessary hardware. Therefore, as... Figure 6 As shown, the technical solution according to the embodiments of the present invention can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.) or on a network, and includes several commands to cause a computing device (such as a personal computer, server, or network device, etc.) to execute the above-described method according to the embodiments of the present invention.
[0158] The software product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0159] The computer-readable storage medium may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, capable of transmitting, propagating, or transmitting programs for use by or in connection with a command execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0160] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0161] The aforementioned computer-readable medium carries one or more programs (e.g., computer executable programs) that, when executed by a device, cause the computer-readable medium to implement the methods of this disclosure.
[0162] Those skilled in the art will understand that the above modules can be distributed in the device as described in the embodiments, or they can be modified accordingly and placed in one or more devices that are unique to this embodiment. The modules in the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.
[0163] Through the description of the above embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions of the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, portable hard drive, etc.) or on a network, including several commands to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of the present invention.
[0164] Exemplary embodiments of the present invention have been specifically shown and described above. It should be understood that the present invention is not limited to the detailed structures, arrangements, or implementations described herein; rather, the present invention is intended to cover various modifications and equivalent arrangements contained within the spirit and scope of the appended claims.
Claims
1. A drug risk assessment method based on metoprolol personalized medication regimens, characterized in that, include: Historical medical data on metoprolol use within a specified historical period is collected. This historical medical data includes patient condition data, drug dosage, physiological indicators, and laboratory data. Data is collected at equal intervals for a specific duration, including collecting blood pressure, heart rate, metoprolol maintenance dose, and related adverse reactions every two weeks. The related adverse reactions specifically include cardiovascular and non-cardiovascular adverse reactions. If metoprolol was not discontinued at the last data collection point, the metoprolol dosage is defined as the maintenance dose. Laboratory data includes whether a clinical endpoint event occurred and the date of occurrence. The collected historical medical data underwent data preprocessing, which included calculating the CYP2D6 activity score according to the latest consensus and converting it into a CYP2D6 metabolic phenotype; converting the "activity score" into a categorical variable; converting the "maintenance dose" into a categorical variable; and performing feature screening according to specified screening rules. Feature screening included: calculating the significance of each feature using univariate analysis for the first screening; calculating the feature correlation using the Spellman correlation method for the second screening; calculating the p-value of the significance between each feature and adverse reactions using univariate analysis; and comparing the calculated p-value with a first specific value for the first screening. The specified rules included: removing features with p-values greater than the first specific value, where the first specific value is 0.05–0.
2. The candidate features obtained after the first and second screenings were input into a random forest model to determine the feature combination with the optimal AUC value, resulting in the following nine features: activity score, age, brain natriuretic peptide, cerebrovascular disease, liver disease, troponin I, heart rate, diastolic blood pressure, and atrial fibrillation. Based on the selected features, positive and negative samples are defined, and a training dataset is established. Specifically, before establishing the training dataset, a small number of samples are oversampled to make the number of positive and negative samples equal. Establish a risk prediction model, train the risk prediction model using a training dataset, and obtain a well-trained risk prediction model. The drug dosage and test data of the subjects to be evaluated are input into the trained risk prediction model to obtain the risk assessment results, which will guide the individualized metoprolol medication regimen.
2. The drug risk assessment method based on metoprolol personalized medication regimens according to claim 1, characterized in that, After the first screening, the Spellman correlation method is used to calculate the feature correlation. The absolute value of the correlation between the calculated features is compared with a second specific value for the second screening. The specified filtering rules include: when the absolute value of the correlation between the selected features is greater than a second specific value, the selected features will be deleted.
3. The drug risk assessment method based on metoprolol personalized medication regimens according to claim 2, characterized in that, The first specific value is 0.1, and the second specific value is 0.
7.
4. The drug risk assessment method based on metoprolol personalized medication regimens according to claim 1, characterized in that, The nine identified features are divided into continuous variables and discrete variables; Use the SMOTE method to synthesize negative samples of all continuous variables; For each synthesized negative sample, find its k nearest neighbor negative samples for each discrete variable and calculate the Hamming distance of these negative samples on the class features. Select the nearest neighbor sample with the smallest Hamming distance as a reference and synthesize the class features of the new sample by randomly selecting a class value different from the reference sample. By combining the type characteristics of the synthesized discrete variables with the values of the continuous variables, a complete synthesized sample is obtained.
5. The drug risk assessment method based on metoprolol personalized medication regimens according to claim 1, characterized in that, The data preprocessing of the collected historical medical data includes: Transform the "activity score" into a categorical variable; Transform "maintenance dose" into a categorical variable; The following expression is used to exponentially convert "brain natriuretic peptide, troponin I, and troponin T"; x i =ln(x i +1e -10 ) Where, x i This indicates that "brain natriuretic peptide, troponin I, troponin T" are converted into exponential form, where i is a positive integer.
6. The drug risk assessment method based on metoprolol personalized medication regimens according to claim 1, characterized in that, include: When applying the risk prediction model, the drug dosage and test data of the object to be evaluated are input into the trained risk prediction model, and the risk assessment value is output to determine the adverse reaction situation or risk situation.
7. A drug risk assessment system based on metoprolol personalized medication regimens, employing the drug risk assessment method based on gene polymorphism as described in claim 1, characterized in that, include: The data collection module collects historical medical data on metoprolol use within a specified historical period. This historical medical data includes patient condition data, drug dosage, physiological indicator data, and laboratory data. Data is collected at equal intervals for a specific duration, including collecting blood pressure, heart rate, metoprolol maintenance dose, and related adverse reactions every two weeks. The related adverse reactions specifically include cardiovascular and non-cardiovascular adverse reactions. If metoprolol was not discontinued at the last data collection point, the metoprolol dosage is defined as the maintenance dose. Laboratory data includes whether a clinical endpoint event occurred and the date of occurrence. Preprocessing includes calculating the CYP2D6 activity score according to the latest consensus and converting it to a CYP2D6 metabolic phenotype; converting the "activity score" and "maintenance dose" into categorical variables. The feature screening module preprocesses the collected historical medical data. This preprocessing includes calculating the CYP2D6 activity score based on the latest consensus and converting it into a CYP2D6 metabolic phenotype; converting the "activity score" into a categorical variable; converting the "maintenance dose" into a categorical variable; and performing feature screening according to specified screening rules. The feature screening includes: calculating the correlation between each feature using univariate analysis for the first screening; and calculating the correlation between features using the Spellman correlation method for the second screening. The first module establishes a training dataset based on multiple selected features, defining positive and negative samples. Specifically, this includes: before establishing the training dataset, oversampling a small number of samples to ensure an equal number of positive and negative samples; using univariate analysis to calculate the p-value of the significance between each feature and adverse reactions; comparing the calculated p-value with a first specific value for initial screening; the specified rule includes: removing features with p-values greater than the first specific value (0.05–0.2); and inputting the candidate features obtained after the first and second screenings into a random forest model to determine the feature combination with the optimal AUC value, resulting in the following nine features: activity score, age, brain natriuretic peptide, cerebrovascular disease, liver disease, troponin I, heart rate, diastolic blood pressure, and atrial fibrillation. The second module establishes a risk prediction model. The risk prediction model is trained using a training dataset to obtain a well-trained risk prediction model. The results output module takes the drug dosage and test data of the object to be evaluated and inputs them into the trained risk prediction model to obtain the risk assessment results.
Citation Information
Patent Citations
Early prediction and evaluation method for predicting curative effect of non-small cell lung cancer immunotherapy
CN116312807A
Untoward drug reaction monitoring and early warning method
CN118280603A