Method for improving drug-adverse reaction signal intensity based on machine learning
By applying machine learning methods in the FAERS database, processing data and building deep learning models, the problem of signal strength accuracy of adverse drug reactions in the database is solved, and more accurate signal strength calculation is achieved, and more accurate clinical decision-making is supported.
Patent Information
- Application Number
- CN202510103524.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
AI Technical Summary
Due to the spontaneous source of data in the FAERS database, there are errors and misreports, which lead to the accuracy of the signal strength of adverse drug reactions, which may mislead doctors' clinical decisions.
Machine learning methods are used to process data in the FAERS database, and deep learning models are built to improve the accuracy of signal strength of adverse drug reactions through steps such as data preprocessing, data balance, model construction and model training.
Through the application of machine learning algorithms, the impact of misreport and underreport is reduced, and the accuracy of signal strength of adverse drug reactions is improved, thereby supporting more accurate clinical decision-making.
Smart Images

Figure CN119939254A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of drug adverse reaction monitoring, and in particular to a method and system for improving the accuracy of drug adverse reaction signal strength in the U.S. Food and Drug Administration (FDA) Adverse Event Reporting System (FAERS) database by using a machine learning algorithm. Background Art
[0002] As one of the world's largest adverse drug reaction monitoring databases, the FAERS database provides valuable information resources for drug safety monitoring and clinical drug use decisions. However, due to the spontaneity of its data source, the FAERS database has errors and omissions, resulting in certain accuracy problems in the signal strength of adverse drug reactions calculated based on the database (such as report odds ratio (ROR) and proportion report ratio (PRR), which may mislead doctors' clinical decisions. Although traditional statistical methods can improve the calculation or expression of signal strength to a certain extent, they cannot fundamentally solve the problems of misreporting and omissions. Therefore, the present invention aims to provide a new method for improving the accuracy of adverse drug reaction signal strength in the FAERS database using a machine learning algorithm. Summary of the invention
[0003] The present invention uses a machine learning method to process the data in the FEARS database to reduce the impact of false positives and false negatives, thereby improving the accuracy of drug-adverse reaction signal strength:
[0004] Data preprocessing stage: The data is upgraded to be input into the machine learning model, and the data is standardized by subtracting the mean of each eigenvalue and dividing it by its standard deviation to eliminate the impact of different eigenvalue dimensions.
[0005] Data balancing stage: A five-fold cross-validation method is used to divide the data set into a training set and a test set, and the data set is balanced by randomly copying data of fewer categories. The balanced training set is randomly disrupted again to ensure the randomness of the data.
[0006] Model building phase: Build a deep learning model, including convolutional layers, pooling layers, LSTM layers, Dropout layers, and fully connected layers. Use the PaddlePaddle framework to build and train the model.
[0007] Model training phase: Define training parameters, including learning rate, batch size, number of training rounds, etc. Use Adam optimizer and cross entropy loss function for model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 Layered framework diagram. DETAILED DESCRIPTION
[0009] (1) Data preprocessing: Extract data from the FAERS database, including fields such as drug name, adverse reaction terminology, and report date; deduplicate the data and remove duplicate report records; standardize drug names and adverse reaction terms to ensure data consistency and standardization; handle missing values and fill in or delete missing data fields.
[0010] (2) Construct and train machine learning algorithms: Using a combination of deep learning convolutional neural networks (CNN) and long short-term memory networks (LSTM) to process data in the FAERS database can effectively extract local features and time series features of the data, thereby screening out poisoned data. Use cross-validation methods to train the algorithm to avoid overfitting. Adjust model parameters, optimize model performance, and determine the optimal model parameters. (3) Model application: Use the trained model to process data in the FAERS database and identify deviations in the data. Calculate the drug adverse reaction signal strength of the processed data: report odds ratio (ROR) and proportion report ratio (PRR), and compare them with the original signal strength.
[0011] (4) Evaluation and optimization: Evaluate the performance of the model to ensure its accuracy and stability when processing large-scale data sets. Further optimize the model based on the evaluation results to improve the accuracy of signal strength. Update the model regularly to adapt to new data and changing medical environments.
Claims
1. The purpose of the present invention is to provide a method based on a machine learning algorithm to improve the accuracy of the signal strength of adverse drug reactions in the U.S. Food and Drug Administration (FDA) Adverse Event Reporting System (FAERS) database. The specific process of the scheme is as follows: (1) Data preprocessing: Extract data from the FAERS database, including fields such as drug name, adverse reaction term, and report date, deduplicate the data, remove duplicate report records, standardize drug names and adverse reaction terms to ensure data consistency and standardization, handle missing values, and fill or delete missing data fields; (2) Construct and train machine learning algorithms: A combination of convolutional neural networks (CNNs) and long short-term memory networks (LSTMs) in deep learning was used to process data in the FAERS database, which can effectively extract local features and time series features of the data, thereby screening out poisoned data. The algorithm was trained using a cross-validation method to avoid overfitting, adjust model parameters, optimize model performance, and determine the optimal model parameters. (3) Model application: Use the trained model to process the data in the FAERS database, identify deviations in the data, and calculate the adverse drug reaction signal strength of the processed data: report odds ratio (ROR) and proportion reporting ratio (PRR), and compare them with the original signal strength; (4) Evaluation and optimization: Evaluate the performance of the model to ensure its accuracy and stability when processing large-scale data sets. Further optimize the model based on the evaluation results to improve the accuracy of signal strength. Regularly update the model to adapt to new data and changing medical environments.