Adverse drug reaction monitoring method and system based on big data
Through the feature extraction and fusion processing of drug clinical feedback big data, the clinical feedback characteristics of drug are generated, which solves the shortcomings of traditional drug adverse reaction monitoring methods in long-term use and in special populations, and accurately predicts and evaluates drug adverse reactions, improving the accuracy and efficiency of monitoring.
Patent Information
- Application Number
- CN202510025811.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-08
AI Technical Summary
Traditional drug adverse reaction monitoring methods have problems such as long cycles, limited sample size, and limited drug use in drug research and development and clinical applications, and it is difficult to fully reveal the adverse reactions of drugs in long-term use, special populations or drug use.
By obtaining clinical feedback big data of multiple target drugs, extracting drug action characteristics data and patient response characteristics data, and merging them to generate drug clinical feedback characteristics, grouping and labeling based on these characteristics to achieve accurate prediction and evaluation of drug adverse reactions.
It significantly improves the accuracy and efficiency of monitoring of adverse drug reactions, reduces the risk of adverse reactions during drug use, and ensures the safety of patients' medication.
Smart Images

Figure CN119418951B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of big data technology, and more specifically, to a method and system for monitoring adverse drug reactions based on big data. Background Art
[0002] In the process of drug development and clinical application, the monitoring and management of adverse drug reactions is a key link to ensure patient safety and improve drug efficacy. Traditionally, the identification of adverse drug reactions mainly relies on animal experiments and Phase I, II, and III clinical trials. Although these processes can collect certain adverse reaction event information, they are limited by factors such as long trial cycles, limited sample sizes, and restrictions on combined medications. It is often difficult to fully reveal the adverse reactions of drugs in long-term use, special populations, or combined medications.
[0003] Especially in public emergencies, it is particularly important to quickly and accurately identify adverse drug reactions, and traditional methods are incapable of doing so in such situations. In addition, the strict inclusion and exclusion criteria of clinical trials may result in some adverse events not being fully exposed during the trial phase, further increasing the difficulty of monitoring adverse reactions after drug marketing. Summary of the invention
[0004] In view of the above-mentioned problems, in combination with the first aspect of the present application, the present application embodiment provides a drug adverse reaction monitoring method based on big data, the method comprising:
[0005] Obtaining big data on clinical feedback generated from drug clinical trials for multiple target drugs;
[0006] Extracting drug action characteristic data and patient response characteristic data from the clinical feedback big data of each of the target drugs;
[0007] The drug action characteristic data and the patient response characteristic data corresponding to the same target drug are respectively integrated to generate drug clinical feedback characteristics of each target drug;
[0008] Searching for a combination drug clinical feedback feature that meets the drug combination application condition with the target drug clinical feedback feature among the drug clinical feedback features of each of the target drugs, and grouping the target drug clinical feedback feature with the combination drug clinical feedback feature that meets the drug combination application condition to generate a feature group;
[0009] The adverse drug reaction labels corresponding to the target drugs are determined based on the feature grouping.
[0010] In a possible implementation of the first aspect, the drug action characteristic data and the patient response characteristic data are obtained using an autoencoder; and before obtaining the clinical feedback big data generated by conducting drug clinical trials on multiple target drugs, the method further includes:
[0011] Obtain clinical feedback big data of each template generated by conducting drug clinical trials on multiple template drugs;
[0012] Extracting the first sample patient response characteristic data and the first sample drug action characteristic data of each of the template clinical feedback big data using a reference autoencoder, and blending the first sample patient response characteristic data and the first sample drug action characteristic data of each of the template clinical feedback big data to generate each first sample drug clinical feedback feature;
[0013] Searching for a combination sample drug clinical feedback feature that meets the drug combination application condition with the target sample drug clinical feedback feature among the first sample drug clinical feedback features;
[0014] Determining a combined error parameter based on the first label heat map of the target sample drug clinical feedback feature and the first label heat map of the combined sample drug clinical feedback feature;
[0015] Determining a balanced distribution error parameter based on a first label heat map of the clinical feedback characteristics of the target sample drug;
[0016] Based on the combined error parameter and the balanced distribution error parameter, the parameter information of the reference autoencoder is updated to generate the autoencoder.
[0017] In a possible implementation of the first aspect, after obtaining the clinical feedback big data of each template generated by conducting drug clinical trials on multiple template drugs, the method further includes:
[0018] Performing feature expansion on the template clinical feedback big data to generate each expanded clinical feedback big data of the template clinical feedback big data;
[0019] Extracting the second sample drug clinical feedback features of the template clinical feedback big data and the first extended sample drug clinical feedback features of each of the extended clinical feedback big data using a pre-trained autoencoder;
[0020] Determining a drug action matching error parameter based on the drug action matching between the second sample drug clinical feedback feature of the template clinical feedback big data and each of the first extended sample drug clinical feedback features;
[0021] Updating parameter information of the pre-trained autoencoder based on the drug action matching error parameter to generate the reference autoencoder;
[0022] In a possible implementation of the first aspect, the pre-trained autoencoder includes a pre-trained patient response encoder and a pre-trained drug action encoder, and the template clinical feedback big data includes template drug feedback data and template drug action state data;
[0023] The method of extracting the second sample drug clinical feedback features of the template clinical feedback big data and the first extended sample drug clinical feedback features of each of the extended clinical feedback big data by using the pre-trained autoencoder includes:
[0024] Using the pre-trained patient response encoder, respectively encode the template drug feedback data and each of the extended clinical feedback big data to generate sample drug feedback feature data and each of the sample extended feedback feature data;
[0025] extracting sample drug action characteristic data of the template drug action state data using the pre-trained drug action encoder;
[0026] The sample drug action characteristic data are respectively integrated with the sample drug feedback characteristic data and each of the sample extended feedback characteristic data to generate a second sample drug clinical feedback characteristic of the template clinical feedback big data and a first extended sample drug clinical feedback characteristic of each of the extended clinical feedback big data.
[0027] For example, in a possible implementation of the first aspect, after updating the parameter information of the reference autoencoder based on the combined error parameter and the balanced distribution error parameter to generate the autoencoder, the method further includes:
[0028] Performing feature expansion on each of the template clinical feedback big data to generate a plurality of expanded clinical feedback big data of each of the template clinical feedback big data;
[0029] Extracting the third sample drug clinical feedback feature of each of the template clinical feedback big data and the second extended sample drug clinical feedback feature of each of the extended clinical feedback big data by using the autoencoder;
[0030] Determine a second label heat map of the clinical feedback characteristics of each of the third sample drugs, and a second label heat map of the clinical feedback characteristics of each of the second extended sample drugs;
[0031] Based on the set threshold value and each of the second label heat maps, determine the fuzzy prediction results corresponding to each of the second label heat maps;
[0032] Polling each of the second label heat maps to generate a target label heat map, determining a first scale of the fuzzy prediction results associated with the target label heat map, and determining a second scale of each of the second label heat maps in the fuzzy prediction results associated with the target label heat map;
[0033] Based on the first scale and the second scale, an update coefficient of the target label heat map is determined; based on the target label heat map and each of the second label heat maps in the fuzzy prediction results associated with the target label heat map, a thermal divergence error parameter of the target label heat map is determined; based on the thermal divergence error parameters and the update coefficient of each of the target label heat maps generated by polling, parameter information of the autoencoder is updated to generate a trained autoencoder.
[0034] On the other hand, an embodiment of the present application also provides a drug adverse reaction monitoring system based on big data, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.
[0035] Based on the above aspects, the embodiment of the present application significantly improves the accuracy and efficiency of adverse drug reaction monitoring by comprehensively collecting and analyzing the clinical feedback big data generated by multiple target drugs in clinical trials. Specifically, by accurately extracting drug action characteristic data and patient response characteristic data from clinical feedback big data, and innovatively blending these characteristic data, a more comprehensive drug clinical feedback feature is generated, which lays a solid foundation for subsequent adverse drug reaction analysis. Further, it is possible to intelligently search and identify feature combinations that meet the application conditions of drug combinations in the huge clinical feedback features, and realize accurate prediction and evaluation of adverse reactions under combined use of drugs. Finally, based on these feature groupings, it is possible to automatically generate adverse drug reaction labels corresponding to target drugs, provide data support for drug development, clinical application and supervision, effectively reduce the risk of adverse reactions during drug use, and ensure the safety of patients' medication. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a schematic diagram of the execution flow of the drug adverse reaction monitoring method based on big data provided in an embodiment of the present application.
[0037] Figure 2 It is a schematic diagram of the hardware architecture of the drug adverse reaction monitoring system based on big data provided in an embodiment of the present application. DETAILED DESCRIPTION
[0038] The present application will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flow chart of a method for monitoring adverse drug reactions based on big data provided by an embodiment of the present application. The method for monitoring adverse drug reactions based on big data is introduced in detail below.
[0039] Step S110, obtaining clinical feedback big data generated by conducting drug clinical trials on multiple target drugs.
[0040] In this embodiment, the server first connects to the data management system of the drug clinical trial to obtain the clinical feedback big data generated by the clinical trials of multiple target drugs. These clinical feedback big data are collected in a strict clinical trial environment to ensure the accuracy and reliability of the data. During the clinical trial, for each target drug, the various reactions of the patient after taking the medicine and the effects of the drug are recorded in detail. These records include changes in the patient's physiological indicators, symptom improvement, the occurrence of any adverse reactions, etc. The server extracts the required clinical feedback big data from these records to prepare for subsequent analysis and processing.
[0041] Specifically, the server will track the clinical trials of each target drug and record the drug feedback events of each patient. These drug feedback events include the time of taking the drug, dosage, immediate reaction after taking the drug, etc. At the same time, the server will also obtain the test data of the test instruments, such as the readings of electrocardiograms, sphygmomanometers and other instruments, which provide objective evidence of the drug's action status. Combining this information, the server generates a large data set of clinical feedback on each target drug.
[0042] Step S120, extracting drug action characteristic data and patient response characteristic data from the clinical feedback big data of each of the target drugs.
[0043] In this embodiment, the server then conducts in-depth mining and analysis of the clinical feedback big data of each target drug. First, the server integrates the first drug action state data and the second drug action state data to generate action state data for each target drug. These data describe the mechanism and effect of the drug in the patient's body. Then, the server uses the drug action encoder to encode and represent these action state data to generate drug action characteristic data. These drug action characteristic data represent the various action characteristics of the drug in a numerical form.
[0044] At the same time, the server also uses the patient response encoder to extract the patient response feature data from the drug feedback data. These patient response feature data describe the various physiological and psychological reactions of patients after taking the drug. By extracting these drug action feature data and patient response feature data, the server provides a basis for the generation of subsequent drug clinical feedback features.
[0045] Step S130, respectively integrating the drug action characteristic data and the patient response characteristic data corresponding to the same target drug to generate drug clinical feedback characteristics of each target drug.
[0046] In this embodiment, after extracting the drug action characteristic data and the patient response characteristic data, the server then blends the drug action characteristic data and the patient response characteristic data. For each target drug, the server determines the drug action characteristic data and the patient response characteristic data extracted based on the same drug feedback event. These data describe the mechanism of action of the drug and the patient's response, respectively. Then, the server blends the drug action characteristic data and the patient response characteristic data to generate drug clinical feedback characteristics corresponding to the drug feedback event in the same target drug. These drug clinical feedback characteristics combine the effects of the drug and the patient's response, providing a basis for subsequent analysis of drug combination applications.
[0047] Step S140, searching for the combination drug clinical feedback characteristics that meet the drug combination application conditions between the target drug clinical feedback characteristics, and grouping the target drug clinical feedback characteristics with the combination drug clinical feedback characteristics that meet the drug combination application conditions to generate a feature grouping.
[0048] In this embodiment, after generating the drug clinical feedback characteristics of each target drug, the server then searches among these drug clinical feedback characteristics. The server polls the clinical feedback characteristics of each target drug and uses them as the target drug clinical feedback characteristics. Then, the server determines the drug action matching degree between the target drug clinical feedback characteristics and the clinical feedback characteristics of other drugs respectively. This drug action matching degree is calculated based on the similarity between the drug action mechanism and the patient's response. By calculating the matching degree, the server can find the combination drug clinical feedback characteristics that meet the drug combination application conditions with the target drug clinical feedback characteristics.
[0049] After finding these combination drug clinical feedback features, the server groups them with the target drug clinical feedback features to generate feature groups. These feature groups represent possible drug combination application scenarios and provide a basis for the subsequent determination of drug adverse reaction labels.
[0050] Step S150, determining the adverse drug reaction label corresponding to each of the target drugs based on the feature grouping.
[0051] In this embodiment, finally, the server determines the adverse drug reaction labels corresponding to each target drug based on the generated feature grouping. First, the server determines the target feature grouping for each target drug based on the scale of the feature grouping. This scale is calculated based on the number of clinical feedback features of the drug in the feature grouping. Then, the server determines the adverse drug reaction label corresponding to the target feature grouping of each target drug. This adverse drug reaction label is determined based on the commonalities and characteristics of the clinical feedback features of the drug in the feature grouping. Finally, the server uses the adverse drug reaction label corresponding to the target feature grouping of each target drug as the adverse drug reaction label corresponding to the target drug. Through these adverse drug reaction labels, doctors and researchers can more accurately understand the adverse reactions that may be caused by each target drug, and thus make more reasonable medication decisions.
[0052] Based on the above steps, the embodiment of the present application significantly improves the accuracy and efficiency of adverse drug reaction monitoring by comprehensively collecting and analyzing the clinical feedback big data generated by multiple target drugs in clinical trials. Specifically, by accurately extracting drug action characteristic data and patient response characteristic data from clinical feedback big data, and innovatively blending these characteristic data, a more comprehensive drug clinical feedback feature is generated, which lays a solid foundation for subsequent adverse drug reaction analysis. Further, it is possible to intelligently search and identify feature combinations that meet the application conditions of drug combinations in the huge clinical feedback features, and realize accurate prediction and evaluation of adverse reactions under combined use of drugs. Finally, based on these feature groupings, it is possible to automatically generate adverse drug reaction labels corresponding to target drugs, provide data support for drug development, clinical application and supervision, effectively reduce the risk of adverse reactions during drug use, and ensure the safety of patients' medication.
[0053] In a possible implementation, step S110 includes:
[0054] Step S111, when conducting drug clinical trials on multiple target drugs, tracking and recording drug feedback events of each of the target drugs, and generating drug feedback data of each of the target drugs.
[0055] In this embodiment, the server first establishes a connection with the data management system of the drug clinical trial through a secure network connection. This data management system is an integrated platform responsible for collecting, storing and managing all data about the drug clinical trial. The server interacts with the data management system through an API interface or a database connection string to ensure the real-time and security of data transmission.
[0056] Once the connection is successful, the server begins to track and record the clinical trials of each target drug. Clinical trials are usually conducted simultaneously in multiple research centers, and each research center has a dedicated data entry clerk responsible for recording patients' drug feedback events. These drug feedback events include but are not limited to:
[0057] Time of taking medication: Record the specific time when the patient takes medication each time.
[0058] Medication dosage: Record the medication dosage each time the patient takes it to ensure data consistency.
[0059] Immediate reactions after taking the medicine: Record any immediate reactions that the patient experiences after taking the medicine, such as nausea, vomiting, dizziness, etc.
[0060] Changes in physiological indicators: Regularly record the patient's physiological indicators, such as blood pressure, heart rate, blood sugar, etc., to evaluate the impact of drugs on the patient's physiological state.
[0061] Symptom improvement: Record whether the patient's symptoms improve after taking the medication, and the extent of the improvement.
[0062] Occurrence of adverse reactions: Record in detail the time of occurrence, symptoms, severity and treatment of any adverse reactions.
[0063] The server captures these records from the data management system through polling or real-time monitoring, and stores them in the local database in preparation for subsequent analysis and processing.
[0064] Step S112, parsing the first drug action status data from the drug feedback data of each of the target drugs.
[0065] In this embodiment, in addition to directly recording the patient feedback events, the server also needs to parse the first drug action status data from the drug feedback data. These first drug action status data are usually hidden in the patient's physiological index changes and symptom improvement, and need to be extracted through professional medical knowledge and data analysis technology.
[0066] For example, for a blood pressure medication, the server will pay special attention to the patient's blood pressure change data. By analyzing the patient's blood pressure readings before and after taking the medication, the server can calculate the blood pressure-lowering effect of the medication and use it as part of the first medication action status data. This data is crucial for evaluating the efficacy and safety of the medication.
[0067] Step S113, acquiring the test instrument detection data of each of the target drugs, and extracting the second drug action status data from the test instrument detection data.
[0068] Wherein, the clinical feedback big data includes the drug feedback data, the first drug action status data and the second drug action status data.
[0069] In addition to the subjective feedback from patients, various instruments are used during the trial to objectively monitor the drug's action status. These instruments include but are not limited to electrocardiographs, blood pressure monitors, blood glucose meters, etc. The server obtains the test data of these instruments through the data management system and extracts the second drug action status data from them.
[0070] For example, for a heart treatment drug, the server will extract the patient's electrocardiogram waveform changes from the electrocardiograph's detection data to evaluate the drug's effect on heart function. These data are corroborated with the patient's subjective feedback data, and together they form a comprehensive description of the drug's action status.
[0071] By integrating patient feedback data, first drug action status data, and second drug action status data, the server generates a large clinical feedback data set for each target drug. These data sets contain rich information and can fully reflect the performance of the drug in clinical trials. The server stores these data sets in a dedicated database for subsequent in-depth analysis and processing.
[0072] As a concrete example, suppose a new antidepressant drug is being tested in a clinical trial. Here is an example of what the server does when executing the steps above:
[0073] The server establishes a secure connection with the data management system through the HTTPS protocol and uses the preset API key for authentication and authorization. The server periodically polls the "Patient Feedback Event" table in the data management system. It captures all feedback event records about new antidepressants in the last 24 hours. These records are sorted according to fields such as patient ID, drug ID, feedback time, etc., and stored in the local database.
[0074] The server extracts data on changes in patient physiological indicators of new antidepressants from the local database. Machine learning algorithms (such as regression analysis) are used to analyze these data and calculate the antidepressant effect score of the drug. The score result is used as the first drug action state data and associated with the corresponding feedback event record. The server obtains the test data of instruments such as electrocardiographs and sphygmomanometers through the data management system interface. Data on the effects of new antidepressants on cardiac function and blood pressure are extracted from these test data. These data are used as the second drug action state data and associated with the corresponding feedback event record.
[0075] The server integrates the patient feedback data, the first drug action status data, and the second drug action status data. A complete clinical feedback big data set is generated for each target drug, including multiple fields such as patient ID, drug ID, feedback time, feedback content, physiological index changes, instrument detection data, etc. These big data sets are stored in a dedicated database for further analysis and processing.
[0076] Through the above steps, the server obtains the clinical feedback big data generated by the clinical trials of multiple target drugs. These clinical feedback big data provide a solid foundation for the subsequent extraction of drug action characteristics, patient response characteristics, drug clinical feedback characteristics generation and drug adverse reaction label determination.
[0077] In a possible implementation, step S120 includes:
[0078] Step S121, integrating the first drug action state data and the second drug action state data to generate target drug action state data.
[0079] Step S122, using a drug action encoder to encode the target drug action state data to generate drug action characteristic data.
[0080] Step S123, extracting patient response characteristic data of the drug feedback data using a patient response encoder.
[0081] In this embodiment, after obtaining the clinical feedback big data of multiple target drugs, the server needs to extract drug action characteristic data and patient response characteristic data from these big data. These data are the basis for subsequent analysis of drug effects and adverse reactions.
[0082] The server first processes data about the drug action status. These data come from two main channels: the first drug action status data parsed from patient feedback data, and the second drug action status data extracted from test instrument detection data.
[0083] First, drug action status data: usually includes the patient's subjective feelings and some quantifiable changes in physiological indicators, such as symptom improvement, pain score, blood pressure changes, etc. These data reflect the direct effect of the drug on the patient's symptoms.
[0084] Second, drug action status data: obtained through objective testing with test instruments such as electrocardiograms, sphygmomanometers, and blood glucose meters, providing a more accurate description of the drug's effect on the patient's physiological functions. For example, for cardiovascular drugs, these data may include changes in heart rate and electrocardiogram waveforms.
[0085] The server integrates these two parts of data through a specific data processing algorithm. The integration process may include steps such as data cleaning (removing outliers, missing value processing, etc.), data standardization (ensuring that data from different sources are comparable on the same scale) and data fusion (merging related data items into a comprehensive indicator). Finally, the server generates comprehensive action status data for each target drug, namely, target drug action status data.
[0086] After obtaining the target drug action state data, the server then uses the drug action encoder to encode these data to generate drug action feature data. The drug action encoder is a trained machine learning model that can capture the complex features of the drug action state and convert them into numerical feature vectors.
[0087] In practical applications, drug action encoders are usually trained on a large amount of historical drug trial data. These data contain records of the effects of various drugs on different patients. By learning the patterns in these records, the encoder learns how to convert drug action states into useful feature representations. During the encoding process, the server inputs the target drug action state data into the drug action encoder. The encoder processes the input data, extracts key feature information, and outputs a fixed-length feature vector as drug action feature data. This feature vector contains a comprehensive description of the drug's effects and can be used for subsequent drug effect evaluation and adverse reaction prediction.
[0088] Similar to the process of extracting drug action characteristic data, the server also needs to extract patient response characteristic data from drug feedback data. These patient response characteristic data describe the various physiological and psychological reactions of patients after taking drugs, and are an important basis for evaluating drug safety and tolerability. Similar to the drug action encoder, the patient response encoder is also a trained machine learning model. It is specifically used to process patient feedback data and extract key features about patient responses. During the extraction process, the server first preprocesses the patient feedback data, including steps such as text cleaning, word segmentation, and removal of stop words. Then, the preprocessed data is input into the patient response encoder. The encoder processes the input data, identifies key information in the patient response (such as symptom description, emotional changes, etc.), and converts it into a numerical feature vector as patient response characteristic data. These patient response characteristic data, together with the drug action characteristic data, constitute a comprehensive description of the performance of the target drug in clinical trials, providing important input information for subsequent drug effect evaluation and adverse reaction prediction.
[0089] For example, suppose a new antibiotic is being tested in a clinical trial. Here is an example of what the server would do as it performs the steps above:
[0090] The server extracts the first drug action state data (such as changes in patient temperature, changes in white blood cell count, etc.) and the second drug action state data (such as bacterial culture results, drug sensitivity test results, etc.) about the new antibiotics from the clinical feedback big data. Through data cleaning and standardization, the server integrates these data into a comprehensive target drug action state data set. The server loads a pre-trained drug action encoder model. The target drug action state data set is input into the encoder, and the encoder outputs a fixed-length drug action feature vector. The server extracts patient feedback data from the clinical feedback big data and performs text preprocessing. The pre-processed patient feedback data is input into the pre-trained patient response encoder. The encoder outputs a fixed-length patient response feature vector, which contains a comprehensive description of the various reactions of patients after taking new antibiotics.
[0091] Through the above steps, the server extracted drug action characteristic data and patient response characteristic data from the clinical feedback big data, laying a solid foundation for subsequent drug effect evaluation and adverse reaction prediction.
[0092] In a possible implementation, step S130 includes:
[0093] Step S131, among the drug action characteristic data and patient response characteristic data corresponding to the same target drug, determine the target drug action characteristic data and target patient response characteristic data extracted based on the same drug feedback event in the target drug.
[0094] Step S132, the target drug action characteristic data and the target patient response characteristic data are integrated to generate a drug clinical feedback characteristic corresponding to the drug feedback event in the same target drug.
[0095] In this embodiment, after extracting the drug action characteristic data and patient response characteristic data of the target drug, the server needs to merge these two parts of data to generate comprehensive drug clinical feedback characteristics. These drug clinical feedback characteristics will comprehensively reflect the effect of the drug in clinical trials and the patient's response, providing an important basis for subsequent drug evaluation and adverse reaction prediction.
[0096] Specifically, the server first needs to find the feature data extracted based on the same drug feedback event among the drug action feature data and patient response feature data corresponding to the same target drug. This means that the server needs to identify which drug action feature data and which patient response feature data are derived from the same clinical feedback event.
[0097] The server matches the drug action characteristic data and patient response characteristic data according to the feedback events by comparing the timestamp, patient ID, drug ID and other key information in the data set. For example, if a patient has symptoms of decreased body temperature and reduced pain after taking the drug, the server needs to find the drug action characteristic data (such as changes in body temperature, decreased inflammatory indicators, etc.) and patient response characteristic data (such as reduced pain scores, improved comfort, etc.) corresponding to the event. During the matching process, the server will also screen the data to ensure that the extracted characteristic data is accurate, complete and relevant. For data with missing values, outliers or inconsistencies, the server will pre-process or eliminate them.
[0098] After determining the target drug action characteristic data and target patient response characteristic data based on the same drug feedback event, the server will then merge these two parts of data. The purpose of the fusion is to combine the drug's effect and the patient's response to form a comprehensive drug clinical feedback characteristic representation.
[0099] In detail, the server uses a specific feature fusion algorithm to process these two parts of data. This algorithm may involve multiple technologies such as feature splicing, feature fusion, and feature conversion. For example, the server can directly splice the target drug action feature vector and the target patient response feature vector to form a longer feature vector; or use a fusion mechanism (such as attention mechanism, gating mechanism, etc.) to weightedly fuse the two parts of the feature to generate a comprehensive feature representation. After the fusion process, the server generates drug clinical feedback features corresponding to each drug feedback event in the same target drug. These features combine the effects of the drug and the patient's response, providing more comprehensive and in-depth information for subsequent drug effect evaluation and adverse reaction prediction.
[0100] For example, suppose a new blood pressure medication is being tested in a clinical trial. Here is an example of what the server would do when executing the steps above:
[0101] The server retrieves all drug action characteristic data and patient response characteristic data about new antihypertensive drugs from the database. By comparing information such as timestamps, patient IDs, and drug IDs, the server matches the characteristic data belonging to the same drug feedback event. For example, if a patient experiences a drop in blood pressure and dizziness after taking antihypertensive drugs, the server will find the drug action characteristic data (such as changes in blood pressure readings) and patient response characteristic data (such as dizziness descriptions, comfort scores, etc.) corresponding to the event.
[0102] The server loads the predefined feature fusion algorithm model. The matched target drug action feature vector and target patient response feature vector are input into the fusion algorithm model. The fusion algorithm model processes and fuses these two features to generate a comprehensive drug clinical feedback feature vector. This drug clinical feedback feature vector contains both the drug's blood pressure lowering effect information and the patient's dizziness reaction information after taking the drug.
[0103] Through the above steps, the server integrates the drug action characteristic data and patient response characteristic data corresponding to the same target drug, and generates comprehensive drug clinical feedback characteristics. These characteristics will be used in subsequent drug effect evaluation and adverse reaction prediction work, providing important decision support for pharmaceutical companies.
[0104] In a possible implementation, step S140 includes:
[0105] Step S141, polling the drug clinical feedback characteristics of each of the target drugs to generate a target drug clinical feedback characteristic.
[0106] Step S142, respectively determining the drug effect matching degree between the clinical feedback characteristics of the target drug and the clinical feedback characteristics of other drugs.
[0107] Step S143, based on the drug action matching degree, determine the combined drug clinical feedback characteristics that meet the drug combination application conditions with the target drug clinical feedback characteristics from the other drug clinical feedback characteristics.
[0108] In this embodiment, after generating the drug clinical feedback features of each target drug, the server needs to further analyze these drug clinical feedback features to explore the possible synergistic or complementary effects between different drugs. This usually involves searching for the combination drug clinical feedback features that match the target drug clinical feedback features among a large number of drug clinical feedback features, that is, looking for other drugs that may have a positive effect with the target drug in the drug combination application.
[0109] In detail, the server first traverses or polls the drug clinical feedback features of all target drugs. This means that the server processes the clinical feedback features of each target drug one by one and uses them as the clinical feedback features of the target drug currently being analyzed. For example, the server accesses the clinical feedback feature set of each target drug in a cyclic or iterative manner. In each iteration, the server selects a clinical feedback feature as the current clinical feedback feature of the target drug.
[0110] For each target drug clinical feedback feature, the server then needs to evaluate the drug action match between it and all other drug clinical feedback features. This usually involves comparing the similarity, complementarity or synergy of different drugs in clinical feedback features.
[0111] In detail, the server may use a variety of algorithms to calculate the drug action matching, such as cosine similarity, Pearson correlation coefficient, rule-based matching algorithm, etc. These algorithms can quantify the similarity or correlation between different drugs in clinical feedback characteristics. When calculating the matching degree, the server will consider information from multiple dimensions, including the drug's mechanism of action, target, metabolic pathway, patient response type, etc. This information helps to comprehensively evaluate the potential for interaction between different drugs.
[0112] Finally, based on the results of drug action matching, the server will screen out the clinical feedback characteristics of combination drugs that match the clinical feedback characteristics of the target drug from the clinical feedback characteristics of other drugs. These combination characteristics represent other drugs that may produce positive effects with the target drug in drug combination applications.
[0113] Specifically, the server will preset a matching threshold. Only when the matching degree between the clinical feedback characteristics of other drugs and the clinical feedback characteristics of the target drug exceeds this threshold, will it be considered as a combination drug clinical feedback characteristic that meets the conditions for drug combination application. After screening out the combination drug clinical feedback characteristics that meet the conditions, the server will output or store these results for reference in subsequent drug combination research or clinical application.
[0114] For example, suppose that the potential for drug combination application of a new anticancer drug A is being studied. The following is an example of the specific operations of the server when executing the above steps:
[0115] The server traverses the drug clinical feedback feature sets of all target drugs (including new anticancer drug A and other potential combination drugs). In each iteration, the server uses a clinical feedback feature of the new anticancer drug A as the current target drug clinical feedback feature. For the current target drug clinical feedback feature (i.e., a clinical feedback feature of the new anticancer drug A), the server calculates the drug action matching degree between it and the clinical feedback features of all other drugs. The server uses the cosine similarity algorithm to quantify the similarity of different drugs in the clinical feedback feature vector. This algorithm takes into account information in multiple dimensions (such as drug targets, metabolic pathways, patient response types, etc.) and generates a matching score.
[0116] The server presets a matching threshold (for example, 0.7) to screen the clinical feedback characteristics of combination drugs that meet the conditions. The server compares the matching scores between the clinical feedback characteristics of other drugs and the clinical feedback characteristics of the target drug with the threshold. Only when the matching score exceeds the threshold, the clinical feedback characteristic is considered to be a clinical feedback characteristic of the combination drug that is consistent with the clinical feedback characteristic of the target drug. Finally, the server outputs a list of all clinical feedback characteristics of eligible combination drugs. This list provides important reference information for subsequent drug combination studies. For example, researchers can further explore the synergistic mechanism or complementary effect of the new anticancer drug A and other drugs based on this list.
[0117] In a possible implementation, the scale of the feature grouping is multiple. Step S150 includes:
[0118] Step S151, determining a target feature grouping of each of the target drugs according to the scale of the drug clinical feedback feature of each of the target drugs in the plurality of feature groups.
[0119] Step S152, determining the adverse drug reaction label corresponding to the target feature group of each target drug.
[0120] Step S153, using the adverse drug reaction labels corresponding to the target feature groups of each of the target drugs as the adverse drug reaction labels corresponding to each of the target drugs.
[0121] In this embodiment, after completing the search and combination of drug clinical feedback features, the server needs to determine the drug adverse reaction label corresponding to each target drug based on these feature groups. Since the scale of feature groups may be multiple, the server needs an effective method to filter out the most representative or most relevant feature groups from these groups, and determine the drug adverse reaction labels based on these groups.
[0122] In detail, the server first analyzes the scale of the drug clinical feedback features in each feature grouping. This scale refers to the number, diversity, or correlation with other features within the grouping. By comparing the scales of different groups, the server can preliminarily screen out target feature groupings that may contain important information. The server calculates the number of features in each feature grouping and may further analyze the diversity (such as different reaction types, different patient groups, etc.) and correlation (such as co-occurrence frequency, mutual dependence, etc. between features) of these features. Based on the results of the scale assessment, the server adopts a screening mechanism to determine the target feature grouping. This mechanism may involve setting thresholds (such as the minimum number of features, the minimum diversity index, etc.) or applying a sorting algorithm (such as sorting based on feature importance) to screen out the most representative or relevant feature groupings.
[0123] For each target feature group, the server then needs to determine its corresponding adverse drug reaction label. This usually involves comprehensive analysis and interpretation of the features within the group to identify the common adverse reaction types.
[0124] In detail, the server conducts an in-depth analysis of each drug clinical feedback feature in the target feature grouping to identify key information related to adverse reactions (such as symptom descriptions, changes in physiological indicators, patient feedback, etc.). Based on the results of the feature analysis, the server applies a label generation algorithm to generate corresponding drug adverse reaction labels for each target feature grouping. This algorithm may involve a variety of technologies such as keyword extraction, cluster analysis, and pattern recognition to extract concise and accurate adverse reaction descriptions from complex clinical feedback features.
[0125] Finally, the server uses the adverse drug reaction labels corresponding to the target feature groups of each target drug as the final adverse drug reaction labels of the target drug. These labels will be used in subsequent drug evaluation, regulatory reporting, patient notification and other links.
[0126] For example, if a target drug corresponds to multiple target feature groups (i.e., there are multiple possible adverse reaction types), the server may need to further integrate the label information of these groups to generate a comprehensive and accurate adverse reaction description. The server outputs or stores the finalized adverse drug reaction labels in an appropriate format for subsequent use. These labels may be presented in text form or may be associated with a specific coding system (such as MedDRA) for international communication and comparison.
[0127] For example, suppose you are determining an adverse drug reaction label for a new cardiovascular drug B. Here is an example of what the server does when executing the above steps:
[0128] The server first analyzed all feature groups for the new cardiovascular drug B and calculated the number and diversity of drug clinical feedback features in each group. By setting thresholds (such as the minimum number of features = 10, the minimum diversity index = 0.5) and applying sorting algorithms (such as descending order based on feature importance), the server screened out two target feature groups G1 and G2. These two groups contain key clinical feedback features about abnormal heart rate and blood pressure fluctuations, respectively.
[0129] For target feature group G1 (abnormal heart rate), the server analyzed the feature descriptions within the group, such as "the patient's heart rate dropped significantly after taking the drug", "the electrocardiogram showed bradycardia", etc., and extracted the keywords "heart rate drop" and "bradycardia". Using the label generation algorithm, the server generated the drug adverse reaction label "bradycardia" for G1. Similarly, for target feature group G2 (blood pressure fluctuations), the server generated the adverse reaction label "unstable blood pressure" by analyzing the features and extracting keywords.
[0130] Since the new cardiovascular drug B corresponds to two target feature groups G1 and G2, the server integrates the adverse reaction labels of these two groups to generate the final adverse drug reaction description: "New cardiovascular drug B may cause bradycardia and blood pressure instability." The server outputs this description in text form and associates it with the MedDRA coding system for international communication and comparison.
[0131] In a possible implementation, step S152 includes:
[0132] Step S1521, performing self-attention processing in the target feature grouping to generate self-attention drug clinical feedback features.
[0133] Step S1522, classifying the adverse drug reactions of the self-attention drug clinical feedback features, and generating adverse drug reaction labels corresponding to the target feature groups.
[0134] In this embodiment, after determining the target feature groups of each target drug, the server needs to further analyze the clinical feedback features of the drugs in these groups to generate corresponding adverse drug reaction labels. To achieve this goal, the server uses a self-attention mechanism to process the data in the feature groups and uses a classification algorithm to identify the adverse reaction type.
[0135] The server first performs self-attention processing on the drug clinical feedback features within each target feature group. The self-attention mechanism allows the model to consider the information of all other features in the group when processing each feature, thereby capturing the intrinsic connections and interdependencies between features.
[0136] In detail, the server first converts the drug clinical feedback features within the group into numerical feature vector representations. These feature vectors may contain information in multiple dimensions, such as changes in physiological indicators, symptom descriptions, patient feedback, etc. For each feature vector within the group, the server calculates the attention weight between it and all other feature vectors in the group. These weights reflect the degree of correlation or importance between different features. By weighted summation, the server generates a self-attention representation of each feature, that is, an enhanced feature vector that takes into account the information of other features in the group. After self-attention processing, the server obtains a new set of feature vectors, namely, self-attention drug clinical feedback features. These features not only contain the information of the original features, but also incorporate the correlation and importance information of other features in the group, providing a richer and more comprehensive input for the subsequent classification of adverse drug reactions.
[0137] After generating the self-attention drug clinical feedback features, the server then uses a classification algorithm to classify these features to generate adverse drug reaction labels corresponding to the target feature groupings.
[0138] In detail, the server selects a suitable classification model to perform the adverse reaction classification task. This model may be an algorithm based on machine learning (such as support vector machine, random forest) or deep learning (such as neural network, convolutional neural network). According to the characteristics of the data and the requirements of the task, the server can choose different model structures and parameter configurations. Before classification, the server needs to use known adverse drug reaction data to train the classification model. These data usually contain a series of drug clinical feedback features and corresponding adverse reaction labels. Through the training process, the model can learn the mapping relationship from features to labels and optimize its internal parameters to improve classification accuracy. After the training is completed, the server inputs the self-attention drug clinical feedback features into the classification model and performs classification processing. The model calculates the corresponding category probability or score based on the input features, and selects the most likely category as the adverse reaction label according to the preset threshold or rule. Finally, the server generates the corresponding adverse drug reaction label for each target feature group. These labels accurately describe the adverse reaction type pointed to by the drug clinical feedback features in the group, providing an important basis for drug evaluation and supervision.
[0139] For example, suppose you are determining adverse drug reaction labels for target feature group G3 of a new antidepressant drug C. Here is an example of what the server does when executing the above steps:
[0140] The server first converts the drug clinical feedback features in group G3 (such as mood change scores, sleep quality scores, patient self-reported symptoms, etc.) into numerical feature vector representations. Next, the server calculates the attention weights between each feature vector and other feature vectors in the group. These weights reflect the degree of correlation or importance between different features. By weighted summation, the server generates a self-attention representation of each feature, that is, an enhanced feature vector that takes into account other feature information in the group. These vectors constitute the self-attention drug clinical feedback feature set.
[0141] The server selected a deep learning-based neural network classification model to perform the adverse reaction classification task. The model has been trained using a large amount of known adverse drug reaction data. The self-attention drug clinical feedback feature set is input into the classification model, and the model calculates the corresponding category probability or score based on the input features. According to the preset threshold or rules (such as selecting the category with the highest probability as the label), the server generates a drug adverse reaction label "sleepiness" for group G3. This label accurately describes the type of adverse reaction pointed to by the drug clinical feedback features in the group.
[0142] Through the above steps, the server determines the adverse drug reaction label corresponding to the target feature group G3 of the new antidepressant drug C, providing strong support for drug evaluation and supervision.
[0143] In a possible implementation, step S140 may include: searching for a combination drug clinical feedback feature that meets the drug combination application condition with the target drug clinical feedback feature among the drug clinical feedback features of each of the target drugs and the drug clinical feedback features of multiple prior drugs.
[0144] Step S152 includes: obtaining the drug clinical feedback feature belonging to any one of the a priori drugs in the target feature group, and using the drug adverse reaction label of any one of the a priori drugs as the drug adverse reaction label of the target feature group.
[0145] In this embodiment, in the process of drug development, it is very important to understand the effects of combined applications of different drugs and their potential adverse reactions. The server analyzes the drug clinical feedback characteristics of the target drug and combines known prior drug data to search for combined drug clinical feedback characteristics that meet the drug combination application conditions and determine the corresponding drug adverse reaction labels.
[0146] In detail, the server first obtains the drug clinical feedback feature sets of multiple target drugs in the aforementioned embodiments, which contain information such as patient responses, changes in physiological indicators, etc. collected in clinical trials of the target drugs. At the same time, the server also maintains a database containing drug clinical feedback features of multiple prior drugs. These prior drugs are drugs that have been extensively studied and clinically verified, and their drug clinical feedback features and adverse reaction labels are known.
[0147] The server starts to search for matches between the drug clinical feedback features of each target drug and the drug clinical feedback features of the prior drug. This process is similar to finding similar or related items in a large-scale data set. The server uses an efficient search algorithm (such as a similarity-based sorting algorithm, a rule-based matching algorithm, etc.) to compare the features of the target drug and the prior drug one by one. The goal of the search is to find the clinical feedback features of the prior drug that are significantly similar or complementary to the clinical feedback features of the target drug. The basis for judging similarity or complementarity may include the Euclidean distance of the feature vector, cosine similarity, feature importance score, etc. The server will determine which feature combinations meet the conditions for drug combination application based on preset thresholds or rules.
[0148] After searching, the server outputs a set of clinical feedback features of combined drugs that meet the drug combination application conditions. These feature combinations represent possible effective combination application schemes between the target drug and the prior drugs.
[0149] The server has divided the drug clinical feedback features of the target drug into multiple feature groups according to a certain rule or algorithm. Next, the server will review each feature group to determine whether it contains drug clinical feedback features belonging to the prior drug. If the drug clinical feedback features belonging to the prior drug are found in the target feature group, the server will perform a label transfer operation. Specifically, the server will look up the drug adverse reaction label corresponding to the prior drug in the database. Since the adverse reaction labels of the prior drug have been verified and confirmed, the server can directly use these labels as drug adverse reaction labels for the target feature group. Doing so can greatly improve the efficiency and accuracy of label determination.
[0150] If a target feature group contains clinical feedback features of multiple prior drugs (i.e., there are multiple possible labels), the server may need to integrate these labels according to certain rules (such as majority voting, weighted average, etc.) to generate a comprehensive adverse drug reaction label. Finally, the server outputs or stores the adverse drug reaction label corresponding to each target feature group for subsequent drug evaluation, regulatory reporting, patient notification, and other links.
[0151] For example, suppose you are studying a new diabetes treatment drug D, and you want to explore its combined effects and adverse reactions with known glucose-lowering drugs (such as drug E and drug F).
[0152] The server obtains the drug clinical feedback feature set of the new diabetes treatment drug D, and performs a matching search with the drug clinical feedback feature database of drug E and drug F. By calculating the cosine similarity of the feature vectors, the server finds that some clinical feedback features of drug D are highly similar to some features of drug E, indicating that they may overlap or complement each other in terms of mechanism of action or target of action. Therefore, the server combines these similar features and outputs them as the clinical feedback features of the combined drug that meet the conditions for drug combination application.
[0153] The server reviewed the target feature groups of the new diabetes treatment drug D and found that one of the groups contained clinical feedback features similar to drug E. The server looked up the corresponding adverse drug reaction labels for drug E in the database (such as "hypoglycemia", "weight gain", etc.) and passed these labels directly to the target feature group. Ultimately, the server output "hypoglycemia" and "weight gain" as adverse drug reaction labels for this target feature group, prompting researchers and clinicians to pay attention to these potential adverse reactions when using the new diabetes treatment drug D in combination with drug E.
[0154] In a possible implementation, the drug action characteristic data and the patient response characteristic data are obtained using an autoencoder. Before step S110, the step further includes:
[0155] Step S101, obtaining clinical feedback big data of each template generated by conducting clinical trials on multiple template drugs.
[0156] In this embodiment, in the field of drug research and development, in order to more accurately extract drug action characteristic data and patient response characteristic data, researchers often use autoencoders, a deep learning model. Autoencoders can learn effective feature representations from raw data through unsupervised learning. In this scenario, the server will perform a series of steps to use the autoencoder to extract features from the clinical feedback big data of the template drug, and use these features to optimize the autoencoder, which is ultimately used for feature extraction of the target drug.
[0157] In detail, the server first connects to the data management system of drug clinical trials, from which it retrieves the clinical feedback big data generated by multiple template drugs for drug clinical trials. These template drugs have been extensively studied and verified, and their clinical feedback data have high reliability and reference value. The template clinical feedback big data contains multi-dimensional information such as changes in patient physiological indicators, symptom improvement, adverse reaction records, etc. collected in clinical trials of template drugs.
[0158] Step S102, using a reference autoencoder to extract the first sample patient response characteristic data and the first sample drug action characteristic data of each of the template clinical feedback big data, and blending the first sample patient response characteristic data and the first sample drug action characteristic data of each of the template clinical feedback big data to generate each first sample drug clinical feedback feature.
[0159] In this embodiment, the server loads a pre-trained reference autoencoder model. This reference autoencoder already has a certain feature extraction capability, but has not yet been optimized for the current task. The server inputs each template clinical feedback big data into the reference autoencoder, and extracts the first sample patient response feature data and the first sample drug action feature data through the processing of the encoder. These two feature data sets respectively describe the physiological and psychological reactions of the patient after taking the drug, as well as the effect of the drug in the patient's body. The server blends the first sample patient response feature data and the first sample drug action feature data of each template clinical feedback big data to generate each first sample drug clinical feedback feature. These blended features integrate the information of drug action and patient response, providing a basis for subsequent combination search and error parameter calculation.
[0160] Step S103, searching among the first sample drug clinical feedback features for a combination sample drug clinical feedback feature that meets the drug combination application condition with the target sample drug clinical feedback feature.
[0161] In this embodiment, the server executes a search algorithm among all the first sample drug clinical feedback features to find the combined sample drug clinical feedback features that meet the drug combination application conditions with the target sample drug clinical feedback features. The "target sample" here can be a set of clinical feedback features of one or more pre-selected template drugs. The search algorithm matches based on the similarity, complementarity or other preset rules between the features. The qualified combined sample drug clinical feedback features will be used as a reference for subsequent error parameter calculations.
[0162] Step S104: determining a combined error parameter based on the first label heat map of the target sample drug clinical feedback feature and the first label heat map of the combined sample drug clinical feedback feature.
[0163] In this embodiment, the server calculates a combination error parameter based on the difference between the first label heat map of the clinical feedback feature of the target sample drug and the first label heat map of the clinical feedback feature of the combination sample drug. This combination error parameter reflects the degree of deviation between the actual feature and the expected feature in the drug combination application.
[0164] Step S105 , determining a balanced distribution error parameter based on the first label heat map of the clinical feedback characteristics of the target sample drug.
[0165] In this embodiment, at the same time, the server also calculates a balanced distribution error parameter based on the first label heat map of the target sample drug clinical feedback feature. This balanced distribution error parameter is used to evaluate the uniformity of the distribution of the feature in the label space, which helps to prevent overfitting or underfitting.
[0166] Step S106: based on the combined error parameter and the balanced distribution error parameter, update the parameter information of the reference autoencoder to generate the autoencoder.
[0167] In this embodiment, the server uses the calculated combined error parameters and balanced distribution error parameters to update the parameter information of the reference autoencoder. This update process is usually implemented through a back propagation algorithm, which aims to minimize the error parameters and improve the accuracy and generalization ability of feature extraction. After multiple rounds of iterative updates, the server generates an autoencoder model optimized for the current task. This autoencoder will be used for subsequent feature extraction of clinical feedback big data of the target drug.
[0168] In a possible implementation, after step S101, the method further includes:
[0169] Step A110 , performing feature expansion on the template clinical feedback big data to generate each expanded clinical feedback big data of the template clinical feedback big data.
[0170] In this embodiment, after the server obtains the original clinical feedback big data generated by the clinical trials of multiple template drugs, it begins to perform feature expansion on these original clinical feedback big data. The purpose of feature expansion is to increase the diversity and complexity of the data so that the autoencoder can learn a more generalized feature representation. Feature expansion can be carried out in a variety of ways, such as data enhancement (such as adding noise, transforming scales, etc.), feature crossover (such as combining different features to generate new features), feature derivation (such as calculating statistics or trend indicators based on original features), etc. The server selects a suitable expansion method based on specific task requirements and data characteristics. After feature expansion, the server generates each extended clinical feedback big data set of the template clinical feedback big data. These extended clinical feedback big data sets contain richer and more diverse feature information than the original data.
[0171] Step A120, using a pre-trained autoencoder to extract the second sample drug clinical feedback features of the template clinical feedback big data and the first extended sample drug clinical feedback features of each of the extended clinical feedback big data.
[0172] In this embodiment, the server loads a pre-trained autoencoder model. This pre-trained autoencoder already has a certain feature extraction capability, but has not yet been optimized for the expanded data. The server inputs the template clinical feedback big data and each extended clinical feedback big data into the pre-trained autoencoder, and extracts the second sample drug clinical feedback features and the first extended sample drug clinical feedback features through the processing of the encoder. These features respectively describe the performance of the template drug and its extended data in clinical trials.
[0173] Step A130, determining a drug action matching error parameter based on the drug action matching between the second sample drug clinical feedback feature of the template clinical feedback big data and each of the first extended sample drug clinical feedback features.
[0174] In this embodiment, the server calculates the drug action matching degree between the second sample drug clinical feedback feature of the template clinical feedback big data and the first extended sample drug clinical feedback feature of each extended clinical feedback big data. This drug action matching degree reflects the consistency or difference in drug action effects between different data sets. The drug action matching degree can be calculated in a variety of ways, such as calculating the similarity between feature vectors (such as cosine similarity), using a classifier to classify features and compare classification results, etc. The server selects an appropriate matching degree calculation method according to specific task requirements. Based on the matching degree calculation results, the server determines the drug action matching degree error parameter. This drug action matching degree error parameter quantifies the consistency error of the pre-trained autoencoder when extracting features from different data sets, reflecting the insufficiency of the autoencoder in generalization ability.
[0175] Step A140: updating the parameter information of the pre-trained autoencoder based on the drug action matching error parameter to generate the reference autoencoder.
[0176] In this embodiment, the server uses the calculated drug action matching error parameter to update the parameter information of the pre-trained autoencoder. This update process is intended to reduce the value of the error parameter and improve the consistency and accuracy of the autoencoder when extracting features from different data sets. The update method generally involves optimization techniques such as back propagation algorithm and gradient descent algorithm. The server calculates the gradient of the error parameter with respect to the autoencoder parameter and updates the parameter value in the direction of gradient descent to minimize the error parameter. After multiple rounds of iterative updates, the server generates an optimized autoencoder model, namely the reference autoencoder. This reference autoencoder has better generalization ability and accuracy when extracting drug clinical feedback features, and provides more reliable model support for subsequent target drug feature extraction.
[0177] In a possible implementation, the pre-trained autoencoder includes a pre-trained patient response encoder and a pre-trained drug action encoder, and the template clinical feedback big data includes template drug feedback data and template drug action state data.
[0178] Step A120 includes:
[0179] Step A121, using the pre-trained patient response encoder, respectively encode the template drug feedback data and each of the extended clinical feedback big data to generate sample drug feedback feature data and each of the sample extended feedback feature data.
[0180] Step A122, using the pre-trained drug action encoder, extracting sample drug action characteristic data of the template drug action state data.
[0181] Step A123, respectively blending the sample drug action characteristic data with the sample drug feedback characteristic data and each of the sample extended feedback characteristic data to generate a second sample drug clinical feedback characteristic of the template clinical feedback big data and a first extended sample drug clinical feedback characteristic of each of the extended clinical feedback big data.
[0182] In this embodiment, in the process of drug development, in order to accurately extract the clinical feedback characteristics of the drug, researchers often use a pre-trained autoencoder model. This model can learn an effective feature representation method by pre-training on a large amount of data. In this scenario, the pre-trained autoencoder consists of a pre-trained patient response encoder and a pre-trained drug action encoder, which are used to process template drug feedback data and template drug action state data, respectively. The server will use these two encoders to extract the drug clinical feedback characteristics of the template clinical feedback big data and its extended data.
[0183] Pre-trained patient response encoder: This encoder is specially designed to process patient feedback data and can learn key features of patient responses. These features may include symptom descriptions, changes in physiological indicators, patient feelings, etc.
[0184] Pre-trained drug action encoder: This encoder focuses on extracting features from drug action state data, which reflect the effects and mechanisms of the drug in the patient's body. This may include changes in drug concentration, target activity, metabolic pathways, etc.
[0185] Template drug feedback data: Contains patient feedback information collected from clinical trials of multiple template drugs, such as symptom improvement, adverse reaction records, etc.
[0186] Template drug action status data: records the action status information of the template drug in the patient's body, such as drug concentration, target activity, etc. This information is usually obtained through test instrument detection.
[0187] The server first loads the pre-trained patient response encoder. The template drug feedback data is input into the encoder, which encodes the input data and generates sample drug feedback feature data. These feature data capture the key information in the patient feedback.
[0188] At the same time, the server loads the pre-trained drug action encoder. The template drug action state data is input into the drug action encoder, and the encoder extracts the sample drug action feature data. These feature data describe the action state and effect of the drug in the patient's body.
[0189] The server blends the sample drug action feature data with the sample drug feedback feature data. The purpose of blending is to combine the drug action effect and patient feedback to generate a comprehensive drug clinical feedback feature representation. The blending process may involve multiple technologies such as feature splicing, weighted fusion, and attention mechanism. The server selects the appropriate blending method based on the specific task requirements. After the blending process, the server generates the second sample drug clinical feedback feature of the template clinical feedback big data.
[0190] For each extended clinical feedback big data set (extended from the template clinical feedback big data), the server repeats the above process. First, the extended data is encoded and represented using the pre-trained patient response encoder to generate each sample extended feedback feature data. Then, the pre-trained drug action encoder is used to extract the sample drug action feature data of the extended data (if the extended data also contains drug action status information). Finally, the sample drug action feature data is blended with the sample extended feedback feature data to generate the first extended sample drug clinical feedback feature of each extended clinical feedback big data.
[0191] For example, in a possible implementation, after step S106, the following is further included:
[0192] Step S1061 , performing feature expansion on each of the template clinical feedback big data to generate a plurality of expanded clinical feedback big data of each of the template clinical feedback big data.
[0193] Step S1062: using the autoencoder to extract the third sample drug clinical feedback feature of each of the template clinical feedback big data and the second extended sample drug clinical feedback feature of each of the extended clinical feedback big data.
[0194] Step S1063, determining a second label heat map of the clinical feedback characteristics of each of the third sample drugs, and a second label heat map of the clinical feedback characteristics of each of the second extended sample drugs.
[0195] Step S1064: Based on the set threshold value and each of the second label heat maps, determine the fuzzy prediction results corresponding to each of the second label heat maps.
[0196] Step S1065, polling each of the second label heat maps to generate a target label heat map, determining a first scale of the fuzzy prediction results associated with the target label heat map, and determining a second scale of each of the second label heat maps in the fuzzy prediction results associated with the target label heat map.
[0197] Step S1066, determining the update coefficient of the target label heat map based on the first scale and the second scale, determining the thermal divergence error parameter of the target label heat map based on the target label heat map and each of the second label heat maps in the fuzzy prediction results associated with the target label heat map, updating the parameter information of the autoencoder based on the thermal divergence error parameters and the update coefficient of each of the target label heat maps generated by polling, and generating a trained autoencoder.
[0198] In this embodiment, after the autoencoder is initially generated, in order to further improve its accuracy and generalization ability in extracting drug clinical feedback features, the server will perform a deeper feature expansion on the template clinical feedback big data, and use the expanded data to iteratively train the autoencoder. This process is achieved by calculating the thermal divergence error parameter of the label heat map and updating the parameter information of the autoencoder accordingly.
[0199] In detail, the server again performs feature expansion on each template clinical feedback big data, but this time after the autoencoder has been generated. The method of feature expansion may be the same as before, but it may also be adjusted based on the performance feedback of the autoencoder. As a result, the server generates multiple extended clinical feedback big data sets for each template clinical feedback big data, which are richer in feature dimension and complexity.
[0200] Next, the server loads the generated autoencoder model and uses it to extract the third sample drug clinical feedback features of each template clinical feedback big data and the second extended sample drug clinical feedback features of each extended clinical feedback big data. These feature data are represented in the form of numerical vectors, capturing the key information of the drug in clinical trials.
[0201] Then, the server generates a corresponding second label heat map for each third sample drug clinical feedback feature and the second extended sample drug clinical feedback feature. A label heat map is a visualization tool used to show the strength of association between each element in a feature vector and a specific label. The generation of a label heat map may involve inputting the feature vector into a classifier or regressor, calculating the similarity or correlation between each element and the label based on the output result, and mapping the similarity value to a color intensity.
[0202] The server determines the fuzzy prediction result corresponding to each heat map based on the set threshold value and each second label heat map. The fuzzy prediction result is a quantitative representation of the degree of association between the feature vector and the specific label, taking into account the uncertainty of the prediction. The set threshold value is pre-set according to task requirements and data characteristics, and is used to distinguish between clear predictions and fuzzy predictions. When the similarity value of an element in the label heat map exceeds the threshold value, it is considered that there is a clear association between the element and the label; otherwise, the association is considered to be fuzzy. The server polls each second label heat map, selecting one as the target label heat map each time.
[0203] For each target label heat map, the server determines the first scale of its associated fuzzy prediction results (i.e., the number of fuzzy prediction results), and the second scale of each second label heat map in these fuzzy prediction results (i.e., the number of times each second label heat map appears in the fuzzy prediction results). Based on the first scale and the second scale, the server calculates an update coefficient for adjusting the subsequent autoencoder parameter update process. The server calculates the thermal divergence error parameter based on the difference between the target label heat map and each second label heat map in its associated fuzzy prediction results. This parameter quantifies the degree of change of the label heat map during the iterative training process, reflecting the room for improvement in the performance of the autoencoder.
[0204] Next, the server uses the thermal divergence error parameters and corresponding update coefficients of each target label heat map generated by polling to update the parameter information of the autoencoder through optimization techniques such as backpropagation algorithm and gradient descent algorithm. This process may require multiple rounds of iterative training until the thermal divergence error parameter converges to a stable value or reaches the preset training round limit. After multiple rounds of iterative training, the server generates a trained autoencoder model. This model has higher accuracy and generalization ability when extracting drug clinical feedback features.
[0205] Figure 2 The hardware structure of the big data-based adverse drug reaction monitoring system 100 for implementing the above-mentioned big data-based adverse drug reaction monitoring method provided in the embodiment of the present application is shown. Figure 2As shown, the big data-based adverse drug reaction monitoring system 100 may include a processor 110 , a machine-readable storage medium 120 , a bus 130 , and a communication unit 140 .
[0206] In a possible design, the drug adverse reaction monitoring system 100 based on big data can be a single server or a server group. The server group can be centralized or distributed (for example, the drug adverse reaction monitoring system 100 based on big data can be a distributed system). In some embodiments, the drug adverse reaction monitoring system 100 based on big data can be local or remote. For example, the drug adverse reaction monitoring system 100 based on big data can access information and / or data stored in the machine-readable storage medium 120 via a network. For another example, the drug adverse reaction monitoring system 100 based on big data can be directly connected to the machine-readable storage medium 120 to access the stored information and / or data. In some embodiments, the drug adverse reaction monitoring system 100 based on big data can be implemented on the drug adverse reaction monitoring system based on big data. As an example only, the drug adverse reaction monitoring system based on big data can include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an internal cloud, a multi-layer cloud, etc. or any combination thereof.
[0207] The machine-readable storage medium 120 may store data and / or instructions. In some embodiments, the machine-readable storage medium 120 may store data acquired from an external terminal. In some embodiments, the machine-readable storage medium 120 may store data and / or instructions that the big data-based adverse drug reaction monitoring system 100 uses to execute or use to complete the exemplary methods described in this application.
[0208] During the specific implementation process, one or more processors 110 execute computer executable instructions stored in the machine-readable storage medium 120, so that the processor 110 can execute the drug adverse reaction monitoring method based on big data in the above method embodiment. The processor 110, the machine-readable storage medium 120 and the communication unit 140 are connected through the bus 130, and the processor 110 can be used to control the sending and receiving actions of the communication unit 140.
[0209] The specific implementation process of the processor 110 can refer to the various method embodiments executed by the above-mentioned big data-based drug adverse reaction monitoring system 100. The implementation principles and technical effects are similar, and this embodiment will not be repeated here.
[0210] In addition, an embodiment of the present application also provides a readable storage medium, in which computer executable instructions are preset. When a processor executes the computer executable instructions, the above-mentioned drug adverse reaction monitoring method based on big data is implemented.
[0211] It should be noted that in order to simplify the description disclosed in the present application and thereby facilitate the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present application, various features are sometimes grouped into one embodiment, drawing or description thereof.
Claims
1. A method for monitoring adverse drug reactions based on big data, characterized in that: The method comprises: Obtaining big data on clinical feedback generated from drug clinical trials for multiple target drugs; Extracting drug action characteristic data and patient response characteristic data from the clinical feedback big data of each of the target drugs; The drug action characteristic data and the patient response characteristic data corresponding to the same target drug are respectively integrated to generate drug clinical feedback characteristics of each target drug; Searching for a combination drug clinical feedback feature that meets the drug combination application condition with the target drug clinical feedback feature among the drug clinical feedback features of each of the target drugs, and grouping the target drug clinical feedback feature with the combination drug clinical feedback feature that meets the drug combination application condition to generate a feature group; Determining the adverse drug reaction label corresponding to each of the target drugs based on the feature grouping; The drug action characteristic data and the patient response characteristic data are obtained by using an autoencoder; before obtaining the clinical feedback big data generated by conducting drug clinical trials on multiple target drugs, it also includes: Obtain clinical feedback big data of each template generated by conducting drug clinical trials on multiple template drugs; Extracting the first sample patient response characteristic data and the first sample drug action characteristic data of each of the template clinical feedback big data using a reference autoencoder, and blending the first sample patient response characteristic data and the first sample drug action characteristic data of each of the template clinical feedback big data to generate each first sample drug clinical feedback feature; Searching for a combination sample drug clinical feedback feature that meets the drug combination application condition with the target sample drug clinical feedback feature among the first sample drug clinical feedback features; Determining a combined error parameter based on the first label heat map of the target sample drug clinical feedback feature and the first label heat map of the combined sample drug clinical feedback feature; Determining a balanced distribution error parameter based on a first label heat map of the clinical feedback characteristics of the target sample drug; Based on the combined error parameter and the balanced distribution error parameter, the parameter information of the reference autoencoder is updated to generate the autoencoder.
2. The method for monitoring adverse drug reactions based on big data according to claim 1, characterized in that: The acquisition of clinical feedback big data generated by conducting drug clinical trials on multiple target drugs includes: When conducting drug clinical trials on multiple target drugs, tracking and recording drug feedback events of each of the target drugs to generate drug feedback data for each of the target drugs; parsing first drug action state data from drug feedback data of each of the target drugs; Acquire the test instrument detection data of each of the target drugs, and extract the second drug action state data from the test instrument detection data; Wherein, the clinical feedback big data includes the drug feedback data, the first drug action state data and the second drug action state data; And, extracting drug action characteristic data and patient response characteristic data from the clinical feedback big data of each target drug includes: Integrating the first drug action state data and the second drug action state data to generate target drug action state data; Using a drug action encoder to encode the target drug action state data to generate drug action characteristic data; A patient response encoder is used to extract patient response characteristic data of the drug feedback data.
3. The method for monitoring adverse drug reactions based on big data according to claim 1, characterized in that: The drug action characteristic data and patient response characteristic data corresponding to the same target drug are respectively integrated to generate drug clinical feedback characteristics of each target drug, including: Determine, among the drug action characteristic data and patient response characteristic data corresponding to the same target drug, the target drug action characteristic data and the target patient response characteristic data extracted based on the same drug feedback event in the target drug; The target drug action characteristic data and the target patient response characteristic data are integrated to generate a drug clinical feedback characteristic corresponding to the drug feedback event in the same target drug.
4. The method for monitoring adverse drug reactions based on big data according to claim 1, characterized in that: The step of searching for the combination drug clinical feedback characteristics that meet the drug combination application conditions among the drug clinical feedback characteristics of each of the target drugs includes: Polling the drug clinical feedback characteristics of each of the target drugs to generate a target drug clinical feedback characteristic; Determining the drug effect matching degree between the clinical feedback characteristics of the target drug and the clinical feedback characteristics of other drugs respectively; Based on the drug action matching degree, the clinical feedback characteristics of the combination drugs that meet the drug combination application conditions with the clinical feedback characteristics of the target drug are determined from the clinical feedback characteristics of the other drugs.
5. The method for monitoring adverse drug reactions based on big data according to claim 1, characterized in that: The scale of the feature grouping is multiple; and the step of determining the adverse drug reaction label corresponding to each target drug based on the feature grouping includes: Determining a target feature grouping for each of the target drugs according to the scale of the drug clinical feedback feature of each of the target drugs in the plurality of feature groups; Determine the adverse drug reaction labels corresponding to the target feature groups of each of the target drugs; The adverse drug reaction labels corresponding to the target feature groups of each of the target drugs are used as the adverse drug reaction labels corresponding to each of the target drugs.
6. The method for monitoring adverse drug reactions based on big data according to claim 5, characterized in that: The step of determining the adverse drug reaction labels corresponding to the target feature groups of the target drugs comprises: Performing self-attention processing in the target feature grouping to generate self-attention drug clinical feedback features; The self-attention drug clinical feedback features are used to classify adverse drug reactions, and adverse drug reaction labels corresponding to the target feature groups are generated.
7. The method for monitoring adverse drug reactions based on big data according to claim 5, characterized in that: The step of searching for the combination drug clinical feedback characteristics that meet the drug combination application conditions among the drug clinical feedback characteristics of each of the target drugs includes: Searching for a combination drug clinical feedback feature that meets the drug combination application condition with the target drug clinical feedback feature among the drug clinical feedback features of each of the target drugs and the drug clinical feedback features of multiple prior drugs; The step of determining the adverse drug reaction labels corresponding to the target feature groups of the target drugs comprises: Acquiring a drug clinical feedback feature belonging to any one of the prior drugs in the target feature group; Any adverse drug reaction label of the prior drugs is used as the adverse drug reaction label of the target feature grouping.
8. The method for monitoring adverse drug reactions based on big data according to claim 1, characterized in that: After obtaining the clinical feedback big data of each template generated by conducting drug clinical trials on multiple template drugs, the method further includes: Performing feature expansion on the template clinical feedback big data to generate each expanded clinical feedback big data of the template clinical feedback big data; Extracting the second sample drug clinical feedback features of the template clinical feedback big data and the first extended sample drug clinical feedback features of each of the extended clinical feedback big data using a pre-trained autoencoder; Determining a drug action matching error parameter based on the drug action matching between the second sample drug clinical feedback feature of the template clinical feedback big data and each of the first extended sample drug clinical feedback features; Updating parameter information of the pre-trained autoencoder based on the drug action matching error parameter to generate the reference autoencoder; Wherein, the pre-trained autoencoder includes a pre-trained patient response encoder and a pre-trained drug action encoder, and the template clinical feedback big data includes template drug feedback data and template drug action state data; The method of extracting the second sample drug clinical feedback features of the template clinical feedback big data and the first extended sample drug clinical feedback features of each of the extended clinical feedback big data by using the pre-trained autoencoder includes: Using the pre-trained patient response encoder, respectively encode the template drug feedback data and each of the extended clinical feedback big data to generate sample drug feedback feature data and each of the sample extended feedback feature data; extracting sample drug action characteristic data of the template drug action state data using the pre-trained drug action encoder; The sample drug action characteristic data are respectively integrated with the sample drug feedback characteristic data and each of the sample extended feedback characteristic data to generate a second sample drug clinical feedback characteristic of the template clinical feedback big data and a first extended sample drug clinical feedback characteristic of each of the extended clinical feedback big data.
9. A drug adverse reaction monitoring system based on big data, characterized in that: The big data-based adverse drug reaction monitoring system includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the big data-based adverse drug reaction monitoring method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Untoward drug reaction monitoring and early warning method
CN118280603A