Drug risk monitoring method based on multi-source data fusion

Through the multi-source data fusion and real-time stream processing framework, data heterogeneity and information island problems in drug risk monitoring are solved, efficient, accurate monitoring and timely early warning of drug risks are achieved, and public drug safety is ensured.

CN120234773AActive Publication Date: 2025-07-01ZHUHAI FOOD & DRUG INSPECTION INSTITUTE (ZHUHAI FOOD & DRUG (MEDICAL DEVICES) ADVERSE REACTION MONITORING CENTER

Patent Information

Application Number
CN202510702720.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Existing drug risk monitoring methods rely on a single data source, resulting in narrow information coverage, limited risk signal capture capabilities, and lack of effective integration methods when processing multi-source data, resulting in the accuracy and timeliness of information island problems and monitoring results being restricted.

Method used

Through multi-source data fusion, including electronic medical records of medical institutions, pharmacy sales records and social media feedback, the weighted collaborative filtering algorithm and Bayesian network are adopted, combined with a real-time stream processing framework, data cleaning, standardization, feature extraction and risk assessment are realized, and risk assessment results are dynamically adjusted.

Benefits of technology

It has achieved comprehensive coverage of drug risk monitoring, improved data coverage and accuracy, significantly improved monitoring timeliness and reliability, and supported real-time dynamic monitoring and early warning of drug safety supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234773A_ABST
    Figure CN120234773A_ABST
Patent Text Reader

Abstract

The invention discloses a drug risk monitoring method based on multi-source data fusion, and relates to the technical field of data analysis. Through multi-source data fusion and processing, the problems of data isomerism and information islands are solved, a comprehensive and high-quality data basis is provided for drug risk monitoring, a weighted collaborative filtering algorithm and a Bayesian network are adopted, potential risk signals are accurately mined, real-time dynamic monitoring and early warning of drug risks are achieved, and the drug risk monitoring and early warning efficiency is improved. The timeliness and the reliability of monitoring are obviously improved; meanwhile, the real-time stream processing framework is used for carrying out incremental updating on newly-added data, dynamically adjusting a risk assessment result, and carrying out automatic early warning when a risk value exceeds a standard, so that the contradiction between a complex process and a rapid early warning demand is effectively solved, the efficiency and accuracy of drug safety supervision are improved, and powerful support is provided for guaranteeing the public drug safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and in particular to a drug risk monitoring method based on multi-source data fusion. Background Art

[0002] Drug risk monitoring is a crucial research direction in the medical field, which is directly related to the safety of public drug use and health protection. With the expansion of the scope of drug use, timely and accurately identifying and evaluating drug risks has become the key to maintaining the stability of the medical system and improving the quality of patients' lives. However, current drug risk monitoring methods mostly rely on a single data source, such as adverse reaction reports from medical institutions or production records of pharmaceutical companies. This single reliance leads to a narrow information coverage and limited risk signal capture ability, making it difficult to comprehensively reflect the risk characteristics of the entire drug chain from production to use. In addition, when dealing with multi-source heterogeneous data, existing methods often suffer from the problem of information silos due to the lack of effective integration means, and the accuracy and timeliness of monitoring results are significantly restricted.

[0003] From a technical perspective, the core challenge in the field of drug risk monitoring lies in how to efficiently fuse multi-source data and mine potential risk signals from it. Specifically, first is the problem of data heterogeneity. Data such as electronic medical records of medical institutions, pharmacy sales data, and social media feedback vary greatly in structure and format, and are difficult to clean and standardize. Second is the insufficient accuracy of fusion algorithms. Traditional methods are difficult to balance the credibility and relevance of different data sources, resulting in the possibility that the fusion results may mask key risk signals. Finally, the real-time requirement poses a higher challenge to existing analysis models, and there is a contradiction between complex data processing processes and the need for rapid early warning. Without solving these technical factors, it has become a difficult problem to comprehensively monitor and timely warn of drug risks. Summary of the Invention

[0004] In order to solve the above-mentioned existing technical problems, the present invention provides a drug risk monitoring method based on multi-source data fusion.

[0005] The technical solution of the present invention is realized as follows: A drug risk monitoring method based on multi-source data fusion includes the following steps: Obtain multi-source data such as electronic medical records of medical institutions, pharmacy sales records, and social media feedback, perform structured processing on the original data through preset format conversion rules to obtain a first data set with unified coding, and then, according to the missing values and outliers in the first data set, use statistical imputation methods and outlier detection algorithms for cleaning and repair to obtain a second data set that has been preprocessed and has consistent fields; Obtain the source attributes of the second dataset, calculate the sample size and update frequency of each data source, determine the credibility weight of each data source, obtain the third dataset with weight tags, then extract features from the third dataset, and convert heterogeneous features into a standardized fourth dataset through a preset feature mapping table; According to the fourth dataset, use the weighted collaborative filtering algorithm to fuse multi-source features. When the weight of a certain data source is lower than the preset threshold, reduce its impact on the fusion result to obtain the fifth dataset containing potential risk signals; Obtain the time series data in the fifth dataset, calculate the change trend of each adverse drug reaction through the sliding window technique to obtain the sixth dataset. According to the abnormal trend in the sixth dataset, use the Bayesian network to perform probability inference on the risk signals to determine high-risk drugs and their quantitative evaluation values, and obtain the seventh dataset; Extract the feature vectors of high-risk drugs from the seventh dataset, perform incremental updates on the new data through the real-time stream processing framework to obtain the dynamically adjusted eighth dataset, and generate the ninth dataset containing drug names and risk levels according to the risk assessment values in the eighth dataset.

[0006] Furthermore, the obtaining of the second dataset that has been preprocessed and has consistent fields specifically includes: Obtain the missing value distribution from the first dataset through statistical methods, use the mean imputation method to fill in the missing values to obtain a preliminary filled dataset, and then identify the abnormal data in the preliminary filled dataset through the isolation forest algorithm to determine the positions of the abnormal points; Obtain the outlier feature from the position of the abnormal point, repair the abnormal data through the median replacement method to obtain the repaired dataset, and then unify the field format through standardization processing according to the field attributes of the repaired dataset to obtain a dataset with consistent format. When there are still outliers in the dataset with consistent format, detect the outlier data through the clustering algorithm to determine whether further repair is required; Obtain the normal data range in the clustering result, use the elimination method to remove the outliers beyond the range to obtain the preprocessed dataset, and determine the field consistency through the integrity check of the preprocessed dataset to obtain the second dataset.

[0007] Furthermore, the obtaining of the third dataset with weight tags specifically includes: Extract the data source and sample size through the source attributes of the second dataset, use statistical methods to calculate the distribution characteristics of the sample size to obtain the sample size indicators of each data source, obtain the frequency level according to the update frequency of the data source. When the frequency level exceeds the preset threshold, it is determined as a high-frequency data source to obtain the classification result of the update frequency. According to the sample size indicators and the update frequency classification, obtain the credibility scores from the pre-established credibility evaluation model to judge the credibility level of each data source; Calculate the weight value using the linear regression algorithm through the credibility level and sample size indicators to obtain a preliminary weight label. Obtain the preliminary weight label and the update frequency classification. When the credibility level matches the frequency level, adjust the weight value to obtain an optimized weight label data set. Then use the clustering algorithm to divide the data source categories to determine the third data set with weight labels. Verify the consistency of the calculation results through the weight labels of the third data set and judge the integrity of the label generation process to obtain the final third data set.

[0008] Furthermore, convert heterogeneous features into a standardized fourth data set through a preset feature mapping table, specifically including: Parse the drug name, adverse reaction description, and timestamp from the third data set, generate a preliminary feature set using the text extraction algorithm, match the heterogeneous features in the preliminary feature set through the preset feature mapping table to obtain a mapped feature set. When there are missing values in the mapped feature set, associate the adverse reaction description through the timestamp and use the content filling algorithm to generate a complete feature set. Compare the complete feature set with the preset table and use the format verification algorithm to judge whether all heterogeneous features are converted into a standardized format to obtain a standardized feature set; Group the drug name and adverse reactions according to the standardized feature set to generate a classification feature set. According to the timestamp in the classification feature set, use the distribution analysis algorithm to calculate the distribution pattern of adverse reactions under different drug names to obtain a distribution feature set. Then use the association analysis algorithm to calculate the association strength among the drug name, adverse reaction, and timestamp to generate the fourth data set.

[0009] Furthermore, obtain a fifth data set containing potential risk signals, specifically including: Obtain multi-source feature data through the fourth data set, calculate the weight values of each data source using the weighted collaborative filtering algorithm to obtain a weight set. Then extract the weight values of each data source according to the weight set. When the weight of a certain data source is lower than the preset threshold, reduce the fusion coefficient of its corresponding feature to obtain an adjusted weight set; Perform weighted fusion processing on the multi-source features using the adjusted weight set to generate a fusion feature set, extract potential risk signals through the fusion feature set to obtain an intermediate feature set containing risk labels; Then obtain the intermediate feature set, analyze the significance of the risk signals using the logistic regression algorithm to obtain a significant risk feature set, and optimize the fusion result according to the significant risk feature set combined with the weighted collaborative filtering algorithm to generate a fifth data set containing potential risks; Analyze the distribution characteristics of the risk signals through the fifth data set using an anomaly detection tool to obtain a risk distribution data set.

[0010] Furthermore, obtain a sixth data set, specifically including: Obtain time series data through the fifth data set, calculate the set of data points in each time period using the preset sliding window technique to obtain the preliminary change trend, and then calculate the sliding window mean and standard deviation of each adverse drug reaction according to the preliminary change trend to obtain the quantified change trend index; According to the quantified change trend index, determine whether the change trend index of a certain drug exceeds the preset threshold. If it exceeds, mark it as abnormal to obtain the preliminary judgment result of the abnormal trend; Extract the corresponding drug list through the preliminary judgment result of the abnormal trend, group the abnormal drugs using the clustering algorithm to obtain the classified set of abnormal drugs, and then obtain the time series characteristics of each group of adverse reactions through the classified set of abnormal drugs to judge the difference in the change trend between groups and obtain the refined analysis result of the abnormal trend; Generate a data set containing time series characteristics and a list of abnormal drugs according to the refined analysis result of the abnormal trend to obtain the sixth data set.

[0011] Furthermore, obtain the seventh data set, specifically including: Extract the time series characteristics of the abnormal trend through the sixth data set, construct a probability model of the risk signal using the Bayesian network to obtain the initial probability distribution data, and then calculate the conditional probability of each adverse drug reaction according to the initial probability distribution data to determine the preliminary intensity of the risk signal and obtain the signal intensity distribution set; Extract the abnormal signals exceeding the preset threshold through the signal intensity distribution set, calculate the risk quantification value of the drug using the weighted average method to obtain the preliminary list of high-risk drugs, and then obtain the corresponding time series data and analyze the significance of the change trend using statistical tools to obtain the set of quantified evaluation values; Judge the risk level of each drug through the set of quantified evaluation values. If the quantified evaluation value exceeds the preset threshold, mark it as a high-risk drug to obtain the classified drug list. Extract the feature vector of the change trend between groups according to the classified drug list and perform secondary probability inference using the Bayesian network to obtain the optimized risk probability distribution; Calculate the final quantified evaluation value of the high-risk drug through the optimized risk probability distribution, generate the seventh data set using the data integration tool to obtain the complete risk assessment data, and then obtain the distribution characteristics of the abnormal trend and verify the data consistency using the loop comparison method to obtain the confirmed set of risk signals; Extract the dynamic characteristics of the high-risk drug through the confirmed set of risk signals and fuse the time series data using the incremental update algorithm to obtain the adjusted risk assessment result.

[0012] Furthermore, obtain the dynamically adjusted eighth data set, specifically including: Extract feature vectors of high-risk drugs from the seventh data set to obtain an initial feature set, and then use a real-time stream processing framework to analyze the newly added data to determine whether it contains high-risk drug information to obtain an updated data subset; The initial feature set is fused through the incremental update algorithm to generate an adjusted feature vector set. When the adjusted feature vector set exceeds the preset threshold, the feature weight is updated through the dynamic adjustment mechanism to obtain the optimized feature set. Based on the optimized feature set, the eighth data set is constructed using the data generation module to complete the dynamic update of the data set.

[0013] Furthermore, a ninth data set containing drug names and risk levels is generated, including: Obtaining the risk assessment value from the eighth data set, determining whether it exceeds a preset threshold, obtaining an exceeding-standard record, and then extracting the name record of the exceeding-standard drug according to the exceeding-standard record to determine the trigger state; Generate a temporary data set containing drug names through the trigger state, obtain the integrity of the temporary data set, use the random forest algorithm to process the temporary data set, extract risk level features, and obtain the level generation result; The ninth data set is constructed through the level generation results and name records, the data integrity is determined, and then the consistency of the risk level in the ninth data set with the assessment of the eighth data set is verified to obtain the verification result; The data association in the ninth data set is updated through the verification results, and the updated data set is determined. Then, combined with the exceeding standard records, data including the drug name and risk level are generated.

[0014] Beneficial effects of the present invention: The present invention successfully solves the problem of data heterogeneity through the collection, structured processing, cleaning and repair, and standardized transformation of multi-source data, breaks the information island, and realizes the efficient integration and unified standardization of multi-source data; it provides a comprehensive and high-quality data foundation for drug risk monitoring, so that it can fully cover the risk characteristics of the entire drug chain, and significantly improve the data coverage and accuracy of drug risk monitoring; The present invention adopts advanced technologies such as weighted collaborative filtering algorithm and Bayesian network to accurately integrate multi-source data features and mine potential risk signals, and dynamically monitor the trend of adverse drug reactions through time series analysis. This process not only improves the accuracy of risk signal mining and avoids the concealment of key risk signals, but also realizes real-time dynamic monitoring and early warning of drug risks, providing strong support for comprehensive monitoring and timely early warning of drug risks, and significantly improving the timeliness and reliability of drug risk monitoring. Incremental updates are performed on newly added data through a real-time stream processing framework, the risk assessment results are dynamically adjusted, and an early warning mechanism is automatically triggered when the risk value exceeds the threshold, achieving real-time dynamic updates and efficient early warnings for risk assessment, being able to promptly reflect changes in drug risks, improving the efficiency and accuracy of drug safety supervision, effectively resolving the contradiction between complex data processing processes and rapid early warning requirements, providing strong technical support for drug safety supervision, and helping to ensure the safety of public drug use. Description of the Drawings

[0015] Figure 1 It is a schematic flow chart of a drug risk monitoring method based on multi-source data fusion provided by this application; Figure 2 It is a schematic flow chart of obtaining a fifth data set containing potential risk signals by a drug risk monitoring method based on multi-source data fusion provided by this application; Figure 3 It is a schematic flow chart of obtaining a seventh data set by a drug risk monitoring method based on multi-source data fusion provided by this application. Detailed Embodiments

[0016] To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0017] As Figures 1-3 shown, this embodiment provides a drug risk monitoring method based on multi-source data fusion, including the following steps: S100. Obtain multi-source data such as medical institution electronic medical records, pharmacy sales records, and social media feedback, perform structured processing on the original data through preset format conversion rules to obtain a first data set with unified coding, and then, according to the missing values and abnormal values in the first data set, use statistical imputation methods and anomaly detection algorithms for cleaning and repair to obtain a second data set that has been preprocessed and has consistent fields; Further, in step S100, obtaining the first data set with unified coding specifically includes: Obtain electronic medical records and sales records through the interfaces of medical institutions and pharmacy systems, collect social feedback through social media platforms to obtain multi-source data, and then perform format conversion on the multi-source data using preset rules to generate raw data with structured processing; apply a unified coding rule to the raw data to obtain a first data set, and then group the first data set through a clustering algorithm to obtain classified data. When there are outliers in the classified data, filter the outliers through a preset threshold to obtain cleaned data, calculate the feature distributions of each group based on the cleaned data to obtain a feature set, and use a decision tree algorithm to analyze the feature set to determine the association patterns between the data and obtain pattern results.

[0018] Specifically, analyze the feature set through a decision tree algorithm to determine the association patterns between the data and obtain pattern results. These association relationships are specifically manifested as the internal connections between different data sources (such as electronic medical records, pharmacy sales records, and social media feedback). For example, the positive correlation between the frequency of occurrence of a certain disease in electronic medical records and the sales volume of the corresponding drug in pharmacies, or the association between negative feedback on a certain drug on social media and the adverse reaction records of the drug in electronic medical records. This association pattern can help medical institutions better understand the association between diseases and treatments, optimize treatment plans; assist pharmaceutical companies in analyzing market demand and drug effects, and adjusting R & D and sales strategies; and at the same time provide data support for public health departments to monitor drug safety and disease epidemic trends, thereby improving the overall decision-making efficiency and service quality of the medical industry.

[0019] Further, in step S100, obtain a second data set that has been preprocessed and has consistent fields, specifically including: Obtain the missing value distribution from the first data set through statistical methods, fill in the missing values using the mean imputation method to obtain a preliminary filled data set, and then identify the abnormal data in the preliminary filled data set through the isolation forest algorithm to determine the positions of the abnormal points; Obtain the outlier feature from the position of the abnormal point, repair the abnormal data through the median replacement method to obtain a repaired data set, and then unify the field format through standardization processing according to the field attributes of the repaired data set to obtain a data set with consistent format. When there are still outliers in the data set with consistent format, detect the outlier data through a clustering algorithm to determine whether further repair is required; Obtain the normal data range in the clustering result, remove the outliers beyond the range through the elimination method to obtain a preprocessed data set, and determine the field consistency through the integrity check of the preprocessed data set to obtain the second data set.

[0020] S200. Calculate the sample size and update frequency of each data source according to the source attributes of the second data set, determine the credibility weight of each data source, obtain the third data set with weight tags, and then extract features such as drug names, adverse reaction descriptions, and timestamps from the third data set, and convert heterogeneous features into a standardized fourth data set through a preset feature mapping table; Further, in step S200, obtaining the third data set with weight tags specifically includes: Extract the data source and sample size through the source attributes of the second data set, calculate the distribution characteristics of the sample size using statistical methods, obtain the sample size indicators of each data source, obtain the frequency level according to the update frequency of the data source, and if the frequency level exceeds a preset threshold, it is determined as a high-frequency data source, obtain the classification result of the update frequency, and obtain the credibility score from a pre-established credibility evaluation model according to the sample size indicator and the update frequency classification, and judge the credibility level of each data source; Through the credibility level and sample size indicators, calculate the weight value using a linear regression algorithm to obtain a preliminary weight tag, obtain the preliminary weight tag and the update frequency classification, and if the credibility level matches the frequency level, adjust the weight value to obtain an optimized weight tag data set, and then use a clustering algorithm to divide the data source categories to determine the third data set with weight tags, verify the consistency of the calculation results through the weight tags of the third data set, and judge the integrity of the tag generation process to obtain the final third data set.

[0021] Specifically, the process of obtaining the preliminary weight tag includes using the credibility level and sample size indicators as independent variables and the credibility score of the data source as the dependent variable to construct a linear regression model. The linear regression model assumes that the credibility score is a linear combination of the credibility level and sample size indicators. By training this model, the weight value of each data source can be obtained. The weight value reflects the contribution degree of the credibility level and sample size indicators to the credibility of the data source. Finally, the calculated weight value is marked on each data source to form the third data set with weight tags.

[0022] Among them, the matching rule between the update frequency and the credibility level is set based on the logical relationship between the update frequency of the data source and the credibility level. Specifically, if the update frequency of the data source is higher than the preset threshold, it is determined as a high-frequency data source. Generally, the credibility level of such data sources will be improved because high-frequency updated data sources can reflect the latest situation in a timely manner, reduce the risk of data obsolescence, and thus improve the credibility of the data. On the contrary, if the update frequency of the data source is lower than the preset threshold, it is determined as a low-frequency data source, and the credibility level of such data sources may need to be comprehensively evaluated through other indicators (such as sample size, historical data quality, etc.); the connection between the update frequency and the credibility lies in that high-frequency updated data sources not only have more advantages in timeliness, but also can provide more complete information, reduce the possibility of data missing, and further enhance the credibility of the data. Therefore, the update frequency is one of the important dimensions for evaluating the credibility of data sources.

[0023] In terms of weight adjustment, first, through the linear regression algorithm, calculate the preliminary weight value of each data source according to the credibility level and the sample size index. Then, match and verify the preliminary weight label with the update frequency classification. If the credibility level matches the update frequency (for example, a high-frequency data source has a high credibility level), it is considered that the weight calculation is reasonable and no adjustment is required. If the credibility level does not match the update frequency (for example, a low-frequency data source has a high credibility level), the weight value needs to be adjusted according to the preset rules to ensure the reasonableness of the weight label. The adjusted weight label forms the third data set with weight labels.

[0024] Furthermore, in step S200, convert the heterogeneous features into a standardized fourth data set through a preset feature mapping table, specifically including: Parse the drug name, adverse reaction description, and timestamp from the third data set, generate a preliminary feature set using a text extraction algorithm, match the heterogeneous features in the preliminary feature set through a preset feature mapping table to obtain the mapped feature set. When there are missing values in the mapped feature set, associate the adverse reaction description through the timestamp, generate a complete feature set using a content filling algorithm, compare the complete feature set with the preset table, and use a format verification algorithm to determine whether all heterogeneous features are converted into a standardized format to obtain a standardized feature set; Group the drug name and adverse reaction according to the standardized feature set to generate a classification feature set. According to the timestamp in the classification feature set, use a distribution analysis algorithm to calculate the distribution pattern of adverse reactions under different drug names to obtain a distribution feature set, and then use an association analysis algorithm to calculate the association strength among the drug name, adverse reaction, and timestamp to generate the fourth data set.

[0025] Exemplarily, the drug name, adverse reaction description and timestamp are parsed from the third data set, and a regular expression-based text extraction algorithm is used to match structured fields such as "drug: aspirin", "adverse reaction: headache", "time: 2023-05-12", etc., to generate a preliminary feature set containing drug name, adverse reaction description and timestamp; for heterogeneous features in the preliminary feature set, such as the presence of two expressions "Aspirin" and "Aspirin" in the drug name, standardized matching is performed through a preset feature mapping table, and "Aspirin" is mapped to "Aspirin", and the mapped feature set; if there are missing values ​​in the mapped feature set, such as a record missing the adverse reaction description, the adverse reaction descriptions with the same timestamp in other records are associated through the timestamp, and the content filling algorithm based on the time window is used to fill the missing content in the nearest neighbor matching method to generate a complete feature set. By comparing the complete feature set with the preset table, the format verification algorithm based on the rule engine is used to check whether the field conforms to standard formats such as "drug name: string", "adverse reaction description: text", and "timestamp: YYYY-MM-DD", and to determine whether all heterogeneous features are converted to obtain a standardized feature set. According to the standardized feature set, the K-means clustering algorithm is used to group the drug names and adverse reactions, and the number of clusters K is set to 5. The similarity is calculated based on the TF-IDF feature to generate a classification feature set; the monthly distribution frequency of adverse reactions under different drug names is statistically analyzed using the time series distribution analysis algorithm through the timestamp in the classification feature set to obtain the distribution feature set; according to the distribution feature set, the Apriori association analysis algorithm is used to calculate the confidence of the drug name and the adverse reaction, such as the confidence of "aspirin → headache" is 85%, to generate the final data set; the cosine similarity algorithm is used to calculate the vector matching degree, and the threshold is set to 0.8 to determine the feature subset of high-risk drugs. According to the feature subset of high-risk drugs, the incremental update algorithm of the Flink stream processing framework is used to integrate the high-risk drug records whose matching degree exceeds the threshold in the newly added data in real time to generate the fourth data set.

[0026] S300, based on the fourth data set, a weighted collaborative filtering algorithm is used to fuse multi-source features, and when the weight of a data source is lower than a preset threshold, its influence on the fusion result is reduced, so as to obtain a fifth data set containing potential risk signals; Furthermore, in step S300, a fifth data set containing potential risk signals is obtained, which specifically includes: S310, obtaining multi-source feature data through the fourth data set, using a weighted collaborative filtering algorithm to calculate the weight value of each data source to obtain a weight set, and then extracting the weight value of each data source according to the weight set. If the weight of a data source is lower than a preset threshold, the fusion coefficient of the corresponding feature is reduced to obtain an adjusted weight set; S320. Use the adjusted weight set to perform weighted fusion processing on multi-source features to generate a fused feature set, extract potential risk signals from the fused feature set, and obtain an intermediate feature set containing risk markers; S330. Then obtain the intermediate feature set, use the logistic regression algorithm to analyze the significance of risk signals to obtain a significant risk feature set, and optimize the fusion result according to the significant risk feature set in combination with the weighted collaborative filtering algorithm to generate a fifth data set containing potential risks; S340. Analyze the distribution characteristics of risk signals through the fifth data set using an anomaly detection tool to obtain a risk distribution data set, and then use a clustering algorithm to divide risk levels to obtain a classified risk data set; extract high-risk feature vectors from the classified risk data set, and combine with a real-time stream processing framework to fuse new data to obtain an updated feature set.

[0027] Exemplarily, extract multi-source feature data such as user ratings, purchase records, and browsing behaviors from the fourth data set, calculate the initial weights of each data source using the Pearson correlation coefficient in the weighted collaborative filtering algorithm, set a threshold of 0.3, and if the weight of social network data is 0.25, reduce its fusion coefficient by 50%. Among them, the Pearson correlation coefficient determines the weight by measuring the linear correlation between two data sources, and its value ranges from -1 to 1. The closer the value is to 1 or -1, the stronger the correlation. Set the threshold to 0.3. If the weight of a certain data source (such as social network data) is lower than this threshold (such as 0.25), it is considered that its correlation is weak and its reliability may be insufficient. Therefore, reduce its fusion coefficient by 50%. This rule of reducing weights adjusts its contribution degree in the fusion process based on the reliability and correlation of the data source to ensure that the fused feature set more accurately reflects user preferences. Perform weighted summation on the features according to the adjusted weight set to generate a fused feature set containing user preferences; detect abnormal rating patterns from the fused features through the Isolation Forest algorithm, mark potential risk users to form an intermediate feature set, and use logistic regression to analyze risk markers, and select features with a p-value less than 0.05 as significant risk features. Input the significant features into an improved collaborative filtering model, use the SGD optimizer to adjust the weights, and output a fifth data set containing high-risk user labels; use the LOF algorithm to calculate the anomaly scores of each sample in the fifth data set to generate a heat map of risk density distribution; based on the distribution data, use K-means clustering (k = 3) to divide high / medium / low risk levels, and extract the cluster center points as feature vectors; when real-time stream data enters, calculate the cosine similarity with the feature vectors through the Flink window, and if it exceeds 0.8, trigger an incremental update of the feature set.

[0028] Specifically, through the weighted collaborative filtering algorithm, the present invention not only considers the credibility of each data source, but also ensures that key risk signals will not be masked by low-quality data sources by dynamically adjusting weights, solving the problem that traditional methods often struggle to balance the credibility and relevance of different data sources when dealing with multi-source data, resulting in the fusion result possibly masking key risk signals. This method significantly improves the accuracy and reliability of the fusion result and can more effectively mine potential risk signals.

[0029] S400. Obtain the time series data in the fifth dataset, calculate the change trends of each adverse drug reaction through the sliding window technique, determine the list of drugs with abnormal trends to obtain the sixth dataset, and perform probability inference on the risk signals using a Bayesian network based on the abnormal trends in the sixth dataset to determine the high-risk drugs and their quantitative evaluation values, obtaining the seventh dataset; Furthermore, in step S400, obtaining the sixth dataset specifically includes: Obtain the time series data through the fifth dataset, calculate the set of data points in each time period using the preset sliding window technique to obtain the preliminary change trend, and then calculate the sliding window mean and standard deviation of each adverse drug reaction based on the preliminary change trend to obtain the quantified change trend index; According to the quantified change trend index, determine whether the change trend index of a certain drug exceeds the preset threshold. If it exceeds, mark it as abnormal to obtain the preliminary judgment result of abnormal trends; Extract the corresponding drug list through the preliminary judgment result of abnormal trends, group the abnormal drugs using a clustering algorithm to obtain the classified set of abnormal drugs, and then obtain the time series characteristics of the adverse reactions in each group through the classified set of abnormal drugs to determine the difference in the change trends between groups, obtaining the refined analysis result of abnormal trends; Generate a dataset containing time series characteristics and a list of abnormal drugs based on the refined analysis result of abnormal trends to obtain the sixth dataset, and calculate the fluctuation amplitude of each abnormal drug using statistical methods through the time series characteristics in the sixth dataset to obtain the trend fluctuation quantification value.

[0030] Exemplarily, time series data of adverse drug reactions are extracted from the fifth dataset. A sliding window with a window size of 30 days is used, and the mean of the data points within each window is calculated with a step size of 7 days to obtain a preliminary change trend. According to the preliminary change trend, the mean and standard deviation of each drug within the sliding window are calculated. If the standard deviation of a certain drug exceeds 2 times the historical mean, it is marked as abnormal. The DBSCAN clustering algorithm is used to group the abnormal drugs, and the neighborhood parameter eps = 0.5 and the minimum number of samples min_samples = 3 are set to obtain a set of classified abnormal drugs. The time series features of each group of drugs are extracted, and the average volatility within the group and the difference degree between groups are calculated. If the difference degree between groups exceeds 0.7, it is determined to be a significant difference; based on the results of the abnormal trend analysis, a sixth dataset containing drug names, time windows, means, standard deviations, and group labels is generated.

[0031] Further, in step S400, a seventh dataset is obtained, specifically including: S410: Extract the time series features of the abnormal trend through the sixth dataset, construct a probability model of risk signals using a Bayesian network to obtain initial probability distribution data, and then calculate the conditional probability of each adverse drug reaction based on the initial probability distribution data to determine the preliminary intensity of the risk signal and obtain a set of signal intensity distributions; S420: Extract abnormal signals exceeding the preset threshold through the set of signal intensity distributions, calculate the risk quantification value of the drug using the weighted average method to obtain a preliminary list of high-risk drugs, then obtain the corresponding time series data, and use statistical tools to analyze the significance of the change trend to obtain a set of quantitative evaluation values; S430: Judge the risk level of each drug through the set of quantitative evaluation values. When the quantitative evaluation value exceeds the preset threshold, it is marked as a high-risk drug to obtain a list of classified drugs. Extract the feature vectors of the change trend between groups according to the list of classified drugs, and perform secondary probability inference using a Bayesian network to obtain an optimized risk probability distribution; S440: Calculate the final quantitative evaluation value of high-risk drugs through the optimized risk probability distribution, use a data integration tool to generate a seventh dataset to obtain complete risk assessment data, then obtain the distribution characteristics of the abnormal trend, and use a cyclic comparison method to verify the data consistency to obtain a confirmed set of risk signals; S450: Extract the dynamic features of high-risk drugs through the confirmed set of risk signals, and use an incremental update algorithm to fuse time series data to obtain an adjusted risk assessment result.

[0032] Exemplarily, time series features of abnormal trends are extracted from the sixth data set, a probability model of risk signals is constructed using a Bayesian network, the prior probability is set to 0.3, and the Markov Chain Monte Carlo method is used for parameter estimation to obtain the initial probability distribution data; the conditional probabilities of each drug adverse reaction are calculated based on the initial probability distribution data, and the Bayesian formula is used in combination with the adverse reaction frequencies in the historical data. For example, P(ADR|Risk) of a certain drug A is 0.85 to determine the preliminary intensity of the risk signal and obtain the signal intensity distribution set; abnormal signals exceeding the preset threshold are extracted from the signal intensity distribution set, the threshold is set to 2 times the standard deviation, and the weighted average method is used to calculate the risk quantification value of the drug, with the weight distribution being 60% for time series volatility and 40% for adverse reaction frequency, to obtain the preliminary high-risk drug list; the corresponding time series data is obtained according to the preliminary high-risk drug list, and the t-test is used to analyze the significance of the change trend, with the p-value < 0.05 set as the significant standard to obtain the set of quantitative evaluation values; the risk levels of each drug are judged through the set of quantitative evaluation values. If the evaluation value of a certain drug B exceeds the threshold of 7.5, it is marked as a high-risk drug to obtain the classified drug list; the eigenvectors of the inter-group change trend are extracted according to the classified drug list, and the K-means clustering algorithm is used to divide the abnormal drugs into 3 groups. Each group is input into the Bayesian network for secondary probability inference, and the learning rate is set to 0.01 to obtain the optimized risk probability distribution; the final quantitative evaluation value of the high-risk drug is calculated through the optimized risk probability distribution, and the SQL database integration tool is used to generate the seventh data set, with the fields including drug ID, risk value, and time stamp to obtain the complete risk assessment data; the distribution characteristics of the abnormal trend are obtained according to the seventh data set, and the hash check algorithm is used to compare the data versions before and after. If the difference rate < 1%, it passes the verification to obtain the confirmed risk signal set; the dynamic characteristics of the high-risk drug are extracted from the confirmed risk signal set, and the online learning algorithm is used to fuse the real-time stream data, with the window size set to 30 days and the sliding step size of 1 day to obtain the adjusted risk assessment result.

[0033] Specifically, the present invention performs probability inference on risk signals through a Bayesian network, which can fully consider various uncertain factors and provide more accurate risk assessment results; in addition, the Bayesian network can also dynamically update the risk assessment model according to new data to ensure the timeliness and accuracy of risk assessment; it solves the problem that traditional methods often lack effective handling of uncertainties during risk assessment, resulting in inaccurate risk assessment results. This method not only improves the scientificity and reliability of risk assessment but also provides strong support for the timely warning of drug risks.

[0034] S500. Extract the feature vectors of high-risk drugs from the seventh dataset, perform incremental updates on the new data through a real-time stream processing framework to obtain a dynamically adjusted eighth dataset, and based on the risk assessment values in the eighth dataset, determine whether the risk value of a certain drug exceeds a preset threshold. When it exceeds, trigger the automatic warning mechanism to generate a ninth dataset containing the drug name and risk level.

[0035] Further, in step S500, to obtain the dynamically adjusted eighth dataset, it specifically includes: Extract the feature vectors of high-risk drugs from the seventh dataset to obtain an initial feature set, and then use a real-time stream processing framework to analyze the new data to determine whether it contains high-risk drug information to obtain an updated data subset; Fuse the initial feature set through an incremental update algorithm to generate an adjusted feature vector set. When the adjusted feature vector set exceeds a preset threshold, update the feature weights through a dynamic adjustment mechanism to obtain an optimized feature set. Then, based on the optimized feature set, use a data generation module to construct the eighth dataset to complete the dynamic update of the dataset.

[0036] Exemplarily, extract the feature vectors of high-risk drugs from the seventh dataset, use the principal component analysis (PCA) algorithm for dimensionality reduction, and extract the first 10 principal components as the initial feature set. Use the real-time stream processing framework Apache Kafka to analyze the new data, and judge whether the number of adverse drug reaction reports of a drug exceeds 50 cases through a rule engine to obtain an updated data subset. For the updated data subset, fuse the initial feature set through the incremental update algorithm OnlinePCA to generate an adjusted feature vector set. If the variance contribution rate of a certain feature in the adjusted feature vector set exceeds 15%, update the feature weights through a dynamic adjustment mechanism, and use the gradient descent method to optimize the weight parameters to obtain an optimized feature set. Based on the optimized feature set, use a data generation module to construct the eighth dataset, and use the HDFS storage format to complete the dynamic update of the dataset. After obtaining the eighth dataset, detect the change trend of high-risk drugs through a risk identification process, and use the time series analysis algorithm ARIMA to predict the risk distribution in the next 30 days to obtain the risk distribution result.

[0037] Further, in step S500, to generate a ninth dataset containing the drug name and risk level, it specifically includes: Obtain the risk assessment value from the eighth dataset, determine whether it exceeds the preset threshold to obtain an over-standard record, and then extract the name record of the over-standard drug according to the over-standard record to determine the trigger status; Generate a temporary dataset containing the drug name through the trigger status, obtain the integrity of the temporary dataset, use the random forest algorithm to process the temporary dataset, and extract the risk level features to obtain the level generation result; Construct the ninth dataset by generating result and name records at different levels, judge the data integrity, and then verify the consistency of the evaluation exceeding the standard with the eighth dataset according to the risk levels in the ninth dataset to obtain the verification result; Update the data association in the ninth dataset based on the verification result to determine the updated dataset, and then combine the records of exceeding the standard to generate data containing drug names and risk levels.

[0038] Exemplarily, extract the drug risk assessment values from the eighth dataset. If the risk value of a certain drug exceeds the preset threshold of 0.8, it is marked as a record of exceeding the standard. According to the drug ID field in the record of exceeding the standard, match the drug information table to obtain the corresponding drug name, generate a name record and set the trigger status to "warning", store the name record in the temporary dataset, and check whether there are null values or duplicate items in the temporary dataset to ensure data integrity. Analyze the temporary dataset using the random forest algorithm, set the number of decision trees to 100 and the maximum depth to 10, extract the risk level features through feature importance ranking, and output the classification results of high, medium, and low levels; combine the classification results with the drug name records to construct the ninth dataset, and check the field matching and the proportion of missing values. Compare the risk levels in the ninth dataset with the original evaluation values in the eighth dataset. If the original evaluation values of high-risk drugs are all greater than 0.8, the verification passes; correct the drug-level mapping relationship in the ninth dataset according to the verification result and update the data association index. Integrate the updated dataset and the records of exceeding the standard to generate a structured data table containing drug names (such as "Drug A") and risk levels (such as "high risk"); process the final dataset using the K-means clustering algorithm, set the number of clusters to 3, divide the grade intervals based on the risk value distribution, and output the grading results.

[0039] The specific embodiments of the invention have been described in detail above, but they are only examples. The invention is not limited to the specific embodiments described above. Those skilled in the art of this industry should understand that the above embodiments and the descriptions in the specification only illustrate the principles of the invention. Without departing from the spirit and scope of the invention, the invention will have various changes and improvements, and these changes and improvements all fall within the scope of the invention claimed. The scope of the invention claimed is defined by the appended claims and their equivalents.

Claims

1. A drug risk monitoring method based on multi-source data fusion, characterized in that: It includes the following steps: Obtain multi-source data, perform structured processing on the original data through preset format conversion rules to obtain a first data set with unified encoding, and then, according to the missing values and outliers in the first data set, use statistical imputation methods and outlier detection algorithms for cleaning and repair to obtain a second data set that has been preprocessed and has consistent fields; Obtain the source attributes of the second data set, calculate the sample size and update frequency of each data source, determine the credibility weights of each data source to obtain a third data set with weight tags, and then extract features from the third data set and convert heterogeneous features into a standardized fourth data set through a preset feature mapping table; According to the fourth data set, use a weighted collaborative filtering algorithm to fuse multi-source features. When the weight of a certain data source is lower than a preset threshold, reduce its impact on the fusion result to obtain a fifth data set containing potential risk signals; Obtain the time series data in the fifth data set, calculate the change trends of each drug adverse reaction through a sliding window technique to obtain a sixth data set, and according to the abnormal trends in the sixth data set, use a Bayesian network to perform probability inference on the risk signals to determine high-risk drugs and their quantitative evaluation values to obtain a seventh data set; Extract the feature vectors of high-risk drugs from the seventh data set, perform incremental updates on new data through a real-time stream processing framework to obtain a dynamically adjusted eighth data set, and generate a ninth data set containing drug names and risk levels according to the risk assessment values in the eighth data set.

2. The method for monitoring drug risks based on multi-source data fusion according to claim 1, wherein: The obtaining of the second data set that has been preprocessed and has consistent fields specifically includes: Obtain the missing value distribution from the first data set through statistical methods, use the mean imputation method to fill in the missing values to obtain a preliminary filled data set, and then identify the abnormal data in the preliminary filled data set through the Isolation Forest algorithm to determine the positions of the abnormal points; Obtain the outlier feature from the positions of the abnormal points, repair the abnormal data through the median replacement method to obtain a repaired data set, and then, according to the field attributes of the repaired data set, use standardization processing to unify the field format to obtain a data set with consistent format. When there are still outliers in the data set with consistent format, detect the outlier data through a clustering algorithm to determine whether further repair is required; Obtain the normal data range in the clustering result, use the elimination method to remove the outliers outside the range to obtain a preprocessed data set, and determine the field consistency through the integrity check of the preprocessed data set to obtain the second data set.

3. The method for monitoring drug risks based on multi-source data fusion according to claim 1, characterized in that: The obtaining of the third data set with weight tags specifically includes: Extract the data source and sample size through the source attributes of the second data set, use statistical methods to calculate the distribution characteristics of the sample size to obtain the sample size indicators of each data source, obtain the frequency level according to the update frequency of the data source, and when the frequency level exceeds a preset threshold, it is determined as a high-frequency data source to obtain the classification result of the update frequency. According to the sample size indicators and the update frequency classification, obtain the credibility scores from a pre-established credibility evaluation model to judge the credibility levels of each data source; Calculate the weight value using the linear regression algorithm through the credibility level and sample size metrics to obtain a preliminary weight label. Obtain the preliminary weight label and the update frequency classification. When the credibility level matches the frequency level, adjust the weight value to obtain an optimized weight label dataset. Then use the clustering algorithm to classify the data source categories and determine the third dataset with weight labels. Verify the consistency of the calculation results through the weight labels of the third dataset and judge the integrity of the label generation process to obtain the final third dataset.

4. A method for monitoring drug risks based on multi-source data fusion according to claim 1, characterized in that: The conversion of heterogeneous features into a standardized fourth dataset through a preset feature mapping table specifically includes: Parse the drug name, adverse reaction description, and timestamp from the third dataset. Use the text extraction algorithm to generate a preliminary feature set. Match the heterogeneous features in the preliminary feature set through the preset feature mapping table to obtain a mapped feature set. When there are missing values in the mapped feature set, associate the adverse reaction description with the timestamp and use the content filling algorithm to generate a complete feature set. Compare the complete feature set with the preset table and use the format verification algorithm to judge whether all heterogeneous features have been converted into a standardized format to obtain a standardized feature set; Group the drug name and adverse reactions according to the standardized feature set to generate a classification feature set. According to the timestamp in the classification feature set, use the distribution analysis algorithm to calculate the distribution pattern of adverse reactions under different drug names to obtain a distribution feature set. Then use the association analysis algorithm to calculate the association strength among the drug name, adverse reaction, and timestamp to generate the fourth dataset.

5. A drug risk monitoring method based on multi-source data fusion according to claim 1, characterized in that: The obtaining of the fifth dataset containing potential risk signals specifically includes: Obtain multi-source feature data from the fourth dataset. Use the weighted collaborative filtering algorithm to calculate the weight values of each data source to obtain a weight set. Then extract the weight values of each data source according to the weight set. When the weight of a certain data source is lower than the preset threshold, reduce the fusion coefficient of its corresponding feature to obtain an adjusted weight set; Perform weighted fusion processing on the multi-source features using the adjusted weight set to generate a fusion feature set. Extract potential risk signals from the fusion feature set to obtain an intermediate feature set with risk labels; Then obtain the intermediate feature set, use the logistic regression algorithm to analyze the significance of the risk signals to obtain a significant risk feature set, and optimize the fusion result according to the significant risk feature set combined with the weighted collaborative filtering algorithm to generate a fifth dataset containing potential risks; Use an anomaly detection tool to analyze the distribution characteristics of the risk signals through the fifth dataset to obtain a risk distribution dataset.

6. The method for monitoring drug risks based on multi-source data fusion according to claim 1, wherein: The obtaining of the sixth dataset specifically includes: Obtain time series data from the fifth dataset. Use the preset sliding window technique to calculate the set of data points in each time period to obtain a preliminary change trend. Then, according to the preliminary change trend, calculate the sliding window mean and standard deviation of each drug adverse reaction to obtain a quantified change trend index; According to the quantified change trend index, judge whether the change trend index of a certain drug exceeds the preset threshold. If it exceeds, mark it as abnormal to obtain a preliminary judgment result of trend abnormality; Extract the corresponding drug list according to the preliminary judgment results of trend anomalies, group the abnormal drugs using a clustering algorithm to obtain the classified set of abnormal drugs, then obtain the time series characteristics of adverse reactions in each group from the classified set of abnormal drugs, judge the differences in the change trends between groups, and obtain the refined analysis results of abnormal trends; Generate a data set containing time series characteristics and a list of abnormal drugs according to the refined analysis results of abnormal trends to obtain the sixth data set.

7. A drug risk monitoring method based on multi-source data fusion according to claim 1, characterized in that: The obtaining of the seventh data set specifically includes: Extract the time series characteristics of abnormal trends from the sixth data set, construct a probability model of risk signals using a Bayesian network to obtain the initial probability distribution data, then calculate the conditional probabilities of adverse reactions of each drug according to the initial probability distribution data to determine the preliminary intensity of risk signals and obtain the signal intensity distribution set; Extract abnormal signals exceeding the preset threshold from the signal intensity distribution set, calculate the risk quantification value of drugs using the weighted average method to obtain a preliminary list of high-risk drugs, then obtain the corresponding time series data, and analyze the significance of the change trend using statistical tools to obtain the set of quantification evaluation values; Judge the risk levels of each drug through the set of quantification evaluation values. When the quantification evaluation value exceeds the preset threshold, mark it as a high-risk drug to obtain the classified drug list. Extract the feature vectors of the change trends between groups from the classified drug list and perform secondary probability reasoning using a Bayesian network to obtain the optimized risk probability distribution; Calculate the final quantification evaluation value of high-risk drugs through the optimized risk probability distribution, generate the seventh data set using a data integration tool to obtain the complete risk assessment data, then obtain the distribution characteristics of abnormal trends, and verify the data consistency using a loop comparison method to obtain the confirmed set of risk signals; Extract the dynamic characteristics of high-risk drugs from the confirmed set of risk signals and fuse the time series data using an incremental update algorithm to obtain the adjusted risk assessment results.

8. A drug risk monitoring method based on multi-source data fusion according to claim 1, characterized in that: The obtaining of the dynamically adjusted eighth data set specifically includes: Extract the feature vectors of high-risk drugs from the seventh data set to obtain the initial feature set, then analyze the new data using a real-time stream processing framework to judge whether it contains high-risk drug information to obtain the updated data subset; Fuse the initial feature set using an incremental update algorithm to generate the adjusted feature vector set. When the adjusted feature vector set exceeds the preset threshold, update the feature weights through a dynamic adjustment mechanism to obtain the optimized feature set, and then construct the eighth data set using a data generation module according to the optimized feature set to complete the dynamic update of the data set.

9. A method for monitoring drug risks based on multi-source data fusion according to claim 1, characterized in that: The generation of the ninth data set containing drug names and risk levels specifically includes: Obtain the risk assessment value from the eighth data set, judge whether it exceeds the preset threshold to obtain the exceeded records, then extract the name records of the exceeded drugs according to the exceeded records to determine the trigger status; Generate a temporary data set containing drug names through the trigger status, obtain the integrity of the temporary data set, process the temporary data set using a random forest algorithm, and extract the risk level characteristics to obtain the level generation result; Construct the ninth dataset by generating result and name records at different levels, judge data integrity, and then verify the consistency of evaluation exceeding the standard with the eighth dataset based on the risk levels in the ninth dataset to obtain the verification result; Update the data association in the ninth dataset based on the verification result, determine the updated dataset, and then combine the records of exceeding the standard to generate data containing drug names and risk levels.

Citation Information

Patent Citations

  • Intelligent identification and prevention system for adverse drug reaction

    CN118053541A

  • Ultra-short-term wind power prediction method based on multi-source data fusion

    CN119965840A

  • Methods to Assess Clinical Outcome Based Upon Updated Probabilities and Treatments Thereof

    US20220392605A1

Cited By

  • Multi-source data intelligent analysis management method and system based on big data

    CN121144267A

  • A big data-based multi-source data intelligent analysis management method and system

    CN121144267B

  • Hazardous chemical substance risk operation behavior early warning method and system

    CN121169083A

  • Drug full life cycle risk monitoring and early warning method based on multi-source data fusion

    CN121215308A

  • A Drug Lifecycle Risk Monitoring and Early Warning Method Based on Multi-Source Data Fusion

    CN121215308B