Data management method and system for drug alert personnel
Through multidimensional quality assessment and risk scoring models, combined with drug type, data source characteristics and timeliness, blockchain resource allocation is dynamically adjusted, which solves the shortcomings of data integration and risk grading in the drug vigilance system and achieves more efficient data management and risk response.
Patent Information
- Application Number
- CN202510800193.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-26
AI Technical Summary
Existing pharmacovigilance systems have significant defects in data integration, quality assessment and risk grading, especially information loss or semantic bias when processing multi-source heterogeneous data. Traditional methods find it difficult to capture the complex nonlinear relationships between drugs, patients and adverse reactions.
A multidimensional quality assessment strategy and a multidimensional risk scoring model are adopted, combined with drug type, data source characteristics and timeliness, through nonlinear exponential weighting functions and dynamic weights, to perform data quality assessment and risk scoring, and dynamically adjust blockchain resource allocation based on the score.
It improves the accuracy of quality assessment and risk level classification of drug vigilance data management, ensures the priority processing of high-risk data and the adaptability of resource allocation, and enhances the data management efficiency and risk response capabilities of the drug vigilance system.
Smart Images

Figure CN120708927A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of big data processing, and in particular relates to a data management method and system for drug vigilance personnel. Background Art
[0002] Pharmacovigilance, a core component of ensuring drug safety, is responsible for monitoring, evaluating, and preventing adverse drug reactions. With the rapid development of the pharmaceutical industry, the volume of data that pharmacovigilance systems must process has exploded, encompassing heterogeneous data from multiple sources, including spontaneous reports, clinical trials, literature reviews, and social media. However, existing technologies present significant shortcomings in practical applications, particularly in data integration, quality assessment, and risk stratification.
[0003] Specifically, existing methods typically use unified standards to process heterogeneous data from different sources (such as structured clinical trial data and unstructured social media text), but lack dynamic adaptation mechanisms. For example, non-standard terms in social media data (such as "headache" and "head discomfort") are difficult to accurately map to standard medical dictionaries (such as MedDRA), resulting in information loss or semantic bias. This problem further exacerbates the complexity of data integration and reduces the credibility of subsequent analysis results.
[0004] Furthermore, traditional quality assessment strategies often focus on a single dimension (such as field completeness), neglecting to analyze the differentiated characteristics of data sources (e.g., hospital data has high completeness but low timeliness, while social media data has high timeliness but low accuracy). For example, anticancer drug data, due to the high severity of adverse reactions, requires a higher accuracy weight. However, existing methods fail to dynamically adjust assessment criteria based on drug type, resulting in high-risk data issues (such as missing key laboratory indicators) not being prioritized for identification.
[0005] Furthermore, existing risk assessment models generally calculate risk scores using simple linear methods, but these models struggle to capture the complex, nonlinear relationships between drugs, patients, and adverse reactions. For example, elderly patients may have an exponentially increasing sensitivity to a particular drug, and linear models cannot effectively quantify this risk. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a data management method and system for pharmacovigilance personnel to improve the quality of pharmacovigilance data management.
[0007] In a first aspect, the present invention provides a data management method for pharmacovigilance personnel, the method comprising the following steps:
[0008] Acquire multi-source heterogeneous pharmacovigilance data and perform de-heterogeneous processing on the data based on the characteristics of the data sources to obtain standardized pharmacovigilance data;
[0009] A multidimensional quality assessment strategy is used to quantitatively score standard pharmacovigilance data and obtain the quality assessment results corresponding to the standard pharmacovigilance data. The multidimensional quality assessment strategy is based on the drug type and data source type in the standard pharmacovigilance data, and the quality assessment results are high or low quality.
[0010] For standard drug vigilance data with a low-quality quality assessment result, a pre-trained multidimensional risk scoring model is used to calculate the risk score corresponding to the standard drug vigilance data, and the risk level of the standard drug vigilance data is divided according to the risk score, and the blockchain resource allocation is dynamically adjusted according to the risk level; the multidimensional risk scoring model is based on the severity of adverse drug reactions in the drug vigilance data, the sensitivity of the patient group, and the urgency of time, and the risk level is high, medium, or low.
[0011] Optionally, multiple dimensions include completeness, accuracy, consistency, and timeliness.
[0012] Optionally, a multidimensional quality assessment strategy is used to quantitatively score the standard pharmacovigilance data to obtain the quality assessment results corresponding to the standard pharmacovigilance data, including:
[0013] The weights corresponding to each dimension of the standard pharmacovigilance data are calculated based on the drug types and data source characteristics included in the data. Drug types include at least anticancer drugs, vaccines, and common drugs. The degree of adverse reactions varies between drug types and is obtained through normalization using expert data. The data source characteristics include volatility. The greater the field missing rate of the data source, the greater the volatility. The weight is positively correlated with the degree of adverse reactions and negatively correlated with volatility.
[0014] The quality score of the standard pharmacovigilance data is calculated by combining the nonlinear exponential weighting function and weights. The quality score is used to measure the quality of the standard pharmacovigilance data.
[0015] The quality assessment result is determined based on the comparison result of the quality score with the preset quality threshold; when the quality score is greater than or equal to the preset quality threshold, the quality assessment result is high quality; otherwise, the quality assessment result is low quality.
[0016] Optionally, standard pharmacovigilance data may be classified into risk levels based on risk scores, including:
[0017] A first risk threshold and a second risk threshold are respectively constructed based on the historical risk score distribution; the first risk threshold is greater than the second risk threshold;
[0018] The risk level is determined based on the relationship between the risk score, the first risk threshold, and the second risk threshold.
[0019] Optionally, the risk level is determined based on the relationship between the risk score, the first risk threshold, and the second risk threshold, including:
[0020] If the risk score is greater than or equal to the first risk threshold, the risk level of the standard pharmacovigilance data is determined to be high;
[0021] If the risk score is less than the first risk threshold and greater than or equal to the second risk threshold, the risk level of the standard pharmacovigilance data is determined to be medium;
[0022] If the risk score is less than the second risk threshold, the risk level of the standard pharmacovigilance data is determined to be low.
[0023] Optionally, dynamically adjust blockchain resource allocation based on risk level, including:
[0024] Calculate the blockchain resource allocation intensity based on risk level, adverse reaction severity, and volatility; blockchain resource allocation intensity represents the processing priority of standard pharmacovigilance data;
[0025] Dynamically adjust blockchain resource allocation according to the intensity of blockchain resource allocation.
[0026] Optionally, blockchain resources include the number of blockchain nodes allocated and the number of data shards.
[0027] Optionally, there is a nonlinear positive correlation between the number of blockchain nodes allocated and the intensity of blockchain resource allocation;
[0028] There is a linear positive correlation between the number of data shards and the intensity of blockchain resource allocation.
[0029] Optionally, after classifying the risk levels of the standard pharmacovigilance data according to the risk scores and dynamically adjusting the allocation of blockchain resources according to the risk levels, the method further includes:
[0030] Trigger data completion for standard drug vigilance data based on risk levels and establish a data modification traceability mechanism.
[0031] In a second aspect, the present invention provides a data management system for pharmacovigilance personnel, comprising:
[0032] A preprocessing module is used to obtain multi-source heterogeneous pharmacovigilance data and perform de-isomerization processing on the multi-source heterogeneous pharmacovigilance data based on the characteristics of the data sources to obtain standard pharmacovigilance data;
[0033] The quality assessment module is used to quantitatively score the standard pharmacovigilance data using a multidimensional quality assessment strategy to obtain the quality assessment results corresponding to the standard pharmacovigilance data. The multidimensional quality assessment strategy is based on the drug type and data source type in the standard pharmacovigilance data, and the quality assessment results are high or low quality.
[0034] The blockchain resource allocation module is used to calculate the risk score corresponding to standard drug vigilance data with a low-quality quality assessment result using a pre-trained multidimensional risk scoring model, divide the risk level of the standard drug vigilance data according to the risk score, and dynamically adjust the blockchain resource allocation according to the risk level; the multidimensional risk scoring model is based on the severity of adverse drug reactions in the drug vigilance data, the sensitivity of the patient group, and the urgency of time, and the risk level is high, medium, or low.
[0035] The beneficial effects of the present invention are:
[0036] The data management method for pharmacovigilance personnel provided by the present invention utilizes a multidimensional quality assessment strategy to quantitatively score standard pharmacovigilance data. Based on the drug type and data source type in the standard pharmacovigilance data, the quality of the standard pharmacovigilance data is assessed from multiple dimensions. The method can dynamically adapt to the drug type and data source type, avoid misjudgment caused by single-dimensional defects, and help improve the accuracy of quality assessment, thereby improving the quality of pharmacovigilance data management. The method also scores the data based on the severity of adverse drug reactions, patient group sensitivity, and time urgency in the pharmacovigilance data, combining the correlation between drugs, patients, and time, which helps improve the accuracy of risk level classification, thereby improving the quality of pharmacovigilance data management. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is a flow chart of a data management method for pharmacovigilance personnel in one embodiment of the present application;
[0038] Figure 2 This is a structural diagram of the data management system for pharmacovigilance personnel in one embodiment of the present application. DETAILED DESCRIPTION
[0039] The following will describe the technical solutions in the embodiments of the present invention in detail with reference to the accompanying drawings. The described embodiments are only a part of the embodiments of the present invention.
[0040] The present invention discloses a data management method for drug vigilance personnel, such as Figure 1 As shown, the method includes the following steps:
[0041] Step 11: Acquire multi-source heterogeneous pharmacovigilance data, and perform de-heterogeneous processing on the multi-source heterogeneous pharmacovigilance data based on the characteristics of the data sources to obtain standard pharmacovigilance data.
[0042] The aforementioned multi-source heterogeneous pharmacovigilance data refers to pharmacovigilance data from different sources. In one embodiment of the present invention, the data sources include spontaneous reports, clinical trials, literature reports, and social media. Spontaneous reports contain patient-reported adverse reaction descriptions, clinical trials provide structured laboratory data, literature reports are semi-structured text, and social media are unstructured user comments.
[0043] In a feasible embodiment, performing de-heterogeneous processing on multi-source heterogeneous pharmacovigilance data based on characteristics of the data sources includes:
[0044] Using a pre-set mapping rule table, heterogeneous data from spontaneous reports, clinical trials, literature reports, and social media is collected. Fields in this heterogeneous data are mapped to standard fields using a field matching algorithm to generate standardized pharmacovigilance data. The mapping rule table can be designed to include three columns: source type, original field, and standard field. For example, "Adverse Reaction Description" in a spontaneous report is mapped to the standard field "AE_DESC." The field matching algorithm can employ a hybrid approach based on string similarity and semantic analysis.
[0045] In another feasible embodiment, performing de-heterogeneous processing on multi-source heterogeneous pharmacovigilance data based on characteristics of the data sources includes:
[0046] Based on the source characteristics of heterogeneous data, a pre-set quality checkpoint configuration table is obtained (the configuration table contains the fields and characteristics that require attention during the data cleansing and verification process, such as field integrity, data type, and value range). Integrity checks and accuracy verification are performed on the standard format data after preliminary mapping. If missing fields or outliers are detected, they are marked as data to be cleaned. Using a cleaning logic library (which contains cleaning rules for different data sources and data types, and selects appropriate cleaning rules based on the data characteristics of different sources. For example, some fields may need to fill missing values, while some outliers may need to be replaced with legal default values), marker fields are extracted from the data to be cleaned (to indicate which fields require cleaning). The corresponding cleaning rules are selected based on the source characteristics. Automated scripts are used to fill missing fields or correct outliers to generate cleaned standard format data. Using consistency verification tools (such as Informatica), the cleaned standard format data is obtained and the field type and value range are determined to determine whether they meet the preset standardization requirements. If not, format conversion is performed to obtain standard pharmacovigilance data.
[0047] Step 12: Use a multidimensional quality assessment strategy to quantitatively score the standard drug vigilance data to obtain a quality assessment result corresponding to the standard drug vigilance data.
[0048] In an embodiment of the present invention, the multidimensional quality assessment strategy assesses the quality of standard pharmacovigilance data from multiple dimensions based on the drug type and data source type within the data, with the quality assessment result being high or low quality. The multiple dimensions include completeness, accuracy, consistency, and timeliness.
[0049] In one feasible embodiment, a multidimensional quality assessment strategy is used to quantitatively score the standard pharmacovigilance data to obtain a quality assessment result corresponding to the standard pharmacovigilance data, including steps 12.1 to 12.3.
[0050] Step 12.1: Calculate the weights corresponding to each dimension of the standard pharmacovigilance data based on the drug types and data source characteristics contained in the standard pharmacovigilance data.
[0051] It should be noted that traditional methods usually use a uniform weight for all drug types, which may ignore the strict data quality requirements of high-risk drugs (such as anti-cancer drugs); at the same time, traditional methods do not take into account the volatility of data sources, which may lead to misjudgment of highly volatile data such as social media.
[0052] To improve the accuracy and adaptability of quality assessment, the present invention calculates the weights corresponding to each dimension of the standard pharmacovigilance data based on the drug types contained in the data and the characteristics of the data source to which it belongs. Through a multidimensional quality assessment strategy, the weights of the four dimensions of completeness, accuracy, consistency, and timeliness are dynamically adjusted. Combined with a nonlinear exponential weighting function and a risk scoring model, accurate quantification of data quality is achieved.
[0053] Specifically, drug types include at least anticancer drug types, vaccine types, and common drug types. The degree of adverse reactions of different drug types is different. The degree of adverse reactions is obtained through expert corpus and normalized. The characteristics of the data source include volatility. The greater the field missing rate of the data source, the greater the volatility. The weight is positively correlated with the degree of adverse reactions and negatively correlated with volatility.
[0054] Exemplarily, calculating the weights corresponding to each dimension of the standard pharmacovigilance data includes:
[0055] By calculating the formula
[0056]
[0057] Get the dynamic weight w of dimension d d ; where d∈{completeness, accuracy, consistency, timeliness}, μ drepresents the pre-set benchmark weight, α d It represents the drug risk sensitivity coefficient, which is determined according to the drug regulatory classification standard. The value of anticancer drugs is 8 to 9, the value of vaccine types is 9 to 10, and the value of common drugs is 3 to 7. k It represents the risk level of drug type k, which is related to the degree of adverse reactions. According to the food and drug supervision standards, the risk level of anticancer drugs is normalized to 0.8 to 0.9, the risk level of vaccines is normalized to 0.9 to 1, and the risk level of common drugs is normalized to 0.3 to 0.7. max Indicates the maximum value of adverse reaction (e.g. 1), β d represents the data source stability coefficient, D s Indicates the volatility of data source s, calculated by the standard deviation of the historical missing rate, α d ∈[0,1],β d ∈[0,1].
[0058] Notably, weights are adjusted based on the risk profile of drug types. Key dimensions (such as completeness and timeliness) for high-risk drugs (e.g., anticancer drugs and vaccines) are prioritized to ensure accurate capture of serious adverse reaction data. Volatility parameters are used to mitigate misjudgment of noisy data, taking into account the varying stability of different data sources, such as social media and clinical trials. For example, high missing rates in social media fields automatically reduce their weight to prevent low-quality data from excessively impacting assessment results. Low-quality data is categorized using a risk scoring model (combining adverse reaction severity, patient sensitivity, and timeliness), and blockchain resources are dynamically adjusted. High-risk data (e.g., vaccine adverse reaction reports) receives a higher number of nodes and data shards, ensuring processing priority and security. A nonlinear weighting function significantly amplifies the impact of low defect rates. For example, when the integrity defect rate of a particular anticancer drug's data is only 5%, the quality score plummets, triggering prioritized processing and preventing potential risks from being missed. This approach addresses issues such as the lag in identifying high-risk data and biased assessments of heterogeneous data sources caused by traditional static weighting, significantly improving the data management efficiency and risk response capabilities of the pharmacovigilance system.
[0059] In step 12.2, the quality score of the standard pharmacovigilance data is calculated by combining the nonlinear exponential weighting function and the weights.
[0060] Aiming at the problem that traditional methods are insufficiently sensitive to low-quality data and ignore the limitations of differences between drugs and data sources, the present invention solves this problem by combining a nonlinear exponential weighting function with a dynamic weight.
[0061] Specifically, by calculating the formula
[0062]
[0063] The quality score Q is obtained to measure the quality of standard pharmacovigilance data; where γ d represents the defect rate nonlinear gain coefficient, γ d ∈[0,1],σ d represents the defect rate of dimension d, σ d,max represents the maximum defect rate of dimension d, ρ dj Represents the dimension correlation coefficient, obtained based on expert experience, d≠j.
[0064] It should be noted that the exponential function It has a nonlinear amplification effect. When the defect rate σ d When the defect rate is low (such as 5%), the slope of the exponential function is large, which makes the quality score Q more sensitive to the change of defect rate. d When the defect rate is high (e.g., more than 30%), the function tends to saturate, avoiding excessive punishment. Under the linear method, an increase in the defect rate from 5% to 10% only leads to a linear decrease in the quality score (e.g., 95→90), while the exponential function can drop the quality score from 85 to 60, significantly amplifying the impact of low defect rates. For example, when the defect rate of a vaccine data integrity is only 3%, the quality score has been greatly increased due to the high weight w. d and exponential functions are significantly lowered, triggering priority processing.
[0065] Moreover, by combining nonlinear exponential weighting functions and weights, the quality score of standard pharmacovigilance data can be calculated to adapt to differences in drug types and data sources: dynamic weight w d Adjust according to the drug risk sensitivity coefficient (such as 8-9 for anticancer drugs and 9-10 for vaccines) and the volatility of the data source (such as the attenuation weight of high volatility of social media). For example, the integrity weight of vaccine data is w d The high-risk sensitivity coefficient is significantly increased, while the weight of social media data is attenuated due to its high volatility, reducing the risk of misjudgment of low-quality data.
[0066] It is worth mentioning that, by combining nonlinear exponential weighting functions and weights, calculating the quality score of standard drug vigilance data can accurately identify low-quality data defects (the nonlinear exponential function significantly amplifies the impact of low defect rates, avoiding the misjudgment of "average pass" under traditional linear weighting), dynamically adapt to differentiated needs (the weights of key dimensions of high-risk drugs such as vaccines (such as integrity) are increased to ensure the accurate capture of serious adverse reaction data; the weight allocation of common drugs is more flexible, saving resources; the weight of high-volatility data sources such as social media is attenuated to reduce the probability of misjudgment of low-quality data), and resource allocation optimization and risk control (the quality score is directly related to the subsequent risk scoring model and blockchain resource allocation. The risk score generated by low-quality data after nonlinear weighting is more accurate, and high-risk data (such as vaccine adverse reaction reports) can obtain a higher number of node allocations and data shards to ensure processing priority and security).
[0067] This approach, through the combination of nonlinear exponential weighting functions and dynamic weights, addresses the limitations of traditional methods, such as insufficient sensitivity to low-quality data and neglect of differences between drugs and data sources. It significantly improves the accuracy and adaptability of quality assessment and provides technical support for the efficient management of drug vigilance systems.
[0068] Step 12.3: Determine the quality assessment result based on the comparison result of the quality score with the preset quality threshold.
[0069] Specifically, when the quality score is greater than or equal to a preset quality threshold, the quality assessment result is high quality; otherwise, the quality assessment result is low quality.
[0070] It should be noted that, in the embodiment of the present invention, high-quality standard drug vigilance data is allocated blockchain resources according to preset standard configurations.
[0071] In another feasible embodiment, the quality threshold may be dynamically adjusted based on a distribution drift perception mechanism of the Wasserstein distance.
[0072] Step 13: For the standard drug vigilance data with a low quality assessment result, a pre-trained multidimensional risk scoring model is used to calculate the risk score corresponding to the standard drug vigilance data, and the risk level of the standard drug vigilance data is divided according to the risk score, and the blockchain resource allocation is dynamically adjusted according to the risk level.
[0073] In an embodiment of the present invention, the multidimensional risk scoring model is based on the severity of adverse drug reactions in the pharmacovigilance data, the sensitivity of the patient population, and the urgency of time, and the risk level is high, medium, or low.
[0074] In a feasible embodiment, the expression of the multidimensional risk scoring model is:
[0075]
[0076] Among them, R represents the risk score, S α It represents the weighted adverse reaction severity coefficient. The greater the severity of the adverse reaction in the standard pharmacovigilance data, the higher the S α The larger the P β It represents the weighted patient group sensitivity coefficient. The greater the sensitivity of the patient group, the higher the P β The larger the patient population, the more common the patient population, including children, pregnant women and the general population. γ represents the weighted time urgency coefficient, δ represents the correction index, which is a constant, α, β, are dimension weight coefficients, both are constants, D low Indicates the proportion of low-quality data, D totalRepresents the total amount of data, S,P,U∈[0,1].
[0077] It should be noted that a nonlinear function is used to construct a risk scoring system by integrating the three core dimensions of severity of adverse drug reactions (S), sensitivity of the patient population (P), and time urgency (U). Among them, α, β, and γ are determined through historical data training, which can amplify the marginal effects of highly sensitive parameters (such as the exponential risk of adverse reactions to vaccines). For example, for anticancer drug data, the model exponentially amplifies the risk of grade 3 liver injury adverse reactions by increasing the weight coefficient of S (α = 1.8), which is significantly different from the simple weighting of the linear model.
[0078] It's worth noting that this approach effectively compensates for the traditional linear model's inadequate response to the exponential growth of drug sensitivity in elderly patients. In one implementation, the sensitivity score for patients over 70 years old jumped from 0.7 in the linear model to 0.92 in the nonlinear model, accurately identifying the exponential risk of elderly patients and effectively capturing nonlinear risk. When a vaccine data set exhibits a time lag (U = 0.3) but severe adverse reactions (S = 0.9), the model ensures that the overall risk score still reaches the advanced warning threshold (R = 7.2) through weight allocation (α = 2.0, ε = 0.8), avoiding misjudgment in a single dimension and achieving dynamic balance across multiple dimensions. Furthermore, model parameters can be dynamically adjusted in response to regulatory policies (such as public health emergencies). For example, during an epidemic, the δ value for vaccine data was increased from 1.5 to 2.0 to prioritize highly sensitive populations.
[0079] In this embodiment, risk levels of standard pharmacovigilance data are divided according to risk scores, including steps 13.1 and 13.2.
[0080] Step 13.1: Construct a first risk threshold and a second risk threshold based on the historical risk score distribution.
[0081] In a feasible embodiment, the expression of the first risk threshold is μ R +2σ R , μ R represents the mean historical risk score, σ R represents the standard deviation of historical risk scores, and the expression of the second risk threshold is μ R +0.5σ R .
[0082] Step 13.2: Determine the risk level based on the relationship between the risk score, the first risk threshold, and the second risk threshold.
[0083] Specifically, if the risk score is greater than or equal to a first risk threshold, the risk level of the standard pharmacovigilance data is determined to be high;
[0084] If the risk score is less than the first risk threshold and greater than or equal to the second risk threshold, the risk level of the standard pharmacovigilance data is determined to be medium;
[0085] If the risk score is less than the second risk threshold, the risk level of the standard pharmacovigilance data is determined to be low.
[0086] In another feasible embodiment, a formula is constructed to calculate the risk score by combining the drug severity coefficient, patient sensitivity coefficient, and reporting timeliness coefficient. The formula weights are determined by expert experience and statistical analysis to ensure that the score reflects the comprehensive impact of risk. This method provides a more comprehensive basis for risk quantification by integrating multi-dimensional information. For example, a decision tree algorithm can be used to determine the risk level.
[0087] In another possible embodiment, the decision tree further subdivides the risk based on the score and other characteristics such as the patient's medical history. For example, if the patient has a history of heart disease, the risk level may be increased to high risk.
[0088] The following describes the process of dynamically adjusting blockchain resource allocation based on risk level in this embodiment.
[0089] Traditional drug vigilance systems use a unified blockchain resource allocation for high-risk data (such as serious adverse reactions to vaccines) and low-risk data (such as mild reactions to common drugs), resulting in insufficient priority for high-risk data processing and waste of low-risk data resources. Furthermore, they do not consider the differences in resource requirements for drug types (such as anticancer drugs and common drugs) and data source characteristics (such as high volatility of social media), resulting in resource allocation that cannot adapt to actual risk needs. Furthermore, the linear resource allocation model makes it difficult to balance node loads, and the data sharding strategy is not linked to the risk level, affecting system stability and data reliability. To address the above-mentioned problems in the prior art, the present invention dynamically adjusts blockchain resource allocation according to risk level, specifically including steps I to II.
[0090] Step I: Calculate the blockchain resource allocation intensity based on risk level, adverse reaction degree, and volatility.
[0091] In an embodiment of the present invention, the blockchain resource allocation intensity represents the processing priority of standard pharmacovigilance data.
[0092] Step II: Dynamically adjust blockchain resource allocation according to blockchain resource allocation intensity.
[0093] Blockchain resources include the number of blockchain nodes allocated and the number of data shards; there is a nonlinear positive correlation between the number of blockchain nodes allocated and the intensity of blockchain resource allocation; there is a linear positive correlation between the number of data shards and the intensity of blockchain resource allocation.
[0094] In a feasible embodiment, the blockchain resource allocation intensity is calculated based on the risk level, adverse reaction degree, and volatility, including:
[0095] By calculating the formula
[0096]
[0097] Get the blockchain resource allocation intensity R blockchain (L, d, s); where L represents the risk level, L∈(1, 2, 3), (high corresponds to 3, medium corresponds to 2, and low corresponds to 1), d′ represents the drug type, d′=1 represents the anticancer drug type, d′=2 represents the vaccine type, and d′=3 represents the common drug type. a, b, and c represent empirical parameters, a+b=1, c∈[0.5, 2], and CV(s) represents the coefficient of variation of the data source. R blockchain (L,d,s)∈[0,1].
[0098] In a feasible embodiment, dynamically adjusting blockchain resource allocation according to blockchain resource allocation intensity includes:
[0099] The number of blockchain node allocations is expressed as The number of data shards is expressed as Among them, D size Indicates data and size, B storage It represents the storage bandwidth, and its size can be configured according to actual needs (such as 500 GB). ρ=3.
[0100] It is worth mentioning that the nonlinear positive correlation between the number of blockchain node allocations and the blockchain resource allocation intensity can ensure the load balancing of distributed storage; the linear positive correlation between the number of data shards and the blockchain resource allocation intensity can improve fault tolerance.
[0101] In another feasible embodiment, a multidimensional quality assessment strategy is used to quantitatively score the standard pharmacovigilance data to obtain a quality assessment result corresponding to the standard pharmacovigilance data, including:
[0102] Obtain the drug severity coefficient, patient sensitivity coefficient, and reporting timeliness coefficient. Using a pre-established assessment matrix, map the drug category to the drug severity coefficient, the patient age and gender to the patient sensitivity coefficient, and the report source to the reporting timeliness coefficient to obtain a coefficient set. Using this coefficient set, calculate the risk score using the formula R = S × P × U, where S represents the drug severity coefficient, P represents the patient sensitivity coefficient, and U represents the reporting timeliness coefficient. Based on this risk score set, a decision tree algorithm is used to determine whether the risk score matches the preset high, medium, and low risk thresholds and determine the risk level.
[0103] For example, aspirin, a nonsteroidal anti-inflammatory drug, could be mapped to a severity coefficient of 0.6. The patient, a 65-year-old male, could be mapped to a sensitivity coefficient of 0.8, taking into account the sensitivity of older adults and males to drug metabolism. The report source, a clinician, updates frequently, so the timeliness coefficient could be mapped to 0.9. It's important to note that the assessment matrix is constructed based on historical pharmacovigilance data and clinical guidelines to ensure that the coefficients reflect actual risk characteristics. This approach enhances the objectivity of risk assessment by quantifying different dimensions.
[0104] In another embodiment of the present invention, after classifying the risk levels of standard drug vigilance data according to the risk scores and dynamically adjusting the blockchain resource allocation according to the risk levels, the method also includes triggering data completion of the standard drug vigilance data based on the risk levels and establishing a data modification traceability mechanism.
[0105] The following technical deficiencies exist in the current field of pharmacovigilance data management:
[0106] 1. The risk response mechanism is rigid.
[0107] The existing system uses a one-size-fits-all approach, uniformly filling or deleting missing data. When addressing fatal adverse reactions (such as fulminant myocarditis caused by immune checkpoint inhibitors), missing key fields (such as medication dosage and concomitant medication records) make it impossible to trace the pathogenic mechanism, delaying the implementation of risk control measures.
[0108] 2. Insufficient data association mining.
[0109] Traditional completion methods rely on single-field matching and fail to fully utilize the multi-dimensional associations in historical data. For example, in a clinical trial of a novel biologic, there was a potential correlation between animal experimental data from the R&D phase and post-marketing allergic reaction reports, but existing systems were unable to automatically identify this cross-stage data connection.
[0110] 3. Modify the missing traceability mechanism.
[0111] Current data revisions generally use an overwriting update method, which lacks complete operational records. For example, the safety data for a certain antidiabetic drug was incorrectly overwritten due to multiple additions, making it impossible to restore the data evolution process.
[0112] The present invention triggers data completion for standard drug vigilance data based on risk level and establishes a data modification traceability mechanism to overcome the above-mentioned defects.
[0113] Specifically, for standard drug vigilance data with a high risk level, key information of similar cases is extracted from the historical database, and the potential relationship between data items is analyzed using association rule mining algorithms, and missing fields are filled in with probabilistic inference; for standard drug vigilance data with a medium or low risk level, batch optimization strategies are applied, and the data are sorted according to the preset processing priority queue, and the original records are traced back from the data source.
[0114] The following describes the process of extracting key information from similar cases from a historical database for high-risk standard pharmacovigilance data, using an association rule mining algorithm to analyze the potential relationships between data items, and probabilistically inferring and filling in missing fields.
[0115] Specifically, the pharmacovigilance data of the standard pharmacovigilance data corresponding to the high risk level is obtained from the historical database. By comparing the consistency between the data items, the preset threshold is used to filter out the key information, and a set of similar cases containing the key information is obtained.
[0116] For a collection of similar cases, an association rule mining algorithm is used to analyze the potential relationships between data items. The strength of association between each data item and other data items is calculated to determine candidate association items for the missing field. If the association strength of the candidate association item exceeds a preset threshold, a probabilistic inference algorithm is used to predict the value of the missing field, obtaining a set of predicted values. Based on the predicted value set, the missing fields are filled in using field filling rules to generate complete high-risk data records, resulting in a completed data set.
[0117] In a feasible embodiment, when obtaining similar data corresponding to high-risk data from a historical database, a database query tool can be used to filter out records that are highly similar to the current high-risk data in terms of drug category, patient characteristics, report source, etc.
[0118] For example, consider high-risk data involving adverse reactions caused by aspirin, with the patient being a 70-year-old woman and the report originating from the emergency department. Cases involving aspirin use, female patients aged 65-75, and gender can be extracted from the historical database to form an initial set of similar data. This approach, by limiting key fields, ensures case relevance and provides a reliable foundation for subsequent analysis.
[0119] In one feasible embodiment, when comparing consistency between data items, a field matching algorithm can be used to calculate the overlap between the current data and similar cases in key fields such as drug dosage and patient medical history. A preset consistency threshold of 80% means that data with a field match below 80% is eliminated. For example, if a similar case shows a patient taking 100 mg of aspirin daily, while the current data shows 75 mg daily, the match is 90%, and the data is retained. This screening ensures high relevance of key information, providing accurate input for subsequent mining.
[0120] In another feasible embodiment, drug vigilance data stored in the vigilance database of historical periods can also be extracted to construct a drug vigilance knowledge graph; the association between drugs, adverse reactions, and patient characteristics can be mined based on the FP-Growth algorithm; and the standard drug vigilance data corresponding to the high risk level can be supplemented based on the association and the similarity between the data.
[0121] The following describes the process of applying a batch optimization strategy to standard pharmacovigilance data with a medium or low risk level.
[0122] Specifically, based on the data priority, the data to be processed is obtained from the preset processing queue, and the data to be processed is arranged according to the priority through a sorting algorithm to obtain an ordered data sequence. For the ordered data sequence, a data tracing tool is used to extract the original data from the source record. If the original data is inconsistent with the ordered data sequence, the difference is determined through a comparison algorithm to obtain the data items that need to be confirmed. Through the automated request tool, a confirmation request containing the data items that need to be confirmed is sent to the data provider, and the returned feedback information is obtained. If the feedback information contains valid updated data, the content of the data after confirmation is determined. Based on the content of the confirmed data, a batch scheduling tool is used to update the data status mark.
[0123] The following describes the data modification traceability mechanism.
[0124] Specifically, multi-fold cross-validation is implemented on the filled high-risk data (standard drug vigilance data with a high risk level). The existing reliable data is randomly divided into k subsets, k-1 subsets are selected each time as training sets, and the remaining one is used as a validation set. After k cycles, the average validation index is calculated, and the reliability level of the filled data is determined according to the preset validation index thresholds (accuracy>85%, precision>80%, recall>75%).
[0125] Based on the cross-validation results, a reliability identifier is added to the padding data. A unique verification hash value is generated for the verified padding data. The verification process parameters and results are recorded in the blockchain audit system, forming a verification chain for the padding data and obtaining a revised high-risk dataset. Specifically, complete high-risk data records are obtained from the padding data. These high-risk data records are processed using a random partitioning algorithm, partitioning them into k subsets to obtain a partition set consisting of a training subset and a validation subset. A cross-validation algorithm is then used to iterate over these partition sets, selecting k-1 training subsets for logistic regression training each time and performing validation on the remaining validation subset, to obtain a validation metric set for each iteration. Based on this validation metric set, the average of accuracy, precision, and recall is calculated to obtain an average metric set. If the accuracy, precision, and recall in this average metric set exceed a preset accuracy threshold, the precision, and recall exceed a preset recall threshold, the padding data is determined to have a high reliability rating, and a reliability rating result is obtained.
[0126] During the data correction process, the digital signature and blockchain traceability mechanism are activated to record the operator identity, modification time, modification content, and modification basis of each data modification, generate a unique operation hash value, and write the modification record into the tamper-proof blockchain ledger.
[0127] A bidirectional index structure is established for the modification records in the blockchain ledger, linking the verification proof chain and the modification operation chain, building a modification-verification relationship graph, and identifying the consistency of data modification and verification through graph analysis. If inconsistencies are found, an abnormal alarm is triggered, forming a complete data modification audit chain.
[0128] Among them, associating the verification proof chain with the modification operation chain includes using a bidirectional index tool to generate an index table containing modification record identifiers and operation chain identifiers to obtain an initialized bidirectional index structure. According to the bidirectional index structure, a graph database tool is used to associate the modification record identifier with the verification proof identifier to generate graph nodes and edges to obtain a relationship graph containing modification-verification associations. If the hash values of the associated nodes of the modification record and the verification proof in the relationship graph match, a consistency verification tool is used to calculate the data consistency between the nodes to obtain a consistency verification result; if there is no match, an exception log is generated. According to the consistency verification result and the exception log, an audit chain generation tool is used to write the modification record, verification proof and exception log into the audit chain, and a new data version is generated through the signature identifier to obtain a complete audit chain record.
[0129] For example, obtaining modification records and verification proofs are key steps in extracting data from blockchain ledgers. Blockchain ledgers store a log of every data modification, including the operator's identity, modification time, and modification content. Verification proofs ensure the authenticity of modification records through digital signatures or hash values.
[0130] In one feasible implementation, a bidirectional indexing tool generates an index table containing modification record identifiers and operation chain identifiers for rapid data location. The bidirectional indexing structure establishes a mapping relationship between record identifiers (e.g., record ID: REC001) and operation chain identifiers (e.g., on-chain transaction ID: TXN789).
[0131] In one feasible implementation, a graph database tool associates modification record identifiers with verification proof identifiers to generate a relationship graph. The graph uses nodes to represent records and proofs, and edges to represent associations.
[0132] For standard pharmacovigilance data with a medium or low risk level, the data is sorted according to a preset processing priority queue and the original records are traced back from the data source. Specifically, the updated data is written to the storage through database operations to obtain a data record with consistency maintained.
[0133] In summary, the data management method for pharmacovigilance personnel provided by the present invention utilizes a multidimensional quality assessment strategy to quantitatively score standard pharmacovigilance data. Based on the drug type and data source type in the standard pharmacovigilance data, the quality of the standard pharmacovigilance data is assessed from multiple dimensions. The method can dynamically adapt to the drug type and data source type, avoid misjudgment caused by single-dimensional defects, and help improve the accuracy of quality assessment, thereby improving the quality of pharmacovigilance data management. The method also scores the data based on the severity of adverse drug reactions, sensitivity of the patient group, and time urgency in the pharmacovigilance data, combining the correlation between drugs, patients, and time, which helps improve the accuracy of risk level classification, thereby improving the quality of pharmacovigilance data management.
[0134] The data management system for pharmacovigilance personnel disclosed in the present invention is described below.
[0135] like Figure 2 As shown, the data management system 200 includes:
[0136] A pre-processing module 201 is used to obtain multi-source heterogeneous pharmacovigilance data and perform de-isomerization processing on the multi-source heterogeneous pharmacovigilance data based on the characteristics of the data sources to obtain standard pharmacovigilance data;
[0137] The quality assessment module 202 is configured to quantitatively score the standard pharmacovigilance data using a multidimensional quality assessment strategy to obtain a quality assessment result corresponding to the standard pharmacovigilance data. The multidimensional quality assessment strategy performs a quality assessment on the standard pharmacovigilance data from multiple dimensions based on the drug type and data source type in the standard pharmacovigilance data, and the quality assessment result is high or low quality.
[0138] The blockchain resource allocation module 203 is used to calculate the risk score corresponding to the standard pharmacovigilance data with a quality assessment result of low quality using a pre-trained multidimensional risk scoring model, classify the risk level of the standard pharmacovigilance data according to the risk score, and dynamically adjust the blockchain resource allocation according to the risk level; the multidimensional risk scoring model is based on the severity of adverse drug reactions in the pharmacovigilance data, the sensitivity of the patient group, and the urgency of timeliness, and the risk level is high, medium, or low.
[0139] It should be noted that the information interaction, execution process and other contents between the above modules are based on the same concept as the embodiment of the method of the present application. Their specific functions and technical effects can be found in the embodiment of the method, and will not be described in detail here. Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above functional units and modules is used as an example. In actual application, the above functions can be distributed to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the above method embodiment, and will not be described in detail here.
[0140] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of protection of the present application is limited to these examples. In line with the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of one or more embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.
[0141] The one or more embodiments of this application are intended to encompass all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of this application should be included in the scope of protection of this application.
Claims
1. A data management method for pharmacovigilance personnel, characterized in that: include: Acquiring multi-source heterogeneous pharmacovigilance data and performing de-isomerization processing on the multi-source heterogeneous pharmacovigilance data based on characteristics of the respective data sources to obtain standardized pharmacovigilance data; A multidimensional quality assessment strategy is used to quantitatively score the standard pharmacovigilance data to obtain a quality assessment result corresponding to the standard pharmacovigilance data; the multidimensional quality assessment strategy is based on the drug type and data source type in the standard pharmacovigilance data, and the quality assessment result is high quality or low quality. For the standard pharmacovigilance data with a low-quality quality assessment result, a risk score corresponding to the standard pharmacovigilance data is calculated using a pre-trained multidimensional risk scoring model, the risk level of the standard pharmacovigilance data is classified according to the risk score, and blockchain resource allocation is dynamically adjusted according to the risk level; The multidimensional risk scoring model is based on the severity of adverse drug reactions in the pharmacovigilance data, the sensitivity of the patient population, and the urgency of time. The risk level is high, medium, or low.
2. The data management method for pharmacovigilance personnel according to claim 1, characterized in that: The multiple dimensions include completeness, accuracy, consistency, and timeliness.
3. The data management method for pharmacovigilance personnel according to claim 2, characterized in that: The method of quantitatively scoring the standard pharmacovigilance data using a multidimensional quality assessment strategy to obtain a quality assessment result corresponding to the standard pharmacovigilance data includes: Calculate the weight corresponding to each dimension of the standard pharmacovigilance data based on the drug types and data source characteristics included in the standard pharmacovigilance data; wherein the drug types include at least anticancer drug types, vaccine types, and common drug types, and the adverse reaction levels of different drug types vary. The adverse reaction levels are obtained by normalizing expert corpus data. The data source characteristics include volatility. The greater the field missing rate of the data source, the greater the volatility. The weight is positively correlated with the adverse reaction level and negatively correlated with the volatility. Calculating a quality score of the standard pharmacovigilance data by combining a nonlinear exponential weighting function and the weight; the quality score is used to measure the quality of the standard pharmacovigilance data; The quality assessment result is determined based on a comparison result between the quality score and a preset quality threshold; when the quality score is greater than or equal to the preset quality threshold, the quality assessment result is high quality; otherwise, the quality assessment result is low quality.
4. The data management method for pharmacovigilance personnel according to claim 3, characterized in that: The risk level classification of the standard pharmacovigilance data according to the risk score includes: Constructing a first risk threshold and a second risk threshold respectively according to the historical risk score distribution; the first risk threshold is greater than the second risk threshold; The risk level is determined according to the relationship between the risk score, the first risk threshold, and the second risk threshold.
5. The data management method for pharmacovigilance personnel according to claim 4, characterized in that: Determining the risk level according to the relationship between the risk score, the first risk threshold, and the second risk threshold includes: If the risk score is greater than or equal to the first risk threshold, determining that the risk level of the standard pharmacovigilance data is high; If the risk score is less than the first risk threshold and greater than or equal to the second risk threshold, determining that the risk level of the standard pharmacovigilance data is medium; If the risk score is less than the second risk threshold, the risk level of the standard pharmacovigilance data is determined to be low.
6. The data management method for pharmacovigilance personnel according to claim 5, characterized in that: The dynamically adjusting blockchain resource allocation according to the risk level includes: Calculating a blockchain resource allocation intensity based on the risk level, the degree of adverse reactions, and the volatility; the blockchain resource allocation intensity represents a processing priority of standard pharmacovigilance data; Dynamically adjust blockchain resource allocation according to the blockchain resource allocation intensity.
7. The data management method for pharmacovigilance personnel according to claim 6, characterized in that: The blockchain resources include the number of blockchain node allocations and the number of data shards.
8. The data management method for pharmacovigilance personnel according to claim 7, characterized in that: There is a nonlinear positive correlation between the number of blockchain node allocations and the intensity of blockchain resource allocation; There is a linear positive correlation between the number of data shards and the intensity of blockchain resource allocation.
9. The data management method for pharmacovigilance personnel according to claim 1, characterized in that: After classifying the risk level of the standard pharmacovigilance data according to the risk score and dynamically adjusting blockchain resource allocation according to the risk level, the method further includes: Trigger data completion for the standard drug vigilance data based on the risk level, and establish a data modification traceability mechanism.
10. A data management system for pharmacovigilance personnel, characterized in that: include: a preprocessing module for acquiring multi-source heterogeneous pharmacovigilance data and performing de-isomerization processing on the multi-source heterogeneous pharmacovigilance data based on the characteristics of the data sources to obtain standard pharmacovigilance data; a quality assessment module, configured to quantitatively score the standard pharmacovigilance data using a multidimensional quality assessment strategy to obtain a quality assessment result corresponding to the standard pharmacovigilance data; the multidimensional quality assessment strategy performs a quality assessment on the standard pharmacovigilance data from multiple dimensions based on the drug type and data source type in the standard pharmacovigilance data, and the quality assessment result is high quality or low quality; a blockchain resource allocation module, configured to calculate, for the standard pharmacovigilance data whose quality assessment result is low quality, a risk score corresponding to the standard pharmacovigilance data using a pre-trained multidimensional risk scoring model, classify the risk level of the standard pharmacovigilance data according to the risk score, and dynamically adjust blockchain resource allocation according to the risk level; The multidimensional risk scoring model is based on the severity of adverse drug reactions in the pharmacovigilance data, the sensitivity of the patient population, and the urgency of time. The risk level is high, medium, or low.