Adverse drug reaction data collecting and processing method and system based on federal learning
By using federated learning to process multi-source adverse drug reaction data, preprocessing and semantic standardization are performed, core feature terms are extracted, and multi-dimensional risk assessment indicators are constructed. This solves the problems of data silos and low assessment accuracy in existing technologies, and achieves efficient, safe and reliable monitoring of adverse drug reactions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING FUTURE JUDIAN INFORMATION TECH CO LTD
- Filing Date
- 2026-03-20
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies rely heavily on centralized data aggregation, making it difficult to achieve collaborative data utilization among multiple institutions while protecting privacy. Multi-source heterogeneous data lacks unified semantic standards, resulting in low accuracy in feature extraction and risk assessment. Early warning and report generation mechanisms are not intelligent enough, making it impossible to achieve precise hierarchical push notifications and failing to meet the needs of efficient, safe, and reliable drug safety monitoring.
A federated learning-based approach to adverse drug reaction (ADR) data collection and processing is adopted. This approach involves acquiring multi-source data, preprocessing and semantic standardization, extracting initial feature terms, classifying niche categories, constructing multi-dimensional risk assessment indicators, performing confidence grading quantification and multi-feature fusion, generating a comprehensive ADR risk assessment result, and developing early warning information based on user roles.
It breaks down data silos while protecting data privacy, enhances the breadth and compliance of data sources, accurately extracts core risk features, improves feature mining efficiency and risk assessment reliability, enables tiered early warning and targeted push notifications, and provides timely early warning support for adverse drug reactions.
Smart Images

Figure CN121885237A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical data processing technology, and more specifically, to a method and system for collecting and processing adverse drug reaction data based on federated learning. Background Technology
[0002] Adverse drug reaction monitoring is a crucial component of drug safety supervision and rational drug use in clinical practice, playing a vital role in ensuring public medication safety. With the continuous improvement of medical informatization, medical institutions and pharmaceutical companies have accumulated a large amount of multi-source, heterogeneous adverse drug reaction data, providing essential data support for risk identification, assessment, and early warning. Federated learning, as a typical privacy-preserving computing technology, enables collaborative modeling of multi-party data without requiring data to leave the local machine, effectively balancing data value mining and information security protection. This provides a feasible technical path for the secure sharing and efficient utilization of adverse drug reaction data across institutions and regions.
[0003] However, existing technologies largely rely on centralized data aggregation, making it difficult to achieve collaborative data utilization among multiple institutions while protecting privacy, and easily leading to data silos. Multi-source heterogeneous data lacks unified semantic standards, resulting in low accuracy in feature extraction and risk assessment. Furthermore, early warning and report generation mechanisms are not intelligent enough to achieve precise tiered delivery, failing to meet the practical needs of efficient, safe, and reliable drug safety monitoring.
[0004] There are currently no effective solutions to the problems in the relevant technologies. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention proposes a method and system for collecting and processing adverse drug reaction data based on federated learning. This solves the problems mentioned in the background, such as the difficulty in achieving collaborative data utilization among multiple institutions while protecting privacy, and the potential for data silos. Furthermore, the lack of unified semantic standards for multi-source heterogeneous data leads to low accuracy in feature extraction and risk assessment. Simultaneously, the early warning and report generation mechanisms are not intelligent enough to achieve precise hierarchical delivery, failing to meet the practical needs of efficient, safe, and reliable drug safety monitoring.
[0006] To achieve the above objectives, the present invention provides the following technical solution: According to one aspect of the present invention, a method for collecting and processing adverse drug reaction data based on federated learning is provided, the method comprising: S1. Obtain adverse drug reaction data from multiple sources and preprocess the adverse drug reaction data from multiple sources; S2. Perform semantic standardization on the preprocessed multi-source adverse drug reaction data to obtain a complete multi-source fusion dataset; S3. Extract the initial feature term population from the complete multi-source fusion dataset, screen out the core seed terms and supplementary seed terms, divide the niche categories, expand the dimensions of all feature seed terms, analyze their potential associations and complete their diffusion and reproduction, and finally form the core feature set that affects the occurrence of adverse drug reactions. S4. Based on the core feature set, construct a multi-dimensional risk assessment index for adverse drug reactions, quantify the credibility of the drug risk assessment index through confidence grading, and obtain the comprehensive risk assessment result of adverse drug reactions through multi-feature fusion calculation. S5. Based on the comprehensive risk assessment results of adverse drug reactions, initiate the corresponding level of risk warning, and formulate and push warning information based on user roles, and generate adverse drug reaction warning reports.
[0007] According to another aspect of the present invention, a federated learning-based adverse drug reaction data collection and processing system is also provided, the system comprising: The data acquisition module is used to acquire adverse drug reaction data from multiple sources and to preprocess the adverse drug reaction data from multiple sources. The data processing module is used to perform semantic standardization on the preprocessed multi-source adverse drug reaction data to obtain a complete multi-source fusion dataset. The feature analysis module is used to extract an initial population of feature terms from the complete multi-source fusion dataset, screen out core seed terms and supplementary seed terms, classify niche categories, expand the dimensions of all feature seed terms, analyze their potential associations and complete their diffusion and reproduction, and finally form a core feature set that affects the occurrence of adverse drug reactions. The risk assessment module is used to construct multi-dimensional risk assessment indicators for adverse drug reactions based on core feature sets, quantify the credibility of drug risk assessment indicators through confidence grading, and obtain the comprehensive risk assessment results of adverse drug reactions through multi-feature fusion calculation. The early warning information generation module is used to initiate corresponding level risk warnings based on the comprehensive risk assessment results of adverse drug reactions, and to push early warning information based on user roles, and generate adverse drug reaction early warning reports.
[0008] The beneficial effects of this invention are as follows: 1. This invention enables the secure collection and preprocessing of multi-source adverse drug reaction data, breaking down data silos while protecting data privacy and improving the breadth and compliance of data sources. Semantic standardization eliminates semantic differences between heterogeneous data, forming a standardized multi-source fusion dataset. Feature term niche segmentation and expansion accurately extract core risk features, improving feature mining efficiency. Combining confidence grading and multi-feature fusion for risk assessment enhances the reliability and objectivity of assessment results. Based on risk levels, tiered early warning and targeted push notifications are implemented, enabling rapid generation of early warning reports and providing strong support for timely early warning and scientific handling of adverse drug reactions.
[0009] 2. This invention can accurately locate high-value risk features through feature term extraction, seed screening, and niche segmentation based on association strength. Combining feature dimension expansion and association diffusion fully explores potential risk associations, improving feature comprehensiveness. Through iterative evaluation and redundancy removal, the core feature set is ensured to be concise and effective, providing reliable support for subsequent risk assessment, thereby improving the accuracy and rationality of adverse drug reaction analysis.
[0010] 3. This invention quantifies the credibility of indicators through confidence grading and combines normalization and average value fusion to perform multi-feature calculations, thereby improving the stability and rationality of risk assessment results. Through iterative optimization and screening of high-credibility indicators, the assessment process can be dynamically optimized, reducing redundant interference and making the comprehensive risk assessment of adverse drug reactions more accurate and objective, providing a reliable basis for subsequent early warning. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart of a method for collecting and processing adverse drug reaction data based on federated learning according to an embodiment of the present invention; Figure 2 This is a schematic diagram of a drug adverse reaction data collection and processing system based on federated learning according to an embodiment of the present invention.
[0013] In the picture: 1. Data acquisition module; 2. Data processing module; 3. Feature analysis module; 4. Risk assessment module; 5. Early warning information generation module. Detailed Implementation
[0014] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0015] In the description of this invention, unless otherwise stated, "a plurality of" means two or more. Furthermore, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0016] According to embodiments of the present invention, a method and system for collecting and processing adverse drug reaction data based on federated learning are provided.
[0017] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments, such as... Figure 1 As shown, according to an embodiment of the present invention, a method for collecting and processing adverse drug reaction data based on federated learning includes: S1. Obtain adverse drug reaction data from multiple sources and preprocess the adverse drug reaction data from multiple sources; Specifically, the multi-source adverse drug reaction data originates from multiple distributed data holders, covering various scenarios such as medical care, drug regulation, and production and distribution. This includes data from medical institutions (patient basic information, clinical medication records, adverse reaction clinical manifestations, and clinical examination data), drug monitoring institutions (adverse drug reaction reporting records, epidemiological statistics, and preliminary risk signal screening results), pharmaceutical manufacturers (drug ingredient information, production process parameters, drug instruction revision records, and post-marketing monitoring data), and retail pharmacies (drug sales records, consumer medication consultation records, and simplified adverse reaction information reported by pharmacies). All of this data is heterogeneous and stored locally at each node, exhibiting inconsistent formats, semantic inconsistencies, and high privacy sensitivity, making direct centralized aggregation and processing impossible. This invention utilizes a distributed collaborative architecture based on federated learning to securely acquire this multi-source data, ensuring that the original data remains locally and that privacy is not compromised throughout the entire process.
[0018] The specific acquisition process is as follows: A federated learning architecture is built, including a federated coordination node and multiple data holder nodes. Each node deploys a federated learning client. The coordination node is responsible for global coordination and scheduling and does not store any raw data. Each data holder node submits a data access application to the coordination node through the federated client and specifies the relevant data information. After the coordination node completes the authentication, it establishes a secure communication link and issues instructions to each node containing only the feature fields to be acquired. Each node filters the corresponding feature data locally and performs preliminary encryption processing. The encrypted feature information is transmitted to the coordination node through the federated coordination protocol. The coordination node summarizes and deduplicates to form a multi-source feature set, while monitoring the transmission status in real time and retaining interaction logs. Preprocessing operations are all completed on the local nodes of each data holder. The raw data or preprocessing intermediate data is not aggregated to the coordination node. Only the standardized feature data after preprocessing participates in the subsequent federated coordination process.
[0019] The specific preprocessing steps are as follows: Each local node first performs data cleaning, removes invalid data, corrects entry errors, then identifies and removes duplicate records, merges some duplicate records with complementary information, fills in missing values using corresponding methods, directly removes records with missing key privacy fields, identifies, verifies, and processes outliers, unifies the data format according to the standard issued by the coordination node, and each node reports the preprocessed feature data statistics to the coordination node. After verification, the preprocessing is completed; if the verification fails, it is reprocessed until the requirements are met.
[0020] S2. Perform semantic standardization on the preprocessed multi-source adverse drug reaction data to obtain a complete multi-source fusion dataset; S3. Extract the initial feature term population from the complete multi-source fusion dataset, screen out the core seed terms and supplementary seed terms, divide the niche categories, expand the dimensions of all feature seed terms, analyze their potential associations and complete their diffusion and reproduction, and finally form the core feature set that affects the occurrence of adverse drug reactions. Specifically, the core characteristic set that influences the occurrence of adverse drug reactions includes: Drug-related characteristics: drug ingredients, drug category, dosage, route of administration, course of treatment, combination therapy, drug properties, etc. Adverse reaction characteristics: type of adverse reaction symptoms, onset time, severity, and correlation of symptoms; Patient population characteristics: age, gender, underlying diseases, allergy history, body type, liver and kidney function, etc. Other related characteristics include: medication use scenario, disease diagnosis results, and adverse reaction management methods.
[0021] S4. Based on the core feature set, construct a multi-dimensional risk assessment index for adverse drug reactions, quantify the credibility of the drug risk assessment index through confidence grading, and obtain the comprehensive risk assessment result of adverse drug reactions through multi-feature fusion calculation. S5. Based on the comprehensive risk assessment results of adverse drug reactions, initiate the corresponding level of risk warning, and formulate and push warning information based on user roles, and generate adverse drug reaction warning reports.
[0022] In this optional embodiment, the preprocessed multi-source adverse drug reaction data undergoes semantic normalization to obtain a complete multi-source fusion dataset, including: S21. Perform terminology alignment on the preprocessed multi-source adverse drug reaction data to unify the expression of drug, symptom, and population-related data and eliminate semantic differences between cross-data sources. S22. Standardize drug entities according to the preset drug coding system, clarify drug identification, classification and attribute information, so that drug data from different sources have a unified basis for identification; S23. Standardize and convert adverse reaction symptoms according to a pre-set authoritative terminology database to unify symptom names and descriptions; S24. Perform semantic analysis on the relationships between data and establish stable correspondence rules between drugs, symptoms, populations and medication information; S25. Integrate the various types of data that have completed semantic standardization processing to form a multi-source fusion dataset of adverse drug reactions with standardized structure, unified terminology, and complete association.
[0023] Specifically, the preprocessed multi-source adverse drug reaction (ADR) data is retrieved, with a focus on extracting drug, symptom, and population-related data. Addressing the issues of synonyms and inconsistent terminology across different data sources, the descriptions of various data types are standardized according to predefined terminology guidelines, eliminating semantic differences across data sources and ensuring consistent representation of the same data object across different data sources. Following the classification and coding system of the WHO Drug Dictionary, all drug entities are standardized, assigning unique identification information to each drug and clarifying its specific classification, dosage form, specifications, and other attributes. This ensures a unified identification basis for drug data from different sources and in different formats, guaranteeing the accuracy of subsequent data matching and analysis. The Med DRA (Medical Dictionary for Regulatory Activities) authoritative terminology database for ADR monitoring is accessed to match non-standard and inconsistent ADR symptom descriptions from various data sources with the standard expressions in the database, completing the standardized conversion of symptom names and clinical manifestation descriptions, ensuring that all ADR symptom descriptions are consistent and standardized. Semantic analysis is performed to examine the inherent relationships between drugs, adverse reaction symptoms, population characteristics, and medication information. Combined with clinical medication knowledge, stable correspondence rules are established between various data types to clarify the logical connections between data and improve data relevance and analyzability. The data, after terminology alignment, drug entity standardization, symptom standardization, and relationship analysis, are integrated according to a pre-defined structured format. Abnormal data identified during the integration process is removed to form a complete multi-source fusion dataset.
[0024] In this optional embodiment, an initial population of feature terms is extracted from the complete multi-source fusion dataset. Core seed terms and supplementary seed terms are selected, and niche categories are defined. The dimensions of all feature seed terms are expanded, their potential associations are analyzed, and diffusion and propagation are completed. Finally, a core feature set influencing the occurrence of adverse drug reactions is formed, including: S31. Extract an initial population of drug-related and adverse reaction-related feature terms from the complete multi-source fusion dataset; S32. Select the most relevant feature terms with the largest preset number of weight values from the initial feature term population related to drugs and adverse reactions as core seed terms, and then randomly select a preset number of potential related supplementary terms from the remaining feature terms as supplementary seed terms. S33. Sort all core feature seed terms and supplementary feature seed terms according to their association strength, and classify them into niche categories with different association dimensions; S34. Based on each microhabitat, expand the feature dimensions of all feature seed entries, analyze the potential association between feature seeds and adverse drug reactions, and complete the growth, reproduction and association diffusion of feature seeds. S35. Reassess the association strength of the expanded feature term population, remove redundant terms with association strength below the preset threshold, and retain high-value associated terms. S36. Count the total number of current feature word populations. If it does not exceed the preset maximum number of core feature populations, return to step S34 to continue expanding. If it has exceeded the limit, select the word with the best fit as the core candidate feature according to the fitness of each niche association from high to low. S37. Summarize all core candidate features to form a core feature set that affects the occurrence of adverse drug reactions.
[0025] Specifically, a complete multi-source fusion dataset is retrieved. This dataset has undergone semantic standardization and covers various standardized data related to drugs, adverse reactions, populations, and medication use. According to preset feature extraction rules, only terms directly related to drug attributes, adverse reaction symptoms, population characteristics, and medication behavior are extracted. Redundant information that has no connection to the occurrence of adverse drug reactions, such as drug manufacturer addresses and irrelevant patient personal information, is removed to ensure that all extracted terms have potential analytical value. All initial feature terms related to drugs and adverse reactions are accurately extracted from the dataset to form an initial feature term population, ensuring that no potentially related feature information is missed. Weight analysis is performed on the extracted initial feature term population. Based on the degree of correlation between each feature term and the occurrence of adverse drug reactions, the weight values of each term are calculated and ranked. A preset number of highly correlated feature terms with the largest weight values are selected (the preset number ranges from 15 to 25, with a preference for 20, which can be flexibly adjusted according to the data scale). These are identified as core seed terms, which are the key core features affecting the occurrence of adverse drug reactions. From the remaining initial feature terms, a preset number of potential correlated supplementary terms are randomly selected as supplementary seed terms (the preset number ranges from 8 to 12, with a preference for 10, and the ratio of the number of supplementary seed terms to the number of core seed terms is controlled at 1:2 to ensure the richness and non-redundancy of supplementary features). Supplementary seed terms can enrich the feature dimensions, avoid the core features being too singular, and ensure the comprehensiveness of subsequent feature expansion.
[0026] All selected core seed terms and supplementary seed terms were evaluated and ranked based on their association strength, arranged in descending order of association strength. Simultaneously, a pre-defined association dimension classification rule was established: terms were divided into four association dimensions based on their attributes: drug-related, adverse reaction symptom-related, population characteristic-related, and medication behavior-related. Each dimension corresponds to a niche category, with terms of similar attributes grouped into the same niche to ensure clustering of similar features and differentiation of dissimilar features. Combining the attributes and association logic of various sub-terms, all seed terms were assigned to niche categories within different association dimensions. Based on these niche categories, feature dimensions were expanded for all feature seed terms. Integrating association information from multi-source fusion datasets, the potential association between each feature seed and adverse drug reactions was analyzed in depth, completing the growth, reproduction, and diffusion of feature seeds, further enriching the number and dimensions of feature terms and improving the comprehensiveness of the features.
[0027] After feature expansion is completed, the association strength of the expanded feature term population is reassessed. A preset association threshold is set (the preset threshold ranges from 0.3 to 0.5, with 0.4 being preferred; terms with an association degree lower than this threshold are considered redundant terms with no practical analytical value). Redundant terms with an association degree lower than this threshold are removed, and only high-value associated terms with high association degree and practical analytical value are retained to reduce interference from redundant information and ensure the effectiveness of feature terms. The total number of currently retained high-value related terms is counted and compared with the preset maximum number of core features (the preset maximum number of core features ranges from 50 to 80, with 60 preferred, balancing feature comprehensiveness and conciseness). If the current total number of terms does not exceed the preset number, return to step S34 to continue expanding the feature dimensions until the number of terms reaches the preset requirement. If the current total number of terms exceeds the preset number, sort the terms by their association fitness in each niche category from high to low, and select the terms with the best fit in each niche as core candidate features to ensure the representativeness and effectiveness of the core candidate features. All core candidate features selected from niches are summarized, and the candidate features are finally verified, eliminating duplicate and conflicting terms. This results in a clear, closely related, comprehensive, and effective core feature set that influences the occurrence of adverse drug reactions, providing core data support for subsequent adverse drug reaction risk assessment.
[0028] Specifically, this invention uses an invasive weed algorithm to screen and form a core feature set that influences the occurrence of adverse drug reactions. The invasive weed algorithm is a swarm intelligence optimization algorithm that simulates the establishment, reproduction, spread, spatial competition and survival of the fittest of wild weeds in the natural environment. The whole process consists of six steps executed sequentially and implemented in a closed loop iteration: initial population construction, fitness calculation, reproduction and spread, niche competition, global iterative elimination, and feature set output.
[0029] In this invention, feature terms are used as individuals in the invasive weed algorithm, and association strength and niche suitability are used as fitness criteria. Global optimal search is achieved through four mechanisms: reproduction, diffusion, competition, and elimination. Initial population construction receives multi-source fusion datasets, extracts all initial feature terms related to adverse drug reactions, completes feature encoding, attribute labeling, initial association strength calculation, and population initialization, forming the initial input population for the invasive weed algorithm. Fitness calculation is the core driving link of the algorithm, combining niche affiliation weights to calculate the comprehensive fitness of each feature term. The fitness value directly determines subsequent reproductive capacity and competition priority. Reproduction... The diffusion process allocates the number of offspring based on fitness values, and completes the feature dimension expansion through normal distribution random diffusion, simulating the offspring diffusion process of weeds to achieve feature association mining. The niche competition is constrained by a preset niche radius, and completes the competition of similar features within four dimensions: drug, symptom, population, and medication. High-fitness features are retained and low-association features are eliminated. The global iterative elimination is performed with a preset maximum population size as the upper limit. When the population size exceeds the limit, the best individuals are retained by sorting by fitness. The feature set output summarizes the high-value features after global screening, and after deduplication and verification, forms the final core feature set, which is used as the output of the invasive weed algorithm.
[0030] The training and iteration process of the "Invading Weeds" algorithm is as follows: The input data for the "Invading Weeds" algorithm comes directly from a complete multi-source fusion dataset after semantic standardization. The data sources include adverse reaction data of desensitized drugs from multiple medical institutions, drug regulatory agencies, and pharmaceutical companies. All data has undergone missing value imputation, outlier removal, duplicate data deletion, and WHODrug and Med... DRA terminology standardization and numerical field min-max normalization ensure unified data format, dimensional standardization, and no leakage of sensitive information, meeting the requirement that federated learning data does not leave the domain. The Invading Weeds algorithm uses a comprehensive fitness function as the evaluation standard for individual performance, with the formula: Comprehensive Fitness = Feature-Adverse Reaction Association Strength × Niche Fit Weight. The feature-adverse reaction association strength is calculated using the point mutual information method, reflecting the co-occurrence degree of the feature and the adverse reaction event. The niche fit weight is dynamically determined based on the representativeness of the feature in the corresponding dimensional cluster, with a value range of 0 to 1, ensuring that the selected features have both strong correlation and dimensional representativeness, without redundancy or bias. The algorithm adopts a hybrid optimization strategy of normal distribution spatial diffusion, niche clustering constraints, and global competitive elimination, which has the advantages of strong global search ability, stable convergence, and low premature convergence. It does not require gradient calculation, is suitable for discrete feature screening scenarios, and is highly compatible with the high-dimensional nonlinear feature patterns of adverse drug reactions.
[0031] The key hyperparameter settings of the invading weed algorithm are preset based on the actual needs of drug adverse reaction feature mining. These include: an initial population size of 92, which is the number of initial feature terms actually extracted from the multi-source fusion dataset; 20 core seed terms, selected based on the top 20 high-value features in terms of association strength; 10 supplementary seed terms, randomly selected from the remaining features to ensure population diversity; an association strength threshold of 0.4, where features below this value are considered redundant and removed; a niche radius of 0.3, used to determine whether features belong to the same association dimension cluster; a maximum population size of 60, which is the upper limit of the total core feature set to avoid overfitting due to too many features; an initial diffusion standard deviation of 0.5, controlling the early feature expansion range; a final diffusion standard deviation of 0.01, which narrows the search range and improves convergence accuracy in the later stages of iteration; a nonlinear adjustment factor of 3, used to balance global search and local refinement; and an iteration termination condition when the feature population size reaches the maximum population size or there is no fitness improvement for two consecutive generations.
[0032] The detailed process of the invasive weed algorithm is as follows: Initial population construction: Receive a multi-source fusion dataset, extract 92 initial feature terms as the initial weed population, and complete feature encoding and attribute binding; Fitness calculation: Traverse all individuals, calculate association strength and comprehensive fitness, and sort in descending order. Select the top 20 high-fitness features as core seeds, and randomly select 10 as supplementary seeds to form 30 initial seed populations; Reproduction and diffusion: Allocate the number of offspring according to various sub-fitnesses. The higher the fitness, the more offspring there are. Complete feature diffusion and dimensional expansion within the corresponding niche, generating potential associated features; Niche competition: With a radius constraint of 0.3, complete the competition of similar features within each dimension, and eliminate low-value features with association strength below 0.4; Global iterative elimination: Count the current population size. When the size exceeds 60, perform global elimination, retaining the best individuals according to comprehensive fitness from high to low; Repeat the above reproduction, diffusion, competition, and elimination process until the iteration termination condition is met; Finally, the feature set output: Deduplicate and verify the selected features to form the core feature set affecting adverse drug reactions.
[0033] The input data for the invasive weed algorithm includes a multi-source fusion dataset, initial feature terms, initial association strength values, normalized feature values, niche classification rules, and a set of preset hyperparameters. All inputs are directly derived from adverse drug reaction monitoring data, are strongly relevant to the application scenario, are reusable, and verifiable. The algorithm output data includes a core feature set, association strength of each feature, fitness value, niche category, feature expansion path, and redundant feature removal records. The output and input form a complete iterative closed loop through fitness calculation, propagation and diffusion, and competitive elimination, which is directly used for the construction of subsequent risk assessment indicators.
[0034] In this optional embodiment, all core feature seed terms and supplementary feature seed terms are sorted according to their association strength, and niche categories with different association dimensions are divided, including: S331. Sort all core feature seed terms and supplementary feature seed terms in descending order according to the association strength value of the feature terms. If the total number of feature terms is greater than the preset maximum population size, only the feature terms with the highest ranking are retained. S332. Set the center of the first niche as the feature term with the highest association strength value, and label the feature term; S333. For the remaining unlabeled feature terms, if the feature similarity between the unlabeled feature term and the current niche center feature term is less than the preset niche radius, then it is included in the niche and labeled; otherwise, it does not belong to the niche. S334. Among the remaining unlabeled feature words, select the unlabeled feature word with the largest association strength value as the new niche center, and at the same time, classify and label the unlabeled feature words whose feature similarity is less than the preset niche radius into the niche. S335. If all feature terms have been labeled, the niche classification ends; otherwise, return to step S334 to continue classification.
[0035] Specifically, all core feature seed terms and supplementary feature seed terms selected in the early stage are retrieved. The association strength value corresponding to each term is obtained, and all terms are sorted in descending order based on the association strength value. Simultaneously, a preset maximum population size is considered (the preset value ranges from 30-40, with 35 preferred to balance sorting efficiency and feature completeness). If the total number of all feature terms exceeds this preset value, only the top-ranked terms are retained, while those ranked lower and with weaker association strength are removed, ensuring that all terms participating in subsequent classification are of high association value. After completing the descending sort, the feature term with the highest association strength value at the top of the ranking is set as the center term of the first niche. This center term is also marked to indicate that it has participated in niche division, avoiding subsequent duplication and providing a core reference for subsequent term classification. For all remaining unlabeled feature seed terms, calculate the feature similarity between each unlabeled term and the current first niche center term. Simultaneously, retrieve the preset niche radius (the preset radius ranges from 0.2 to 0.4, with 0.3 being preferred, to determine whether terms belong to the same association dimension). If the feature similarity of an unlabeled term is less than the preset radius, it indicates that it belongs to the same association dimension as the center term, and it is classified into that niche and labeled. If the similarity is not less than the preset radius, it is determined that it does not belong to that niche, and it is not classified for the time being, leaving it for subsequent processing.
[0036] After classifying the first niche term, select the term with the highest association strength from the remaining unlabeled feature terms and determine it as the new niche center. Repeat step S333 to calculate the feature similarity between the remaining unlabeled terms and the new center term. Unlabeled terms with similarity less than the preset niche radius are assigned to the new niche, and these assigned terms are labeled, completing the construction of the new niche. Check the labeling status of all feature seed terms. If all terms have been labeled, all terms have been assigned to their corresponding niches, and the niche classification work is complete. If there are still unlabeled terms, return to step S334 to continue selecting new niche centers and classifying terms until all feature terms are labeled and classification is complete.
[0037] In this optional embodiment, among the remaining unlabeled feature terms, the unlabeled feature term with the highest association strength value is selected as the new niche center. Simultaneously, unlabeled feature terms with feature similarity less than a preset niche radius are included in this niche and labeled accordingly. S3341. Traverse all unlabeled feature terms, select the feature term with the highest association strength value, determine the term as the central feature term of the new niche, and complete the labeling. S3342. Calculate the feature similarity between the remaining unlabeled feature terms and the feature terms of the new niche center one by one, and obtain the similarity value corresponding to each unlabeled feature term; S3343. Compare the similarity value corresponding to each unlabeled feature term with the preset niche radius to determine whether the unlabeled feature term meets the similarity condition for being classified into the current niche. S3344. Assign unlabeled feature terms that meet the similarity criteria to the current niche, perform labeling operations on the assigned unlabeled feature terms, and complete the classification of the niche.
[0038] Specifically, it iterates through all unlabeled feature terms, reads the association strength value corresponding to each term, and filters out the feature term with the highest association strength value by comparing the values. This term is determined as the central feature term of the current new niche and is then labeled to indicate that the term has completed the center location and classification status registration, thus avoiding repeated participation in the subsequent center selection process.
[0039] After identifying the central feature term of the new niche, the feature similarity between each of the remaining unlabeled feature terms and the central feature term is calculated. Through a comprehensive comparison of feature attributes, association logic, and semantic dimensions, a similarity value is obtained for each unlabeled feature term, providing a quantitative basis for subsequent classification. The calculated similarity value for each unlabeled feature term is compared with a preset niche radius (preset range 0.2-0.4, with 0.3 preferred to distinguish feature boundaries of different association dimensions). Based on the comparison results, it is determined whether each unlabeled feature term meets the similarity condition for inclusion in the current niche, thus clarifying the criteria for term classification.
[0040] Unlabeled feature terms with similarity values less than the preset niche radius and meeting the classification criteria are uniformly assigned to the current new niche. All feature terms assigned to this niche are then labeled, and their classification status is updated. This completes the entire niche construction and term classification process, ensuring that feature terms of the same related dimensions are centrally classified. Before all niches are constructed, the above steps are repeated, continuously selecting new centers, calculating similarities, and completing term classification until all feature terms are labeled and classified.
[0041] In this optional embodiment, based on the core feature set, a multi-dimensional risk assessment index for adverse drug reactions is constructed. The credibility of the drug risk assessment index is quantified by confidence grading. The comprehensive risk assessment result of adverse drug reactions is obtained through multi-feature fusion calculation, including: S41. Set the number of risk assessment indicators, the maximum number of iterations, the initial confidence threshold, and the maximum feature fusion period to determine the basic calculation parameters for risk assessment. S42. Initialize multi-dimensional drug risk assessment indicators based on the core feature set, determine the credibility judgment rules of the indicators, and complete the initial configuration for iterative calculation. S43. Quantify the credibility of each drug risk assessment indicator according to the preset confidence level to obtain the risk confidence value corresponding to each drug risk assessment indicator. S44. Update the strength of the high-confidence risk assessment indicators. If the new combination of risk assessment indicators is better or the confidence of the risk assessment indicators is higher, then adopt the combination of risk assessment indicators and expand the feature dimensions. S45. Update the comprehensive risk calculation results through multi-feature fusion, select and retain high-confidence indicators, and optimize the risk assessment calculation process; S46. Determine whether the maximum feature fusion period has been reached. If it has, reset the feature fusion status and update the confidence level and risk calculation direction of the risk assessment indicators. S47. Determine if the maximum number of iterations has been reached. If it has, output the comprehensive risk assessment result of adverse drug reactions; otherwise, continue iterative calculation.
[0042] Specifically, first, clarify the basic calculation parameters for risk assessment and reasonably set various preset parameters (the number of risk assessment indicators ranges from 10 to 15, with 12 preferred; the maximum number of iterations ranges from 50 to 80, with 60 preferred; the initial confidence threshold ranges from 0.4 to 0.6, with 0.5 preferred; the maximum feature fusion period ranges from 8 to 12 iteration steps, with 10 iteration steps preferred). All parameters can be flexibly adjusted according to the actual assessment accuracy requirements, providing a stable basis for subsequent risk assessment calculations and ensuring that the calculation process is orderly and controllable. Then, retrieve the core feature set generated earlier. Based on the associated features of drugs, symptoms, populations, and medications in the core feature set, initialize multi-dimensional drug risk assessment indicators, covering dimensions such as drug attribute risk, population suitability risk, medication behavior risk, and adverse reaction association risk. Simultaneously, determine the indicator credibility judgment rules, that is, clarify the correspondence between confidence values and credibility levels, for subsequent indicator credibility quantification, completing all initial configurations for iterative calculations and ensuring that subsequent iteration processes have clear execution standards.
[0043] According to the preset confidence level rules, the confidence level is divided into three levels: high, medium, and low, corresponding to different quantification intervals. The credibility of each drug risk assessment indicator after initialization is quantified. Combining the correlation strength information in the core feature set, the specific risk confidence value corresponding to each drug risk assessment indicator is calculated, quantifying the indicator credibility and providing quantitative support for subsequent indicator optimization and risk calculation. Based on the risk confidence value of each indicator, high-confidence risk assessment indicators are updated, with their influence strength not lower than the initial confidence threshold, giving them a higher weight in the comprehensive risk calculation. At the same time, the current combination of risk assessment indicators is compared with the newly generated combination of indicators. If the new combination has better assessment accuracy or higher overall confidence, the new combination is adopted, and relevant feature dimensions are expanded simultaneously to enrich the comprehensiveness of the risk assessment.
[0044] The comprehensive risk assessment of adverse drug reactions is updated by employing multi-feature fusion and combining the effect strength and risk confidence of each indicator. During the calculation process, high-confidence indicators are selectively retained while low-confidence redundant indicators are eliminated, simplifying the calculation process, reducing interference, optimizing the overall risk assessment calculation process, and improving calculation efficiency and accuracy. The iteration step size of the current feature fusion is tracked in real time to determine if the preset maximum feature fusion cycle has been reached. If it has, the feature fusion state is reset, redundant fusion records are cleared, and the confidence values and risk calculation directions of the risk assessment indicators are updated based on the current assessment results to ensure the rationality of subsequent fusion calculations and avoid calculation deviations caused by long-term fusion. The current iteration count is counted and compared with the preset maximum iteration count. If the maximum iteration count has been reached, the iteration calculation stops, and the final comprehensive risk assessment result of adverse drug reactions is output, clarifying the risk level and key influencing factors of adverse drug reactions. If the maximum iteration count has not been reached, the process returns to step S43 to continue confidence quantification, indicator updates, and fusion calculations until the iteration termination condition is met.
[0045] Specifically, this invention obtains the comprehensive risk assessment results of adverse drug reactions through the lightning search algorithm. The lightning search algorithm is a heuristic optimization algorithm that simulates the formation, conduction, step-by-step advancement and target tracking process of lightning in nature. It follows the natural evolution law of leader channel formation, step-by-step leader advancement, stationary point optimization update and global convergence termination. It achieves the global optimal solution through dynamic step size adjustment, multi-directional parallel search, confidence-driven optimization and iterative convergence mechanism.
[0046] In this invention, the Lightning Search algorithm adopts a tiered iterative optimization architecture. It uses drug risk assessment indicators as search units, indicator confidence and risk fitting accuracy as the basis for advancement, and dynamic step size and feature fusion cycle as constraints. Through four stages—leader channel formation, tiered leader advancement, stationary point optimization update, and global convergence termination—it achieves the optimal solution for the comprehensive risk assessment result. The specific usage process and operation details are as follows: Based on the 28 high-value features of the core feature set (covering four dimensions: drug, symptom, population, and medication), the number of risk assessment indicators is set to 12 (3 for each dimension, matching the distribution of the core feature dimensions). At the same time, the maximum number of iterations is preset to 60 (balancing convergence accuracy and computational efficiency), the initial confidence threshold is 0.5 (used to initially distinguish between high / low confidence indicators), and the maximum feature fusion cycle is 10 iteration steps (controlling the feature fusion frequency to avoid redundancy caused by excessive fusion).
[0047] Based on 28 core features of the core feature set, 12 multi-dimensional drug risk assessment indicators were initialized, corresponding to four categories: drug dosage risk, population suitability risk, symptom association risk, and medication behavior risk. At the same time, the confidence judgment rules of the indicators were determined (that is, the correspondence between confidence value and confidence level is clarified, i.e., confidence greater than or equal to 0.7 is high confidence, 0.4-0.7 is medium confidence, and less than 0.4 is low confidence). The initialized indicators and judgment rules were entered into the algorithm to complete all the initial configurations for iterative calculations, and the initial search channel of the lightning search algorithm was constructed to ensure that there are clear execution standards for subsequent iterations. The Lightning Search algorithm invokes preset confidence grading rules to perform confidence quantification calculations on each of the 12 initialized risk assessment indicators. Combining the correlation strength values of each feature in the core feature set, the algorithm obtains the specific risk confidence value corresponding to each risk assessment indicator through its built-in confidence calculation logic (e.g., initial confidence value of 0.52 for drug dosage risk indicator and initial confidence value of 0.48 for population suitability risk indicator). This risk confidence value serves as the core basis for the algorithm's tiered advancement, completing the first round of confidence quantification and providing quantitative support for subsequent indicator optimization.
[0048] The Lightning Search algorithm automatically updates the strength of high-confidence risk assessment indicators (confidence greater than or equal to 0.5, i.e., the initial confidence threshold) based on the risk confidence values of each indicator, increasing the weight of high-confidence indicators by 20% to give them a higher weight in the overall risk calculation. Simultaneously, the algorithm automatically compares the current combination of 12 risk assessment indicators with newly generated indicator combinations (adding potentially high-value features not included in the initial indicators from the core feature set). If the assessment accuracy of the new combination improves by more than 10% or the overall average confidence of the indicators improves by more than 0.05, the new combination is adopted, and related feature dimensions are expanded simultaneously (e.g., expanding from drug dosage risk indicators to drug dosage-liver and kidney function correlation risk indicators), achieving adaptive optimization of indicator combinations and propelling the Lightning Search algorithm towards a better direction. The Lightning Search algorithm uses multi-feature fusion to combine the strength of each indicator's effect and risk confidence level to update the comprehensive risk calculation results of adverse drug reactions. During the calculation process, the Lightning Search algorithm automatically selects high-confidence indicators with a confidence level greater than or equal to 0.5 and removes low-confidence redundant indicators with a confidence level less than 0.4, simplifying the calculation process, reducing interference, and optimizing the overall risk assessment calculation process. For example, after the first round of fusion, two low-confidence indicators are removed, 10 effective indicators are retained, the comprehensive risk value is recalculated, and the comprehensive risk result is optimized and updated.
[0049] The Lightning Search algorithm continuously monitors the iteration step size of the current feature fusion process. After each iteration, the step size is automatically accumulated. When the step size reaches the preset maximum feature fusion cycle (10 iteration steps), the Lightning Search algorithm automatically resets the feature fusion state, clears redundant fusion records, and updates the confidence values of all risk assessment indicators based on the evaluation results of the current 10 iterations (e.g., increasing the confidence of high-performing indicators by 0.03-0.05 and decreasing the confidence of poor-performing indicators by 0.02-0.04). The risk calculation direction is adjusted synchronously to ensure the rationality of subsequent fusion calculations, avoid calculation deviations caused by long-term fusion, and maintain the correct optimization orientation of the Lightning Search algorithm.
[0050] The Lightning Search algorithm continuously counts the current iteration count. After each round of confidence quantification, indicator update, and fusion calculation, the iteration count is incremented by 1. The current iteration count is compared to the preset maximum iteration count (60 times). If 60 iterations have been reached, the Lightning Search algorithm terminates the iteration calculation and automatically outputs the final comprehensive risk assessment result for adverse drug reactions, clearly defining the risk level and key influencing factors (e.g., a final comprehensive risk value of 0.76 corresponds to high risk, with key influencing factors being high drug dosage and high correlation with liver and kidney function). If 60 iterations have not been reached, the Lightning Search algorithm automatically returns to the confidence quantification stage and continues confidence quantification, indicator update, and fusion calculation until the iteration termination condition is met. The iterative process of the Lightning Search algorithm is fully specified. The input data comes directly from the core feature set of adverse drug reactions obtained through screening. The core feature set has undergone semantic standardization, normalization, and association strength calculation. The data sources cover desensitized adverse drug reaction data from multiple medical institutions, drug regulatory agencies, and pharmaceutical companies. All data have undergone missing value imputation, outlier removal, duplicate data deletion, terminology standardization, and numerical normalization to ensure that the input of the Lightning Search algorithm is stable, standardized, and traceable. The Lightning Search algorithm adopts a target evaluation mechanism with confidence level, indicator effect strength, and comprehensive risk fitting accuracy as its core. The optimization objectives are to retain high-confidence indicators, remove low-confidence indicators, and achieve smooth convergence of comprehensive risk values. It can achieve adaptive selection without the need for traditional loss functions. The optimization direction is completely consistent with the business objectives of adverse drug reaction risk assessment. It has the characteristics of high fitting accuracy, fast convergence speed, and strong stability, and is highly matched with the nonlinear coupling law of multidimensional risk factors of adverse drug reactions.
[0051] The key parameters of the Lightning Search algorithm are clearly defined and have sufficient selection criteria. There are 12 risk assessment indicators, a maximum number of iterations of 60, an initial confidence threshold of 0.5, and a maximum feature fusion period of 10 iteration steps. All parameters are determined in combination with the accuracy requirements, computational complexity, and convergence stability of adverse drug reaction risk assessment. They have clear adjustment ranges (e.g., the maximum number of iterations can be adjusted between 50 and 70, and the initial confidence threshold can be adjusted between 0.4 and 0.6) and optimization strategies (the parameters are dynamically adjusted according to the convergence speed of previous iterations).
[0052] The input data for the Lightning Search algorithm includes the core feature set of adverse drug reactions, the number of risk assessment indicators, the maximum number of iterations, the initial confidence threshold, the maximum feature fusion period, and the indicator credibility judgment rules. All inputs are directly derived from the adverse drug reaction monitoring and feature screening process, and are highly relevant to the application scenario, reusable, and verifiable. The output data of the Lightning Search algorithm includes the confidence values of each drug risk assessment indicator, the optimized combination of risk assessment indicators, the record of the comprehensive risk calculation process, and the final comprehensive risk assessment result of adverse drug reactions. The output and input form a complete logical closed loop through confidence grading, effect intensity updating, multi-feature fusion, iterative advancement, and convergence judgment, which is directly used for subsequent adverse drug reaction risk warning and warning report generation.
[0053] In this optional embodiment, the comprehensive risk calculation result is updated through multi-feature fusion, and high-confidence indicators are selectively retained. The optimization of the risk assessment calculation process includes: S451. Normalize the feature values of each risk assessment indicator in the core feature set to unify risk assessment indicators of different dimensions and scales into the same numerical range, ensuring that each risk assessment indicator is under equivalent calculation conditions when participating in the fusion. S452. Incorporate the normalized multiple risk assessment indicators into the same feature combination, determine all effective risk assessment indicators participating in the comprehensive risk calculation, and form a basic feature set for average value fusion. S453. Calculate the arithmetic mean of the feature values of all valid risk assessment indicators in the basic feature set, and use this mean as the current fusion result to obtain the preliminary comprehensive risk calculation value of adverse drug reactions. S454. Based on the confidence level results, retain high-confidence risk assessment indicators, remove low-confidence risk assessment indicators, recalculate the average value based on the selected effective risk assessment indicators, and update and optimize the final comprehensive risk calculation results.
[0054] Specifically, the original feature values of each risk assessment indicator in the core feature set are retrieved. Addressing the issue of varying dimensions and inconsistent units among different risk assessment indicators, a pre-defined normalization rule is employed. This rule, a minimum-maximum normalization method, maps the feature values of each indicator to a uniform 0-1 value range. Normalization is performed on the feature values of all risk assessment indicators. Specifically, this involves: first, traversing each valid risk assessment indicator within the basic feature set, extracting the maximum and minimum values of the original feature value for that indicator across all samples, and constructing a numerical range mapping relationship; then, for the original feature value of that indicator in a single sample, calculating the normalized value using the formula: Normalized Value = (Original Feature Value - Minimum Indicator Value) / (Maximum Indicator Value - Minimum Indicator Value), uniformly mapping the feature values of each indicator to a 0-1 value range. This unifies risk indicators of different magnitudes and attributes to the same numerical range, eliminating calculation biases caused by differences in units, ensuring that each risk assessment indicator has equivalent calculation conditions in subsequent fusion calculations, and guaranteeing the fairness and accuracy of the fusion results.
[0055] After completing the normalization of all risk assessment indicators, and combining the previously determined criteria for valid risk assessment indicators, the normalized multiple risk assessment indicators are incorporated into the same feature set. The validity of each indicator is then verified one by one, and redundant indicators that have been previously determined to be invalid are eliminated. Finally, all valid risk assessment indicators participating in the comprehensive risk calculation are determined. Based on this, a basic feature set for average value fusion is formed, providing a standardized feature input range for subsequent fusion calculations.
[0056] The normalized eigenvalues of all valid risk assessment indicators within the basic feature set are uniformly calculated, and the arithmetic mean of all eigenvalues is obtained. This average is used as the preliminary result of the current multi-feature fusion to calculate the initial value for the comprehensive risk calculation of adverse drug reactions, completing the preliminary fusion process for comprehensive risk calculation and providing basic reference data for subsequent optimization. Based on the previously preset confidence level results (the risk confidence level is divided into three levels: high, medium, and low, where a confidence level of not less than 0.7 is considered a high confidence indicator, and a confidence level of less than 0.4 is considered a low confidence indicator), each risk assessment indicator is screened, retaining only high confidence risk assessment indicators and eliminating low confidence and invalid risk assessment indicators. Based on the remaining valid high confidence risk assessment indicators after screening, the arithmetic mean of their normalized eigenvalues is recalculated to update and optimize the final comprehensive risk calculation result of adverse drug reactions.
[0057] In this optional embodiment, based on the confidence level results, high-confidence risk assessment indicators are retained, low-confidence risk assessment indicators are removed, and the average value is recalculated based on the screened effective risk assessment indicators to update and optimize the final comprehensive risk calculation result, including: S4541. Retrieve the confidence level results of each risk assessment indicator, preset the division criteria for high confidence and low confidence indicators, and distinguish between risk assessment indicators to be retained and those to be removed. S4542. Based on the preset classification criteria, retain all high-credibility risk assessment indicators and remove low-credibility risk assessment indicators to obtain a set of effective risk assessment indicators after screening. S4543. Organize the feature values of the selected effective risk assessment indicators, unify the calculation method, and recalculate the arithmetic mean of all the feature values of the effective risk assessment indicators. S4544. Use the recalculated arithmetic mean as the updated comprehensive risk calculation value, replace the original result, and complete the optimization of the comprehensive risk calculation result for adverse drug reactions.
[0058] Specifically, the confidence levels of previously completed risk assessment indicators are retrieved to clarify the corresponding confidence values and classification levels for each indicator. Simultaneously, a pre-defined standard for classifying high-confidence and low-confidence indicators is established (the pre-defined standard is: confidence values of 0.7 or higher are considered high-confidence indicators; confidence values below 0.4 are considered low-confidence indicators; and values between 0.4 and 0.7 are considered medium-confidence indicators, which are not included in the screening scope for the time being). This standard clearly distinguishes between high-confidence indicators to be retained and low-confidence indicators to be eliminated, providing a clear basis for subsequent indicator screening.
[0059] Based on a pre-defined credibility classification standard, all risk assessment indicators participating in the comprehensive risk calculation are individually screened. Indicators meeting high credibility standards are retained, while low-credibility indicators are completely eliminated. Simultaneously, interference from medium-credibility indicators is excluded, resulting in a final set of effective risk assessment indicators. This ensures that subsequent calculations are based solely on high-value, high-credibility indicators. The feature values of the selected effective risk assessment indicators are then uniformly organized, and the normalization status of each indicator is verified. This ensures that all indicators have undergone minimum-maximum normalization and are within the 0-1 value range, unifying the calculation caliber and avoiding calculation deviations caused by inconsistent calibers. The normalized feature values of each effective indicator are re-retrieved, and the arithmetic mean of all feature values is calculated to ensure a standardized calculation process and accurate results. The recalculated arithmetic mean is used as the updated comprehensive risk calculation value for adverse drug reactions, directly replacing the initial comprehensive risk calculation value obtained from multi-feature fusion. The update process and its basis are recorded simultaneously, optimizing the comprehensive risk calculation results and ensuring that the final output risk calculation results are more credible and valuable.
[0060] In this optional embodiment, the risk warning of the corresponding level is initiated based on the comprehensive risk assessment results of adverse drug reactions, and the warning information is pushed based on user roles to generate an adverse drug reaction warning report, including: S51. Based on the comprehensive risk assessment results of adverse drug reactions, determine the current risk level according to the preset risk level classification rules, and determine the early warning activation conditions and early warning type that match the risk level. S52. Based on the responsibilities of different user roles, formulate differentiated early warning information push strategies and determine the early warning content and information display format required for each user role. S53. In accordance with the established push strategy, push the warning information of the corresponding level to the terminal of the corresponding user role to ensure that the warning information accurately reaches the relevant personnel with the authority to handle the situation. S54. Integrate risk assessment results, risk levels, early warning content, and handling authority information to generate a structured adverse drug reaction early warning report in a fixed format and complete the archiving.
[0061] Specifically, the system retrieves the output comprehensive risk assessment results of adverse drug reactions and classifies risks into four levels—general, moderate, severe, and extremely severe—based on preset risk level classification rules. The current adverse drug reaction risk level is determined by combining the comprehensive risk calculation value. At the same time, the system determines the warning activation conditions and warning types that match each risk level: blue warning for general risk, yellow warning for moderate risk, orange warning for severe risk, and red warning for extremely severe risk. The system clarifies the activation thresholds and response requirements for each level of warning to ensure that the warning level matches the actual risk.
[0062] We identified all relevant user roles and their corresponding responsibilities, developed a differentiated early warning information push strategy, and distinguished between different roles such as clinicians, pharmacists, drug monitoring personnel, and managers. We clarified that clinicians need to receive information on risky drugs, adverse reaction symptoms, and treatment suggestions; pharmacists need to receive medication adjustment prompts; and managers need to receive an overall risk overview. At the same time, we determined the information display format for each role to ensure that the content is relevant to their job responsibilities and that the format is concise and easy to understand.
[0063] According to the established push strategy, specifically: push notifications are categorized by user role, with clear push channels, push time limits, and content priorities. Priority is given to clinicians and pharmacists, with a push time limit not exceeding 5 minutes. Management personnel receive push notifications within the standard time limit. Corresponding warning information is targeted to the work terminals (e.g., mobile phones, office computers) of the relevant user roles, with delivery reminders set up and real-time feedback on reception status. Users who do not receive the notification in time will receive a second push, ensuring that warning information accurately reaches relevant personnel with the authority to handle the situation, thus gaining time for rapid response. The system integrates risk assessment results, current risk level, warning content, warning type, and the handling authority information for each role.
[0064] Following a preset fixed format, specifically: Report Title - Risk Overview - Risk Level and Warning Type - Warning Content - Handling Authority - Assessment Basis - Archiving Identifier, a unified electronic document format is adopted, with standardized content layout and complete elements, generating a structured adverse drug reaction warning report. After generation, it is automatically archived electronically, facilitating subsequent querying, tracing, and review, and ensuring that warning-related information is traceable and manageable.
[0065] According to another embodiment of the invention, such as Figure 2 As shown, a federated learning-based adverse drug reaction data collection and processing system is also provided, which includes: Data acquisition module 1 is used to acquire adverse drug reaction data from multiple sources and to preprocess the adverse drug reaction data from multiple sources. Data processing module 2 is used to perform semantic standardization on the preprocessed multi-source adverse drug reaction data to obtain a complete multi-source fusion dataset; Feature analysis module 3 is used to extract an initial feature term population from the complete multi-source fusion dataset, screen out core seed terms and supplementary seed terms, classify niche categories, expand the dimensions of all feature seed terms, analyze their potential associations and complete their diffusion and reproduction, and finally form a core feature set that affects the occurrence of adverse drug reactions. Risk assessment module 4 is used to construct multi-dimensional risk assessment indicators for adverse drug reactions based on the core feature set, quantify the credibility of the drug risk assessment indicators through confidence grading, and obtain the comprehensive risk assessment result of adverse drug reactions through multi-feature fusion calculation. The early warning information generation module 5 is used to initiate corresponding level risk warnings based on the comprehensive risk assessment results of adverse drug reactions, and to push early warning information based on user roles, and generate adverse drug reaction early warning reports.
[0066] To facilitate understanding of the above technical solutions of the present invention, the following provides a detailed description of the collection and processing of adverse drug reaction data based on federated learning in practice.
[0067] I. Acquisition and preprocessing of adverse drug reaction data from multiple sources.
[0068] By connecting to the adverse reaction data of desensitized drugs from 3 cooperating medical institutions, 1 provincial drug regulatory agency, and 1 core pharmaceutical company through the federated learning node, a total of 5,200 original case data were collected. The data fields cover 12 categories of core information, including basic patient information, medication records, adverse reaction symptoms, diagnosis results, and liver and kidney function indicators.
[0069] Initiate the preprocessing procedure and strictly follow these standardized procedures: Missing value handling: For non-core missing fields such as medication course and liver and kidney function classification, the mode of the same field was used for imputation, filling a total of 386 missing data. For core missing fields such as patient allergy history, 124 data were marked as unknown and were not included in the core feature calculation.
[0070] Outlier removal: Outliers in numerical fields such as daily drug dosage and symptom duration are removed according to the 3σ principle. For example, 45 abnormal medication records with daily dosages exceeding three times the standard deviation of the mean are removed to ensure data rationality.
[0071] Duplicate data removal: Based on triple matching of patient ID, drug name, and adverse reaction symptoms, 280 duplicate cases were removed.
[0072] After preprocessing, 4,795 valid data entries were obtained. All data were stored in the local databases of each partner institution. Only the encrypted feature hash value was uploaded to the federated aggregation node. No original sensitive information was disclosed throughout the process, thus meeting privacy compliance requirements.
[0073] II. Semantic standardization processing and multi-source fusion dataset generation.
[0074] Semantic standardization is performed on the preprocessed multi-source data. The core operations are as follows: Drug entity standardization: The WHO Drug Dictionary is used to encode and map drug names from various institutions, unifying drugs with consistent descriptions across different institutions into corresponding codes, and clarifying drug categories, dosage forms, and specifications.
[0075] Symptom entity standardization: The Med DRA medical dictionary was used to standardize the descriptions of adverse reaction symptoms, and non-standard symptom descriptions were converted into standard terms. A total of 897 symptom descriptions were standardized.
[0076] Numerical field normalization: Min-max normalization is performed on numerical fields such as age, medication dosage, and symptom duration, uniformly mapping them to the numerical range of 0 to 1. The calculation formula is: the normalized value equals the original value minus the field's minimum value, divided by the field's maximum value minus the field's minimum value. Taking the age field as an example, the original age range is 18 to 85 years old, and the normalized range is 0 to 1, eliminating the impact of differences in the units of measurement of different fields on subsequent calculations.
[0077] Federated data fusion: Each institution only uploads the standardized feature vector and encrypted hash value. The federated aggregation node completes the cross-institutional data splicing and finally generates a complete multi-source fusion dataset with a total of 4,795 standardized samples, covering 28 standardized fields in 4 major categories: drugs, symptoms, population, and medication. The dataset structure is shown in Table 1.
[0078] Table 1. Structure of the Multi-Source Fusion Dataset
[0079] III. Feature extraction, niche segmentation and construction of core feature set.
[0080] Feature extraction and core feature selection were performed based on a multi-source fusion dataset. The specific values and operations are as follows: Initial feature extraction: Through the federated semantic association algorithm, 92 initial feature terms related to drugs, symptoms, populations and medication were extracted from 4795 samples, covering core dimensions such as drug category, symptom name, age group and medication frequency.
[0081] Seed Feature Screening: The association strength between each initial feature term and adverse drug reactions was calculated using the point mutual information method. The features were sorted in descending order of association strength, and the top 20 features with the strongest association strength were selected as core seed terms. From the remaining 72 initial feature terms, 10 potential related terms were randomly selected as supplementary seed terms, ensuring a total of 30 core and supplementary seeds, balancing feature coreness and richness.
[0082] Niche segmentation and expansion: A niche radius of 0.3 was preset, and the niches were divided into four categories based on feature attributes: drug-related, symptom-related, population-related, and medication-related. Thirty seed terms were traversed to complete niche clustering. Based on each niche, feature dimensions were expanded. Potential association features of seed terms were mined from multi-source datasets, resulting in 48 expanded features. Association degree filtering was performed, with a preset association degree threshold of 0.4. Twelve redundant features with an association degree less than 0.4 after expansion were removed, ultimately retaining 36 effective expanded features.
[0083] Core Feature Set Generation: The core seed, supplementary seed, and effective extended features are summarized, and duplicate entries are removed to finally form a core feature set that affects the occurrence of adverse drug reactions, consisting of 28 high-value features, as shown in Table 2.
[0084] Table 2 Core Feature Set
[0085] IV. Construction of multi-dimensional risk assessment indicators and comprehensive risk calculation.
[0086] Risk assessment indicators are constructed based on the core feature set and federated iterative calculations are performed. The specific parameters and values are as follows: Basic parameter configuration: 12 preset risk assessment indicators, 60 maximum iterations, 0.5 initial confidence threshold, and 10 iteration steps for maximum feature fusion period. Initialize multi-dimensional risk assessment indicators, covering 4 categories: drug dosage risk, population suitability risk, symptom association risk, and medication behavior risk, for a total of 12 indicators.
[0087] Confidence rating quantification: The confidence rating rules are set in advance, with a confidence rating of 0.7 or higher being high confidence, 0.4 to 0.7 being medium confidence, and less than 0.4 being low confidence. The confidence of 12 risk assessment indicators is quantified. After federated iterative calculation, the confidence results of each indicator are shown in Table 3.
[0088] Table 3 Confidence Levels of Risk Assessment Indicators
[0089] Multi-feature fusion calculation: Step 1: Normalization. The original feature values of the 12 risk assessment indicators are normalized using minimum-maximum normalization to unify them to the range of 0 to 1. Preliminary fusion. The arithmetic mean of the normalized values of the 12 indicators is calculated, yielding a preliminary comprehensive risk value of 0.58. Confidence screening. Low-confidence indicators are removed, retaining 10 high- and medium-confidence indicators. The arithmetic mean is recalculated, resulting in an updated comprehensive risk value of 0.65. Iterative optimization. The confidence quantification, feature fusion, and result update process is repeated 60 times, finally converging to a comprehensive risk assessment result of 0.76 for adverse drug reactions, corresponding to a high risk level. The preset risk levels are: less than 0.4 for low risk, 0.4 to 0.6 for medium risk, 0.6 to 0.8 for high risk, and greater than or equal to 0.8 for extremely high risk.
[0090] V. Risk warning activation, information push and warning report generation.
[0091] Based on the comprehensive risk assessment results, an early warning will be initiated, and information will be pushed out and reports will be generated. The specific steps are as follows: Risk warning determination: Based on the comprehensive risk value of 0.76, and in accordance with the preset risk level classification rules, a high-risk orange warning is activated. The warning activation condition is that the comprehensive risk value is greater than or equal to 0.6, and the warning response time limit is to handle the situation within 15 minutes.
[0092] Differentiated alert push notifications: Distributed push notifications are tailored to user roles, including clinicians, pharmacists, drug regulatory personnel, and pharmaceutical company quality managers, clearly defining the push channels, time limits, and content priorities. Clinicians receive the notifications within 5 minutes, pharmacists within 10 minutes, and drug regulatory personnel and pharmaceutical company quality managers within 15 minutes, achieving 100% push reach.
[0093] Early Warning Report Generation and Archiving: Structured adverse drug reaction early warning reports are generated according to a fixed format. Each report includes elements such as report number, data source, comprehensive risk value, risk level, core risk characteristics, treatment recommendations, push notification history, and archiving time, as shown in Table 4. After generation, reports are automatically encrypted and stored on the federated aggregation node and in the local archives of each institution, ensuring report traceability and verifiability.
[0094]
[0095] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for collecting and processing adverse drug reaction data based on federated learning, characterized in that, The method includes: S1. Obtain adverse drug reaction data from multiple sources and preprocess the adverse drug reaction data from multiple sources; S2. Perform semantic standardization on the preprocessed multi-source adverse drug reaction data to obtain a complete multi-source fusion dataset; S3. Extract the initial feature term population from the complete multi-source fusion dataset, screen out the core seed terms and supplementary seed terms, divide the niche categories, expand the dimensions of all feature seed terms, analyze their potential associations and complete their diffusion and reproduction, and finally form the core feature set that affects the occurrence of adverse drug reactions. S4. Based on the core feature set, construct a multi-dimensional risk assessment index for adverse drug reactions, quantify the credibility of the drug risk assessment index through confidence grading, and obtain the comprehensive risk assessment result of adverse drug reactions through multi-feature fusion calculation. S5. Based on the comprehensive risk assessment results of adverse drug reactions, initiate the corresponding level of risk warning, and formulate and push warning information based on user roles, and generate adverse drug reaction warning reports.
2. The method for collecting and processing adverse drug reaction data based on federated learning according to claim 1, characterized in that, The step of semantically standardizing the preprocessed multi-source adverse drug reaction data to obtain a complete multi-source fusion dataset includes: S21. Perform terminology alignment on the preprocessed multi-source adverse drug reaction data to unify the expression of drug, symptom, and population-related data and eliminate semantic differences between cross-data sources. S22. Standardize drug entities according to the preset drug coding system, clarify drug identification, classification and attribute information, so that drug data from different sources have a unified basis for identification; S23. Standardize and convert adverse reaction symptoms according to a pre-set authoritative terminology database to unify symptom names and descriptions; S24. Perform semantic analysis on the relationships between data and establish stable correspondence rules between drugs, symptoms, populations and medication information; S25. Integrate the various types of data that have completed semantic standardization processing to form a multi-source fusion dataset of adverse drug reactions with standardized structure, unified terminology, and complete association.
3. The method for collecting and processing adverse drug reaction data based on federated learning according to claim 1, characterized in that, The process involves extracting an initial population of feature terms from a complete multi-source fusion dataset, selecting core seed terms and supplementary seed terms, classifying niche categories, expanding the dimensions of all feature seed terms, analyzing their potential associations, and completing their diffusion and propagation. The final result is a core feature set influencing adverse drug reactions, including: S31. Extract an initial population of drug-related and adverse reaction-related feature terms from the complete multi-source fusion dataset; S32. Select the most relevant feature terms with the largest preset number of weight values from the initial feature term population related to drugs and adverse reactions as core seed terms, and then randomly select a preset number of potential related supplementary terms from the remaining feature terms as supplementary seed terms. S33. Sort all core feature seed terms and supplementary feature seed terms according to their association strength, and classify them into niche categories with different association dimensions; S34. Based on each microhabitat, expand the feature dimensions of all feature seed entries, analyze the potential association between feature seeds and adverse drug reactions, and complete the growth, reproduction and association diffusion of feature seeds. S35. Reassess the association strength of the expanded feature term population, remove redundant terms with association strength below the preset threshold, and retain high-value associated terms. S36. Count the total number of current feature word populations. If it does not exceed the preset maximum number of core feature populations, return to step S34 to continue expanding. If it has exceeded the limit, select the word with the best fit as the core candidate feature according to the fitness of each niche association from high to low. S37. Summarize all core candidate features to form a core feature set that affects the occurrence of adverse drug reactions.
4. The method for collecting and processing adverse drug reaction data based on federated learning according to claim 3, characterized in that, The process of sorting all core feature seed terms and supplementary feature seed terms according to their association strength and classifying them into niche categories with different association dimensions includes: S331. Sort all core feature seed terms and supplementary feature seed terms in descending order according to the association strength value of the feature terms. If the total number of feature terms is greater than the preset maximum population size, only the feature terms with the highest ranking are retained. S332. Set the center of the first niche as the feature term with the highest association strength value, and label the feature term; S333. For the remaining unlabeled feature terms, if the feature similarity between the unlabeled feature term and the current niche center feature term is less than the preset niche radius, then it is included in the niche and labeled; otherwise, it does not belong to the niche. S334. Among the remaining unlabeled feature words, select the unlabeled feature word with the largest association strength value as the new niche center, and at the same time, classify and label the unlabeled feature words whose feature similarity is less than the preset niche radius into the niche. S335. If all feature terms have been labeled, the niche classification ends; otherwise, return to step S334 to continue classification.
5. The method for collecting and processing adverse drug reaction data based on federated learning according to claim 4, characterized in that, The step involves selecting the unlabeled feature word with the highest association strength value from the remaining unlabeled feature words as the new niche center, and simultaneously including and labeling unlabeled feature words with feature similarity less than the preset niche radius into this niche. S3341. Traverse all unlabeled feature terms, select the feature term with the highest association strength value, determine the term as the central feature term of the new niche, and complete the labeling. S3342. Calculate the feature similarity between the remaining unlabeled feature terms and the feature terms of the new niche center one by one, and obtain the similarity value corresponding to each unlabeled feature term; S3343. Compare the similarity value corresponding to each unlabeled feature term with the preset niche radius to determine whether the unlabeled feature term meets the similarity condition for being classified into the current niche. S3344. Assign unlabeled feature terms that meet the similarity criteria to the current niche, perform labeling operations on the assigned unlabeled feature terms, and complete the classification of the niche.
6. The method for collecting and processing adverse drug reaction data based on federated learning according to claim 1, characterized in that, The method involves constructing a multi-dimensional risk assessment index for adverse drug reactions based on a core feature set. The reliability of the drug risk assessment index is quantified through confidence grading. The comprehensive risk assessment result for adverse drug reactions is obtained through multi-feature fusion calculation, including: S41. Set the number of risk assessment indicators, the maximum number of iterations, the initial confidence threshold, and the maximum feature fusion period to determine the basic calculation parameters for risk assessment. S42. Initialize multi-dimensional drug risk assessment indicators based on the core feature set, determine the credibility judgment rules of the indicators, and complete the initial configuration for iterative calculation. S43. Quantify the credibility of each drug risk assessment indicator according to the preset confidence level to obtain the risk confidence value corresponding to each drug risk assessment indicator. S44. Update the strength of the high-confidence risk assessment indicators. If the new combination of risk assessment indicators is better or the confidence of the risk assessment indicators is higher, then adopt the combination of risk assessment indicators and expand the feature dimensions. S45. Update the comprehensive risk calculation results through multi-feature fusion, select and retain high-confidence indicators, and optimize the risk assessment calculation process; S46. Determine whether the maximum feature fusion period has been reached. If it has, reset the feature fusion status and update the confidence level and risk calculation direction of the risk assessment indicators. S47. Determine if the maximum number of iterations has been reached. If it has, output the comprehensive risk assessment result of adverse drug reactions; otherwise, continue iterative calculation.
7. The method for collecting and processing adverse drug reaction data based on federated learning according to claim 6, characterized in that, The process of updating the comprehensive risk calculation results through multi-feature fusion, selectively retaining high-reliability indicators, and optimizing the risk assessment calculation includes: S451. Normalize the feature values of each risk assessment indicator in the core feature set to unify risk assessment indicators of different dimensions and scales into the same numerical range, ensuring that each risk assessment indicator is under equivalent calculation conditions when participating in the fusion. S452. Incorporate the normalized multiple risk assessment indicators into the same feature combination, determine all effective risk assessment indicators participating in the comprehensive risk calculation, and form a basic feature set for average value fusion. S453. Calculate the arithmetic mean of the feature values of all valid risk assessment indicators in the basic feature set, and use this mean as the current fusion result to obtain the preliminary comprehensive risk calculation value of adverse drug reactions. S454. Based on the confidence level results, retain high-confidence risk assessment indicators, remove low-confidence risk assessment indicators, recalculate the average value based on the selected effective risk assessment indicators, and update and optimize the final comprehensive risk calculation results.
8. The method for collecting and processing adverse drug reaction data based on federated learning according to claim 7, characterized in that, The process of retaining high-confidence risk assessment indicators and eliminating low-confidence risk assessment indicators based on the confidence level results, recalculating the average value based on the selected effective risk assessment indicators, and updating and optimizing the final comprehensive risk calculation result includes: S4541. Retrieve the confidence level results of each risk assessment indicator, preset the division criteria for high confidence and low confidence indicators, and distinguish between risk assessment indicators to be retained and those to be removed. S4542. Based on the preset classification criteria, retain all high-credibility risk assessment indicators and remove low-credibility risk assessment indicators to obtain a set of effective risk assessment indicators after screening. S4543. Organize the feature values of the selected effective risk assessment indicators, unify the calculation method, and recalculate the arithmetic mean of all the feature values of the effective risk assessment indicators. S4544. Use the recalculated arithmetic mean as the updated comprehensive risk calculation value, replace the original result, and complete the optimization of the comprehensive risk calculation result for adverse drug reactions.
9. The method for collecting and processing adverse drug reaction data based on federated learning according to claim 1, characterized in that, The process of initiating corresponding level risk warnings based on the comprehensive risk assessment results of adverse drug reactions, and generating adverse drug reaction warning reports based on user roles, includes: S51. Based on the comprehensive risk assessment results of adverse drug reactions, determine the current risk level according to the preset risk level classification rules, and determine the early warning activation conditions and early warning type that match the risk level. S52. Based on the responsibilities of different user roles, formulate differentiated early warning information push strategies and determine the early warning content and information display format required for each user role. S53. In accordance with the established push strategy, push the warning information of the corresponding level to the terminal of the corresponding user role to ensure that the warning information accurately reaches the relevant personnel with the authority to handle the situation. S54. Integrate risk assessment results, risk levels, early warning content, and handling authority information to generate a structured adverse drug reaction early warning report in a fixed format and complete the archiving.
10. A federated learning-based adverse drug reaction data collection and processing system, used to implement the federated learning-based adverse drug reaction data collection and processing method according to any one of claims 1-9, characterized in that, The system includes: The data acquisition module is used to acquire adverse drug reaction data from multiple sources and to preprocess the adverse drug reaction data from multiple sources. The data processing module is used to perform semantic standardization on the preprocessed multi-source adverse drug reaction data to obtain a complete multi-source fusion dataset. The feature analysis module is used to extract an initial population of feature terms from the complete multi-source fusion dataset, screen out core seed terms and supplementary seed terms, classify niche categories, expand the dimensions of all feature seed terms, analyze their potential associations and complete their diffusion and reproduction, and finally form a core feature set that affects the occurrence of adverse drug reactions. The risk assessment module is used to construct multi-dimensional risk assessment indicators for adverse drug reactions based on core feature sets, quantify the credibility of drug risk assessment indicators through confidence grading, and obtain the comprehensive risk assessment results of adverse drug reactions through multi-feature fusion calculation. The early warning information generation module is used to initiate corresponding level risk warnings based on the comprehensive risk assessment results of adverse drug reactions, and to push early warning information based on user roles, and generate adverse drug reaction early warning reports.
Citation Information
Patent Citations
Individualized-learning-path optimization method based on lightning searching algorithm
CN108197695A
Explanatable multi-dimensional CI index dynamic scoring and risk assessment method
CN121638942A
Niche Ranking Method
US20220108186A1
Low-altitude air route risk map construction method and system based on hexagonal grid cells
WO2026025602A1