Drug risk analysis method and system based on full life cycle data

By constructing a drug causal network and combining it with a dual closed-loop feedback mechanism, the static and adaptive problems of existing drug risk analysis methods are solved, enabling dynamic management and accurate prediction of drug risks throughout their entire life cycle.

CN122000089APending Publication Date: 2026-05-08TAIJI GROUP SICHUAN MIANYANG PHARMA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TAIJI GROUP SICHUAN MIANYANG PHARMA
Filing Date
2026-01-06
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing drug risk analysis methods are static and have poor adaptability. They cannot self-correct based on real-time data and actual risk events, and it is difficult to quantify and predict the transmission path and cumulative effect of risks in complex systems.

Method used

A causal network for drugs is constructed based on full life-cycle data. Mechanistic models and data-driven causal inference algorithms are used to determine the relationships between nodes, calculate the risk entropy of nodes and simulate risk transmission. The causal network is dynamically corrected through a dual closed-loop feedback mechanism to achieve prediction of risk outbreak paths and model self-adaptation.

Benefits of technology

It enables dynamic correction and adaptive evolution of drug risk analysis, provides objective and quantifiable risk measurement, accurately predicts risk transmission paths and provides targeted decision-making basis, thereby improving the accuracy and adaptability of risk management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122000089A_ABST
    Figure CN122000089A_ABST
Patent Text Reader

Abstract

The invention discloses a drug risk analysis method and system based on full-life-cycle data, and belongs to the technical field of drug quality and safety management.The method comprises the steps that a drug full-life-cycle causal network containing nodes and directed edges is constructed based on historical and real-time full-life-cycle data of drugs; calculating a node risk entropy of each node; performing risk conduction simulation to determine a risk accumulation condition; and predicting and outputting a risk outbreak path. The core of the method is that self-adaptive evolution is realized through a double closed-loop feedback mechanism: a first closed loop corrects the causal network according to a comparison result of a predicted risk outbreak path and an actual risk event, and a second closed loop analyzes and updates an ideal state baseline model according to a global risk state. The technical problem that an existing risk analysis model is static and poor in adaptability is solved, risk analysis capable of self-correction and evolution is achieved, and the long-term accuracy of risk prediction is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pharmaceutical quality and safety management technology, and in particular to a pharmaceutical risk analysis method and system based on full life cycle data. Background Technology

[0002] In the pharmaceutical industry, systematic risk management throughout the entire product lifecycle is a core element in ensuring drug quality, safety, and efficacy, and is also a mandatory requirement of global drug regulatory agencies (such as the U.S. Food and Drug Administration (FDA) and the European Medicines Agency (EMA)). Guidelines such as ICH Q9, "Quality Risk Management," explicitly advocate for a proactive or retrospective approach to identify, assess, control, communicate, and review potential risks throughout the entire product chain, from development and production to distribution and final use.

[0003] Currently, widely used risk assessment tools in the industry, such as Failure Mode and Effects Analysis (FMEA), Fault Tree Analysis (FTA), and Hazard Analysis and Critical Control Points (HACCP), play a crucial role in the initial identification and qualitative assessment of risks. These traditional methods typically rely on the knowledge and experience of expert teams at specific project stages (such as process development or technology transfer) to construct risk models through discussions and scoring. However, the inherent limitations of these existing techniques are becoming increasingly apparent in the face of the increasingly complex and data-driven modern pharmaceutical environment.

[0004] A significant drawback lies in the static nature of existing risk models. Once a risk assessment is completed, its results, such as the Risk Priority Number (RPN), remain unchanged for a long period, lacking dynamic correlation with the massive amounts of real-time data continuously generated during the production process. This leads to a gradual disconnect between risk assessment and actual operating conditions, and the model cannot automatically reflect risk baseline drift caused by changes in raw materials, equipment aging, or process fine-tuning.

[0005] Furthermore, existing methods lack adaptability. Static models cannot learn and evolve when production processes are optimized, the system enters a more stable new normal, or when an unforeseen risk event actually occurs. Any revisions to the model rely on manual, periodic, and delayed review meetings, which are not only resource-intensive but also slow to react and fail to capture early signs of risk evolution. Therefore, these models often become archival compliance documents rather than dynamic management tools that guide daily operations and enable continuous evolution.

[0006] Furthermore, existing technologies fall short in quantifying the transmission and cumulative effects of risks. In a highly interconnected production process, a small fluctuation in an upstream parameter can trigger significant quality deviations downstream through complex causal chains. Traditional methods struggle to accurately simulate the propagation path and aggregation effects of such risks, thus failing to accurately predict the risk transmission path most likely to erupt and cause ultimate harm under the current system state.

[0007] Therefore, how to overcome the static, lagging, and subjective nature of existing risk analysis methods and develop a new technical solution that can deeply integrate full life cycle data, realize dynamic correction and adaptive evolution of risk models, and accurately predict risk outbreak paths has become an urgent technical problem to be solved in this field. Summary of the Invention

[0008] To address the shortcomings of existing technologies, one of the objectives of this invention is to provide a drug risk analysis method based on full life-cycle data, in order to solve the problems of existing risk analysis models being static, having poor adaptability, being unable to self-correct based on real-time data and actual risk events, and having difficulty in quantifying and predicting the transmission path and cumulative effect of risks in complex systems.

[0009] To achieve the above objectives, the present invention provides the following technical solution: A drug risk analysis method based on full life-cycle data includes the following steps: a. Based on historical and real-time full life cycle data of drugs, construct a causal network of the full life cycle of drugs containing multiple nodes and directed edges. The nodes represent entities or states in the life cycle of drugs, the directed edges represent causal relationships between nodes, and a transmission coefficient is configured for the directed edges. b. For each node in the causal network of the entire drug life cycle, the node risk entropy is calculated based on the deviation of the node's current state probability distribution from the preset ideal state baseline model; c. Based on the node risk entropy of each node and the transmission coefficient of the directed edge, perform risk transmission simulation to determine the accumulation of risk in the causal network of the entire drug life cycle; d. Based on the accumulation of the aforementioned risks, predict and output the risk outbreak path; e. Obtain the predicted results of the risk outbreak path and the verification results of the actual risk events, and dynamically correct the causal network of the entire drug life cycle based on the verification results.

[0010] As a preferred technical solution, in step a, The historical and real-time full lifecycle data of the drug includes: Data covering four stages: R&D design, production execution, distribution traceability, and user feedback; as well as timestamp data, temperature chain data, and tracking code data throughout these four stages; Constructing the causal network of the entire drug lifecycle specifically includes:

[0011] A hybrid driving method combining mechanistic models and data-driven causal inference algorithms is used to determine the directed edges between the nodes.

[0012] As a preferred technical solution, the calculation of the node risk entropy in step b specifically includes: Calculate the current state probability distribution of the node. Relative to the ideal state baseline model KL divergence The KL divergence is then compared with a preset severity coefficient that characterizes the inherent risk level of the node. By combining these factors, the node risk entropy of the node can be obtained. The calculation formula can be expressed as: ; in, .

[0013] As a preferred technical solution, step c, which involves simulating risk transmission, specifically includes: Based on the node risk entropy of the upstream node, the transmission coefficient of the directed edge between the upstream node and the downstream node, and a susceptibility coefficient characterizing the downstream node's sensitivity to upstream input risks, the input increment of risk to the downstream node is calculated.

[0014] As a preferred technical solution, the step d of predicting and outputting the risk outbreak path specifically includes: The drug life cycle causal network is transformed into a weighted graph, wherein the travel cost of any directed edge in the weighted graph is set as the reciprocal of the transmission coefficient of the corresponding directed edge in the drug life cycle causal network. A graph search algorithm is used to search and determine the path with the minimum cumulative travel cost from the risk source node to the final hazard node in the weighted graph, which is taken as the risk outbreak path.

[0015] As a preferred technical solution, the dynamic correction of the causal network of the entire drug life cycle in step e specifically includes:

[0016] The predicted risk outbreak path is compared with the actual risk path corresponding to the actual risk event; and based on the comparison result, at least one of the following correction operations is performed: (1) Update the transmission coefficient of the corresponding directed edge in the causal network of the entire life cycle of the drug, wherein if the directed edge in the predicted path exists in the actual risk path, the transmission coefficient of the directed edge is increased; if the directed edge in the predicted path does not exist in the actual risk path, the transmission coefficient of the directed edge is decreased. (2) Add new directed edges to the causal network of the entire life cycle of the drug.

[0017] Furthermore, the drug risk analysis method based on full life-cycle data also includes: Periodically analyze the node risk entropy of all nodes in the causal network of the entire life cycle of the drug to identify the stable region that is in a low risk entropy state for a long time. The ideal state baseline model is updated based on the data from the nodes within the stability domain.

[0018] As a specific implementation, when the ideal state baseline model of a particular node is relaxed for that particular node, the propagation coefficient of all directed edges pointed to by that particular node is increased compensatorily.

[0019] As a preferred technical solution, the present invention achieves adaptive evolution through a dual closed-loop feedback mechanism, wherein the dual closed-loop feedback mechanism includes: A first closed-loop feedback, based on the comparison between the risk outbreak path and the actual risk path, corrects the directed edges and transmission coefficients of the drug's full life cycle causal network; And a second closed-loop feedback, which adjusts the ideal state baseline model based on a global analysis of the node risk entropy of all nodes in the causal network of the entire life cycle of the drug.

[0020] The second objective of this invention is to provide a drug risk analysis system based on full life-cycle data. The drug risk analysis system includes a data acquisition module, a network construction module, a storage module, a risk calculation and transmission module, a path prediction module, and a dynamic correction module that communicate sequentially, as well as a processor that communicates with the data acquisition module, the network construction module, the risk calculation and transmission module, the path prediction module, and the dynamic correction module, respectively.

[0021] This invention provides a drug risk analysis method based on full life-cycle data. It has the following beneficial effects: 1. This invention employs a dual closed-loop feedback mechanism. The first closed-loop feedback dynamically adjusts the directed edges and transmission coefficients of the causal network based on a comparison between the predicted risk outbreak path and the actual risk event. The second closed-loop feedback periodically updates the ideal state baseline model using the risk entropy of the global analysis nodes. This dual adjustment mechanism enables the model to continuously learn and evolve from new data and actual results, effectively responding to dynamically changing system environments and maintaining the long-term validity and accuracy of the analysis results.

[0022] 2. This invention does not employ subjective or fuzzy risk scoring. Instead, it introduces KL divergence from information theory to calculate the deviation of the probability distribution of a node's current state from the ideal baseline model, providing an objective and quantifiable basis for risk measurement. Furthermore, by incorporating a severity coefficient characterizing the inherent risk level of a node, the final calculated node risk entropy more accurately reflects the actual impact of different nodes deviating from their normal state, avoiding errors caused by treating all node deviations equally.

[0023] 3. This invention constructs a causal network of the entire life cycle of a drug and transforms the problem of predicting risk outbreak paths into the problem of searching for the path with the minimum cumulative passage cost on a weighted graph. It can not only predict the cumulative results of risks, but also clearly output a specific risk transmission path. Since this path is generated based on a network of nodes with causal relationships, the prediction results have clear physical meaning and logical interpretability, providing a direct and clear decision-making basis for taking targeted prevention and blocking measures. Attached Figure Description

[0024] Figure 1 This is a functional block diagram of the drug risk analysis system of the present invention; Figure 2 This is a flowchart illustrating the node risk entropy calculation process for a single node in this invention. Figure 3 This is a schematic diagram illustrating the transmission of risks from upstream nodes to downstream nodes in this invention. Figure 4 This is a schematic diagram illustrating the transformation of a causal network into a weighted graph for path search according to the present invention. Figure 5 The flowchart is as follows: This invention provides a dual closed-loop feedback dynamic correction mechanism. Figure 6 This is a simplified causal network diagram of drug production according to the present invention.

[0025] Among them, 100 is the drug risk analysis system; 101 is the data acquisition module; 102 is the network construction module; 103 is the risk calculation and transmission module; 104 is the path prediction module; 105 is the dynamic correction module; 106 is the storage module; and 107 is the processor. Detailed Implementation

[0026] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] See attached document Figure 1 , Figure 1 This is a functional block diagram of a drug risk analysis system according to an embodiment of the present invention. The drug risk analysis system 100 can be used to execute the drug risk analysis method of the present invention, and includes: a data acquisition module 101, a network construction module 102, a risk calculation and transmission module 103, a path prediction module 104, a dynamic correction module 105, a storage module 106, and a processor 107. All modules can communicate and exchange data through an internal bus, and their corresponding functions are executed by the processor 107.

[0028] The present invention provides a drug risk analysis method based on full life cycle data, the specific implementation process of which is as follows: First, step a is performed: based on the historical and real-time full life cycle data of the drug, a causal network of the full life cycle of the drug containing multiple nodes and directed edges is constructed. This step is completed collaboratively by the data acquisition module 101 and the network construction module 102.

[0029] Specifically, the data acquisition module 101 is used to acquire historical and real-time full lifecycle data of the drug from multiple data sources and store it in the storage module 106. The data covers four stages: research and development design, production execution, distribution traceability, and user feedback. For example, research and development design data includes formulation process research data and stability study data; production execution data includes batch production records (BMR), equipment operating parameters, and material inspection reports; distribution traceability data includes warehouse temperature and humidity records and transportation trajectories; and user feedback data includes adverse event reports and patient complaint records.

[0030] In addition, the data also includes timestamp data, temperature chain data, and tracking code data throughout the four stages mentioned above, to ensure the temporal correlation and full traceability of the data.

[0031] The network construction module 102 retrieves data from the storage module 106 to determine the nodes and directed edges in the causal network. In one embodiment, a hybrid driving method combining a mechanistic model and a data-driven causal inference algorithm is employed.

[0032] This hybrid-driven approach first establishes fundamental causal relationships based on mechanistic models. These models are built upon prior knowledge from fields such as pharmacology, chemical engineering, and physics. For example, based on heat transfer principles, a node representing the ambient temperature of the refrigeration equipment is established within the network. Pointing to the node representing the "internal temperature of the drug". The directed edge.

[0033] Subsequently, the method utilizes data-driven causal inference algorithms to discover causal relationships not covered by the mechanistic model from historical data. For example, by employing constraint-based causal discovery algorithms (such as the PC algorithm), historical data in storage module 106 is analyzed to identify conditional independence relationships between variables, thereby inferring the causal structure between nodes. For instance, based on the data analysis results, the algorithm can establish a node representing a supplier's batch number. Pointing to the node representing the purity of the intermediate. The directed edge.

[0034] Finally, the set of directed edges determined by the mechanistic model is merged with the set of directed edges discovered by the causal inference algorithm to form a complete causal network topology structure for the entire life cycle of a drug.

[0035] After determining the nodes and directed edges of the network, the network building module 102 configures a propagation coefficient for each directed edge in the network. For nodes... Pointing to node A directed edge whose conduction coefficient is Characterizes the nodes State changes affect nodes The intensity of the impact of state changes.

[0036] In one specific embodiment, the conduction coefficient Initialization can be performed using linear regression analysis on historical data. Assume the upstream node... and downstream nodes The state values ​​are respectively variables and The relationship between them can be modeled as follows: ; in: It is a downstream node The state value; It is an upstream node The state value; It is the transmission coefficient to be determined, i.e., the regression coefficient; This is the residual term; the entire formula is a univariate linear regression model. Its technical significance lies in the fact that it considers the downstream nodes... State value The changes can be broken down into two parts: one part is caused by its direct upstream node. State value The linearly driven part ( The other part is caused by random disturbances or unmodeled factors. In this invention, this model is applied and fitted using historical data. The purpose of this value is to assign a quantitative initial weight, namely the transmission coefficient, to the directed edges in the causal network based on the statistical regularity of historical data.

[0037] Network building module 102 extracts upstream nodes from storage module 106 and downstream nodes Historical time series data, calculated by performing least squares method The estimated value is used as the initial propagation coefficient of the directed edge. The completed causal network, including its nodes, directed edges and propagation coefficients, is stored in storage module 106 for subsequent steps.

[0038] After constructing the causal network of the entire drug lifecycle, the method proceeds to step b: for each node in the causal network of the entire drug lifecycle, the node risk entropy of the node is calculated. This step is mainly performed by the risk calculation and transmission module 103.

[0039] See attached document Figure 2 , Figure 2 This is a flowchart illustrating the node risk entropy calculation for a single node according to an embodiment of the present invention. The process aims to transform the real-time operational status data of a node into a quantified risk indicator. The specific calculation process for this step is as follows: First, an ideal baseline model is pre-defined for each node in the network. The model It is a probability distribution that characterizes the data distribution characteristics of a node under a risk-free, high-quality operating state. In a specific embodiment, the model is established as follows: Multiple product batches historically confirmed as having qualified quality and excellent production indicators are selected from storage module 106, and specific nodes (e.g., nodes) are extracted from these batches. All historical data of the node. The risk calculation and transmission module 103 performs statistical fitting on these data. If the data is a continuous variable (such as temperature or pressure), it can be fitted with a Gaussian distribution. ),at this time This is the probability density function of the Gaussian distribution. If the data are discrete variables (such as equipment number or operator ID), they can be fitted with a multinomial distribution. This is the probability mass function composed of the frequencies of occurrence of each category. The ideal state baseline model for all nodes is then established. All are stored in storage module 106.

[0040] Secondly, obtain the state probability distribution of each node at the current moment. The risk calculation and transmission module 103 continuously receives real-time data streams from each node through the data acquisition module 101. A sliding time window technique is used to calculate the distribution of the current state. For example, a time window size is set to... (e.g., the past 24 hours) or the number of data points is (For example, the latest 1000 data points), the real-time data within the window is used as a sample set. Subsequently, a baseline model is established. Using the same data fitting method, the sample set is processed to obtain the probability distribution of the current state. .

[0041] In obtaining and Subsequently, the risk calculation and transmission module 103 calculates the node risk entropy of the node. In one specific embodiment, the calculation process includes two sub-steps: The first sub-step is to calculate the probability distribution of the current state. Compared to the ideal baseline model KL divergence KL divergence, also known as relative entropy, is an asymmetric measure used to measure the difference between two probability distributions. Its formula is: ; in: It is the probability distribution of the current state of the node; It is the ideal state baseline model of the node; These are all possible values ​​of the node's state variable. This represents summing over all possible values ​​of the state variable x (for discrete variables) or integrating over all possible values ​​(for continuous variables). This represents the logarithmic function.

[0042] The result of the calculation is a non-negative value. If and only if and When they are exactly the same, the value is 0, indicating that the current state is no different from the ideal state. and The greater the difference, the larger the value, indicating that the current state deviates more from the ideal state.

[0043] The second sub-step is to combine the calculated KL divergence with a preset severity coefficient. Combined. This severity coefficient It is a quantitative indicator used to characterize the severity of the impact of a node's deviation on the quality or safety of the final product, i.e., the inherent risk level of the node. In one embodiment, The values ​​are determined based on Failure Mode and Effects Analysis (FMEA) methods. For example, for nodes corresponding to critical quality attributes (CQA) or critical process parameters (CPP) in the pharmaceutical manufacturing process, a higher severity score in the FMEA analysis will result in a higher assigned severity value. Values ​​(e.g., 0.8-1.0) are used to configure a lower severity score for nodes corresponding to general process monitoring parameters if their severity scores are low. Values ​​(e.g., 0.1-0.3). All nodes. The values ​​are all preset and stored in the configuration table of storage module 106.

[0044] Ultimately, the node's node risk entropy The following formula is used to calculate: ; in, The final calculated node risk entropy is a comprehensive risk indicator value. The severity coefficient represents a preset value, which is a dimensionless weight that characterizes the severity of the impact of node failure on the final result. This represents the KL divergence value calculated by the previous formula, i.e., the state deviation.

[0045] This calculation formula uses multiplication to convert dynamic, data-driven state deviations. With static, prior knowledge-based inherent severity This combination has technical implications, namely, the final node risk entropy. It is a weighted risk metric. It reflects not only how much the node is off-target (by...) The metric also reflects how severe the consequences are after the node has become skewed (by...). (Weighted) This combination means that a small deviation from a high severity node may produce a higher risk entropy than a drastic deviation from a low severity node, thus making the risk assessment closer to the actual quality and safety management needs.

[0046] The calculation yielded This is a dimensionless risk value that combines state deviation and inherent severity. The risk calculation and propagation module 103 periodically performs the above calculation for each node in the network and updates the value accordingly. The value is stored together with the node identifier and used in subsequent risk transmission simulation steps.

[0047] After calculating the node risk entropy of each node in the network, the method then proceeds to step c: based on the node risk entropy of each node and the propagation coefficient of the directed edges, a risk propagation simulation is performed to determine the accumulation of risk in the causal network throughout the entire drug lifecycle. This step is also performed by the risk calculation and propagation module 103.

[0048] See attached document Figure 3 , Figure 3 This is a schematic diagram illustrating the transmission of risk from an upstream node to a downstream node according to an embodiment of the present invention. The process simulates how a risk generated at one node affects other associated nodes through a causal path.

[0049] In one specific embodiment, risk transmission simulation is achieved by calculating the risk input increment. For any downstream node in the network... It receives data from a direct upstream node. Risk input increment It is calculated based on three key parameters: upstream node Node risk entropy Connect to upstream nodes With downstream nodes The transmission coefficient of the directed edge and downstream nodes Susceptibility coefficient .

[0050] Among them, susceptibility coefficient It is a pre-defined representation of downstream nodes. A quantitative indicator of a node's sensitivity or vulnerability to upstream input risks. This coefficient reflects the node's own stability. In one embodiment, The value range is [0,1]. A node with strong robustness or built-in error correction capability has a strong ability to absorb upstream fluctuations and is configured with a lower [value]. A value (e.g., close to 0). Conversely, a node that is inherently sensitive or lacks stable control mechanisms is configured with a higher value. Values ​​(e.g., close to 1). For example, a process parameter node that uses statistical process control (SPC) for real-time monitoring and adjustment, its... The value is low; while a node representing the activity of a biopharmaceutical that is sensitive to temperature changes, its The value is relatively high. The susceptibility coefficient for all nodes. All of them are preset and stored in storage module 106.

[0051] The risk calculation and transmission module 103 uses the following formula to calculate the risk input increment. : ; in, It is an upstream node The node risk entropy represents the magnitude of the risk transmitted from the source; From the upstream node to downstream nodes The propagation coefficient of a directed edge represents the transmission efficiency of risk along the propagation path. It is a downstream node The susceptibility coefficient represents the degree to which downstream nodes absorb or amplify incoming risks.

[0052] The technical implication of this calculation formula is that the impact of a risk generated by an upstream node, when transmitted to a downstream node, is simultaneously modulated by both the efficiency of the transmission path and the stability of the receiving node itself.

[0053] A downstream node Cumulative risks It is its own node risk entropy. The sum of the risk input increments passed to it by all directly upstream nodes. If node The set of direct upstream nodes is The cumulative risk is calculated as follows: ; in, Represents a set Each upstream node Sum the corresponding terms; Representative node From each of its direct upstream nodes The received risk input increment.

[0054] The risk calculation and transmission module 103 calculates the cumulative risk of each node in the network layer by layer, starting from the source node which has no upstream nodes, according to the topological order of the causal network. After the calculation is completed, each node in the network corresponds to a cumulative risk value. This set of values ​​reflects the distribution and accumulation of risk in the causal network of the entire drug life cycle and is stored in storage module 106 for subsequent risk outbreak path prediction.

[0055] After calculating the cumulative risk for each node in the network, the method proceeds to step d: predicting and outputting the risk outbreak path based on the cumulative risk. This step is executed by the path prediction module 104.

[0056] See attached document Figure 4 , Figure 4 This is a schematic diagram illustrating the transformation of a causal network into a weighted graph for pathfinding according to an embodiment of the present invention. The goal of this process is to identify the most likely path for risk to propagate from a high-risk source to the eventual harmful incident.

[0057] First, the path prediction module 104 transforms the drug lifecycle causal network constructed in the storage module 106 into a weighted directed graph for path search. During this transformation, all nodes and directed edges in the causal network are fully preserved. The core of the transformation lies in defining each directed edge in the graph... (From upstream node) Pointing to downstream nodes Assign a passage cost The passage cost is set as the propagation coefficient of the corresponding directed edge in the causal network. The reciprocal of . The formula is as follows: ; in: It is a weighted graph starting from the upstream node to downstream nodes The cost of traveling along a directed edge; The reason is that in the resultant network, from the node To the node The conduction coefficient of directed edges is a formula that defines the mapping rule for converting "conduction efficiency" into "travel cost". Its technical implication is that the conduction coefficient... The larger the value, the easier it is for the risk to travel along that path. In graph search algorithms, this should be reflected as a "shorter" or "better" path. By taking its reciprocal, a high transmission coefficient is transformed into a low passage cost, thus allowing algorithms such as Dijkstra's, which find minimum-cost paths, to be directly used to find the path where the risk is most easily transmitted, i.e., the risk outbreak path.

[0058] The technical implication of this calculation formula lies in the conductivity coefficient. Larger edges indicate higher efficiency in risk transmission, and therefore should be prioritized in path search. By using its reciprocal as the travel cost, a path with high transmission efficiency has a lower travel cost in the weighted graph, thus meeting the goal of graph search algorithms to find the minimum cost path.

[0059] Secondly, the path prediction module 104 determines the risk source nodes and the final hazard nodes in the transformed weighted graph. In a specific embodiment, the risk source nodes... Defined as having the highest cumulative risk in the current network. The path prediction module 104 queries all nodes in the storage module 106. The values ​​are compared to determine the node. Ultimately, the node becomes the source of the harm. These are a type of node predefined during network construction, representing undesirable end events in the drug lifecycle, such as drug recalls, serious adverse reaction reports, or batch obsolescence.

[0060] Finally, the path prediction module 104 uses a graph search algorithm to search and determine the path from the risk source node in the weighted graph. To one or more final hazard nodes The minimum cumulative toll cost path. In one embodiment, Dijkstra's algorithm is used to perform this search. This algorithm starts from... Begin by systematically exploring the nodes in the network, calculating and updating from... Find the shortest path to each node, until a final hazard node is found. The path with the lowest cumulative toll cost.

[0061] The path with the minimum cumulative travel cost found is identified as the predicted risk outbreak path. This path is an ordered sequence of nodes, for example... It clearly indicates the series of intermediate steps through which risk propagates from the current highest risk point to the final harmful event. The path prediction module 104 outputs this path sequence and stores it in the storage module 106 for use in subsequent dynamic correction steps and as a basis for risk intervention decisions.

[0062] After predicting and outputting the risk outbreak path, the method executes step e: obtaining the predicted results of the risk outbreak path and the verification results of the actual risk events, and dynamically correcting the drug's entire life cycle causal network based on the verification results. This step is executed by the dynamic correction module 105, whose core is to enable the entire risk analysis model to have the ability to self-correct and evolve through a dual closed-loop feedback mechanism.

[0063] See attached document Figure 5 , Figure 5 This is a flowchart of a dual closed-loop feedback dynamic correction mechanism according to an embodiment of the present invention, the mechanism including a first closed-loop feedback and a second closed-loop feedback.

[0064] The first closed-loop feedback, based on the verification results of a single risk event, corrects the local structure and parameters of the causal network. When an actual risk event occurs (for example, receiving a report through a drug adverse reaction monitoring system), relevant personnel will trace the event to determine the actual risk path that led to the event. This actual risk path is then input into the dynamic correction module 105.

[0065] The dynamic correction module 105 first compares the previously predicted risk outbreak paths related to the event in the storage module 106 with the actual input risk paths. The comparison is performed edge-by-edge. Based on the comparison results, the dynamic correction module 105 performs at least one of the following correction operations: (1) Update the propagation coefficients of directed edges. For directed edges that coexist in both the predicted path and the actual path. Increase its conductivity This reinforces the verified causal relationship. For directed edges that exist in the predicted path but not in the actual path... Reduce its conduction coefficient This weakens the falsified causal relationship. In a specific embodiment, the update formula is as follows: ; ; in: and These are the updated conduction coefficients; and These are the conduction coefficients before and after the update; and It is a preset positive number, used as the learning rate for updating the step size.

[0066] (2) Add new directed edges. If there exists a node in the actual risk path... To the node The causal relationship is not established, but there is no corresponding directed edge in the current causal network model. The dynamic correction module 105 will add a directed edge from the network. point to The new directed edge. The initial propagation coefficient of this new directed edge can be initially set based on the data of this risk event.

[0067] The second closed-loop feedback adjusts the risk assessment benchmark based on long-term analysis of the system's global state. The dynamic correction module 105 periodically (e.g., quarterly) adjusts the node risk entropy of all nodes in the network. Analysis is performed to identify the stable region. A node is identified as being in the stable region when its node risk entropy exceeds a certain threshold over N consecutive analysis periods. It is consistently below a preset low-risk threshold. .

[0068] After identifying the stable region, the dynamic correction module 105 extracts the original data corresponding to all nodes within the stable region during this time period, and uses this data to establish the ideal state baseline model for these nodes. Update it. For example, if the original baseline model It is a Gaussian distribution Then, the mean and variance are recalculated using the new data from the stability region to obtain new parameters. and This updates the baseline model. This operation allows the risk assessment benchmark to adapt to the new, better stable operating state achieved by the system after process optimization or environmental changes.

[0069] Furthermore, the present invention also includes a compensatory adjustment mechanism. When targeting a specific node... Relax its ideal state baseline model When (for example, after a process change, the range of data fluctuations that allow for normal operation increases), to ensure that the overall risk sensitivity of the system does not decrease, the dynamic correction module 105 will compensatorily increase the value generated by the nodes. All directed edges indicated (i.e., from) arrive All downstream nodes Transmission coefficient In one embodiment, this compensatory adjustment is compared with the baseline model. variance The changes are related, for example: ; in, Representative node A certain outgoing edge (pointing to) downstream nodes Updated conduction coefficient; This represents the propagation coefficient of the edge before the update; Representative node Updated ideal baseline model The variance; Representative node Ideal baseline model before update variance This mechanism ensures that even if the normal fluctuation range of a node is widened, the impact of any abnormal fluctuation on downstream nodes will be amplified proportionally, thereby maintaining the overall balance of risk transmission. All dynamically corrected network parameters, structures, and baseline models will be updated and stored in storage module 106.

[0070] See attached document Figure 6 , Figure 6This is a simplified causal network diagram of drug production according to an embodiment of the present invention. To facilitate understanding of the technical solution of the present invention, a specific application embodiment is given below. This application embodiment describes the application of the method of the present invention in a simplified monoclonal antibody injection production process, but the scope of protection of the present invention is not limited to this embodiment.

[0071] 1. Construction of causal networks In a specific application scenario, the drug risk analysis system 100 first constructs a causal network for a certain monoclonal antibody injection. The network construction module 102 identifies the following six key nodes: Raw material supplier batches (discrete variable). pH value of bioreactor (continuous variable) Purification chromatography column pressure (continuous variable). Sterilization filtration integrity test (discrete variable: pass / fail). Polymer content in finished product (continuous variable, unit: %) Serious adverse event reporting (final hazard point); The directed edges and initial conduction coefficients between nodes were determined using a hybrid driving method. For example, based on biochemical mechanisms, the direction of conduction was determined from... (pH value) indicates The directed edges (of aggregate content) are defined, and their initial conduction coefficients are set through regression analysis of historical data. Through data-driven causal inference algorithms, specific supplier batches were identified. ) and chromatography column pressure ( The anomalies are related, so an additional entry is added from... point to The directed edges are defined. And so on, a complete network containing all nodes, directed edges, and propagation coefficients is constructed and stored.

[0072] 2. Calculation of node risk entropy The risk calculation and transmission module 103 calculates the risk entropy for each node. (Based on the node...) (Taking pH value of a bioreactor as an example): Ideal baseline model Based on data from 100 historical qualified batches, the pH value distribution was fitted to a value with a mean of 7.10 and a variance of [missing value]. Gaussian distribution ; Current state probability distribution pH data from the past 24 hours were collected and fitted to obtain a new Gaussian distribution. ; KL divergence calculation: and Substituting into the KL divergence formula, we can calculate... ; Node risk entropy calculation: pH value is a critical process parameter with a preset severity coefficient. The final calculation yields the nodes. Node risk entropy ; Perform the same calculation on other nodes in the network, and assume that the result is... .

[0073] 3. Risk transmission and accumulation The risk calculation and transmission module 103 then calculates the cumulative transmission of risk. This is done using nodes... Taking (polymer content) as an example, its direct upstream node is and ; Preset Nodes Susceptibility coefficient ; Calculation from Risk input increment: ; Calculation from Risk input increment: (assumption) =0.2); compute nodes Cumulative risks: .

[0074] 4. Prediction of risk outbreak paths Path prediction module 104 performs path prediction: First, the causal network is transformed into a weighted graph, where the travel cost of an edge is the reciprocal of the propagation coefficient. For example, toll costs ; Secondly, identify the risk source nodes and the final hazard nodes. This is done by comparing the node risk entropy of all nodes. The source node with the highest current risk The final hazard node is a preset one. ; Finally, Dijkstra's algorithm is used to search for the source in the weighted graph. arrive The path with the minimum cumulative toll cost. Assume the search result is the path... This path is output by the system as the predicted risk outbreak path.

[0075] 5. Dynamic correction of the model During subsequent operation, a serious adverse event occurred. An investigation determined the root cause to be a specific batch from a particular supplier (…). The raw materials contained an undetected impurity, which directly led to a decrease in the polymer content in the finished product. The increase in ) ultimately led to adverse events ( The actual risk path is as follows: The dynamic correction module 105 receives this verification result and performs a dual closed-loop feedback: First closed-loop feedback: Predicted path Compared with the actual path Perform a comparison; Due to the edge If it exists in the predicted path but not in the actual path, reduce its propagation coefficient: (Assume β = 0.1) Due to the edge Since it exists in both paths, increase its conduction coefficient: (assumption) =0.7, =0.1); Because the actual path contains causal relationships not included in the model. Add a new directed edge to the network. And set an initial transmission coefficient for it based on the data of this event, for example =0.3; Second closed-loop feedback: After one quarter, periodic analysis of the system revealed that due to the replacement of the chromatography column packing material, the nodes... The operating state of (chromatographic column pressure) became very stable, and its node risk entropy... For multiple consecutive periods, the value is below the preset threshold. ; The dynamic correction module 105 will node It was identified as entering the stable region, and its ideal-state baseline model was constructed using the data from that quarter. From the original Updated to be more accurate .

[0076] Through the above steps, this application example fully demonstrates how to build models based on full lifecycle data, quantify and predict risks, and how to achieve adaptive correction and evolution of models by comparing them with actual events and analyzing system stability.

[0077] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A drug risk analysis method based on full life-cycle data, characterized in that, Includes the following steps: a. Based on historical and real-time full life cycle data of drugs, construct a causal network of the full life cycle of drugs containing multiple nodes and directed edges. The nodes represent entities or states in the life cycle of drugs, the directed edges represent causal relationships between nodes, and a transmission coefficient is configured for the directed edges. b. For each node in the causal network of the entire drug life cycle, the node risk entropy is calculated based on the deviation of the node's current state probability distribution from the preset ideal state baseline model; c. Based on the node risk entropy of each node and the transmission coefficient of the directed edge, perform risk transmission simulation to determine the accumulation of risk in the causal network of the entire drug life cycle; d. Based on the accumulation of the aforementioned risks, predict and output the risk outbreak path; e. Obtain the predicted results of the risk outbreak path and the verification results of the actual risk events, and dynamically correct the causal network of the entire drug life cycle based on the verification results.

2. The drug risk analysis method based on full life cycle data according to claim 1, characterized in that, In step a) The historical and real-time full lifecycle data of the drug includes: Data covering four stages: R&D design, production execution, distribution traceability, and user feedback; And timestamp data, temperature chain data, and tracking code data throughout the four stages; Constructing the causal network of the entire drug lifecycle specifically includes: A hybrid driving method combining mechanistic models and data-driven causal inference algorithms is used to determine the directed edges between the nodes.

3. The drug risk analysis method based on full life cycle data according to claim 1, characterized in that, In step b, The calculation of the node risk entropy specifically includes: Calculate the KL divergence of the current state probability distribution relative to the ideal state baseline model, and combine the KL divergence with a preset severity coefficient characterizing the inherent risk level of the node to obtain the node risk entropy.

4. The drug risk analysis method based on full life cycle data according to claim 1, characterized in that, In step c, Conducting risk transmission simulations specifically includes: Based on the node risk entropy of the upstream node, the transmission coefficient of the directed edge between the upstream node and the downstream node, and a susceptibility coefficient characterizing the downstream node's sensitivity to upstream input risks, the input increment of risk to the upstream node is calculated.

5. The drug risk analysis method based on full life cycle data according to claim 1, characterized in that, In step d, Predicting and outputting the risk outbreak path specifically includes: The causal network of the entire drug life cycle is transformed into a weighted graph with the reciprocal of the transmission coefficient as the passage cost; A graph search algorithm is used to search and determine the path with the minimum cumulative travel cost from the risk source node to the final hazard node in the weighted graph, which is taken as the risk outbreak path.

6. The drug risk analysis method based on full life cycle data according to claim 1, characterized in that, In step e Dynamically correcting the causal network of the entire drug lifecycle specifically includes: The predicted risk outbreak path is compared with the actual risk path corresponding to the actual risk event. Based on the comparison results, perform at least one of the following correction operations: Update the transmission coefficient of the corresponding directed edge in the causal network of the entire life cycle of the drug. If the directed edge in the predicted path exists in the actual risk path, increase the transmission coefficient of the directed edge; if the directed edge in the predicted path does not exist in the actual risk path, decrease the transmission coefficient of the directed edge. Add new directed edges to the causal network of the entire drug life cycle.

7. The drug risk analysis method based on full life cycle data according to claim 1, characterized in that, The drug risk analysis method based on full life cycle data also includes: Periodically analyze the node risk entropy of all nodes in the causal network of the entire life cycle of the drug to identify the stable region that is in a low risk entropy state for a long time. The ideal state baseline model is updated based on the data from the nodes within the stability domain.

8. The drug risk analysis method based on full life cycle data according to claim 7, characterized in that, The drug risk analysis method based on full life cycle data also includes: When the ideal state baseline model of a particular node is relaxed for that particular node, the transmission coefficient of all directed edges pointed to by that particular node is increased compensatorily.

9. The drug risk analysis method based on full life cycle data according to claim 1, characterized in that, The drug risk analysis method based on full life-cycle data achieves adaptive evolution through a dual closed-loop feedback mechanism, which includes: A first closed-loop feedback that corrects the directed edges and transmission coefficients of the drug's full life-cycle causal network based on the comparison results between the risk outbreak path and the actual risk path; And a second closed-loop feedback that adjusts the ideal baseline model based on a global analysis of the node risk entropy of all nodes in the causal network of the entire life cycle of the drug.

10. A drug risk analysis system based on full life-cycle data, characterized in that, The drug risk analysis system (100) includes a data acquisition module (101), a network construction module (102), a storage module (106), a risk calculation and transmission module (103), a path prediction module (104), and a dynamic correction module (105) that communicate sequentially, as well as a processor (107) that communicates with the data acquisition module (101), the network construction module (102), the risk calculation and transmission module (103), the path prediction module (104), and the dynamic correction module (105) respectively.