Drug production quality identification method and system

By combining a multimodal edge intelligent sensing network and a digital twin simulation model with blockchain technology, the problems of missed detection, false alarms, and response delays in drug quality testing have been solved, enabling efficient and accurate drug quality monitoring and root cause analysis.

CN121903409APending Publication Date: 2026-04-21JINGHUA PHARMA GRP NANTONG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JINGHUA PHARMA GRP NANTONG
Filing Date
2025-11-11
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Traditional drug quality testing methods suffer from high false negative rates, long response delays, and high false alarm rates, making it difficult to ensure drug quality and production efficiency.

Method used

By employing a multimodal edge intelligent sensing network and a dynamic threshold fusion decision-making mechanism, combined with a digital twin simulation model and blockchain technology, we can achieve real-time non-destructive testing and data-driven root cause analysis throughout the entire process.

Benefits of technology

It enables real-time non-destructive testing in the drug production process, reducing the false alarm rate to 0.1%, the false alarm rate to 5%, the detection response time to 10 seconds, the root cause localization time to 2 hours, and the accuracy rate to 90%, thereby reducing production costs and raw material waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121903409A_ABST
    Figure CN121903409A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of drug production, in particular to a drug production quality identification method and system, and the method comprises the steps: data collection and integration, real-time monitoring and anomaly detection, intelligent quality prediction and root cause analysis, automatic processing and closed-loop control, and feedback optimization closed loop. Compared with the traditional medicine quality detection method which mainly depends on manual sampling inspection or single sensor monitoring, has the defects of high omission ratio, long response delay, frequent misinformation and the like, and causes the risk that unqualified products flow into the downstream and even are recalled in the market, the method adopts a multi-mode edge intelligent sensing network and dynamic threshold fusion decision-making mechanism; according to the technical scheme, full-process real-time nondestructive testing is achieved, an actual production line verifies that the omission ratio is reduced to 0.1%, the false alarm rate is reduced to 5%, the detection response time is shortened to 10 seconds from 4 hours, about 500,000 dollars of quality recall loss caused by omission can be avoided every year, and meanwhile, the reinspection labor cost is reduced by 30%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pharmaceutical manufacturing technology, and in particular to a method and system for identifying pharmaceutical manufacturing quality. Background Technology

[0002] In the pharmaceutical manufacturing sector, drug quality testing is a crucial step in ensuring drug safety and efficacy. It is of paramount importance for safeguarding patient medication safety, maintaining order in the pharmaceutical market, and enhancing the competitiveness of pharmaceutical companies.

[0003] Traditional drug quality testing methods primarily rely on two approaches: manual sampling and single-sensor monitoring. Manual sampling involves selecting a certain percentage of product samples, which are then inspected by quality control personnel according to relevant standards and specifications. However, this method has several limitations. Firstly, the sample size for manual sampling is relatively limited, making it difficult to cover the entire production batch, resulting in a high false negative rate, typically around 5%. This means that some substandard drugs may enter subsequent production stages or even the market, posing a potential risk to patient safety. Secondly, manual sampling requires significant time and manpower, the testing process is cumbersome, and the results are greatly influenced by the experience, skills, and subjective judgment of the quality control personnel, making them prone to human error.

[0004] Single-sensor monitoring utilizes specific types of sensors to monitor certain key parameters in the drug manufacturing process in real time, such as temperature, humidity, and pressure. While this method can improve the automation and real-time performance of detection to some extent, its limited monitoring range means it can only acquire partial information and cannot comprehensively reflect the quality status of the drug manufacturing process. Furthermore, single-sensor monitoring systems typically use fixed threshold judgment standards, making it difficult to adapt to the various complex and changing conditions in the production process. This results in long response delays, often requiring several hours to complete a single detection and feedback, hindering timely problem detection and intervention. In addition, due to the inherent accuracy and stability issues of the sensors themselves, as well as interference from the monitoring environment, single-sensor monitoring is prone to false alarms, with a false alarm rate of approximately 15%. This not only increases the company's production costs and operational burden but may also disrupt normal production operations. Summary of the Invention

[0005] To overcome the problems mentioned in the background art, the present invention proposes a method and system for identifying drug production quality.

[0006] The technical solution of the present invention is: a method for identifying the quality of drug production, comprising the following steps:

[0007] S11: Data acquisition and integration. Input drug production data, testing data and external data, and clean, format and store the data to establish a full-domain data asset catalog. Among them, drug production data includes equipment parameters for drug production, process steps for drug production and real-time sensor data during drug production.

[0008] S12: Real-time monitoring and anomaly detection. Real-time data of finished drug products is collected, and edge computing technology is used to process the collected data in real time. Then, anomaly detection algorithms are used for analysis.

[0009] S13: Intelligent quality prediction and root cause analysis. Establish a digital twin simulation model and input the drug's production data, inspection data, and external data into the established digital twin simulation model to predict drug quality and perform root cause analysis on abnormal drug data.

[0010] S14: Automated processing and closed-loop control, using blockchain technology to record drug production batches and automatically process drugs with quality defects;

[0011] S15: Feedback optimization closed loop, feeding back deviation processing results and customer complaint data to the data lake to optimize the anomaly detection algorithm and digital twin simulation model.

[0012] Preferably, when collecting real-time data on the drug, processing the collected data in real-time using edge computing technology, and then analyzing it using anomaly detection algorithms, the specific steps include:

[0013] S21: Data acquisition, deployment of a multimodal sensor network, and acquisition of real-time data of finished drug products using the multimodal sensor network, including a linear CCD camera, a micro balance, a laser leak detector, and an RFID reader;

[0014] S22: Edge preprocessing, which utilizes edge computing technology to preprocess, extract features, and fuse features from the data collected by the multimodal sensor network;

[0015] S23: Multi-level anomaly detection, which uses multiple anomaly detection algorithms to detect anomalies in the data of finished drug products.

[0016] Preferably, when using edge computing technology to preprocess, extract features, and fuse features from data collected by a multimodal sensor network, the resulting feature data includes:

[0017] A11: Time-domain features, including approximate entropy, where the calculation principle formula is: ApEn(m,r)=□ m (r)-□ m+1(r), where m is the embedding dimension and m = 2, r is the similarity threshold and r = 0.2σ, □ m (r) represents the average log probability of pattern repetition in dimension m, and σ is the standard deviation of the time series.

[0018] A12: Frequency domain features, including wavelet packet energy entropy, where the calculation principle formula is: Among them, W j (n) represents the coefficients of the nth wavelet packet in the j-th layer, E j Let E be the wavelet packet energy of the j-th layer. total H is the sum of the energies of all layers, where H is the energy entropy value and J is the total number of layers in the wavelet packet decomposition.

[0019] A13: Spatial characteristics, including the ratio of surface defect area to drug surface, where the calculation principle formula is: Among them, R defect For the surface area defect of the drug, A defect Let A be the pixel area of ​​the defect region. total This represents the total projected area of ​​the tablet;

[0020] As a preferred method, when using multiple anomaly detection algorithms to detect anomalies in the data of finished drug products, the specific methods include:

[0021] S31: Multi-source data synchronization, time alignment of multi-source data through NTP protocol, and spatial registration through mechanical coordinate and visual coordinate transformation matrix;

[0022] S32: Multi-model parallel detection, which uses three anomaly detection algorithms in parallel: an isolated forest-based anomaly detection algorithm, an autoencoder reconstruction detection algorithm, and a dynamic statistical process control method;

[0023] S33: Confidence fusion decision, which uses a weighted voting mechanism to weight and fuse the data results of the three anomaly detection algorithms;

[0024] S34: Dynamic threshold adjustment, dynamically adjusting the threshold for abnormal alarms based on the slow changes in the test results of the finished drug product.

[0025] As a preferred approach, when using three anomaly detection algorithms in parallel—isolated forest-based anomaly detection algorithm, autoencoder reconstruction detection algorithm, and dynamic statistical process control method—the underlying principle is as follows:

[0026] A21: Anomaly detection algorithm based on isolated forest:

[0027]

[0028] Where s(x,n) represents the anomaly score of sample x, E(h(x)) represents the average path length of sample x across all trees, c(n) represents the path length correction term used to standardize the results under different sample sizes, and n is the subsample size used during training. H(k) is the harmonic number, and H(k)≈ln(k)+0.5772;

[0029] A22: Autoencoder Reconstruction Detection Algorithm

[0030] Alarm condition: □>1.5Q 0.95 ;

[0031] Where x is the original input feature vector, For the autoencoder to reconstruct the output, Q 0.95 The 95th percentile of the reconstruction error of normal samples in the training set;

[0032] A23: Dynamic Statistical Process Control Methods

[0033]

[0034] Where USL is the upper limit of specifications, LSL is the lower limit of specifications, μ is the process mean, and σ is the process standard deviation.

[0035] As a preferred option, when dynamically adjusting the threshold for abnormal alarms based on the slow changes in the test results of the finished drug product, the underlying formula is as follows:

[0036] μ t =λμ t-1 +(1-λ)x t ;

[0037]

[0038] λ = 0.95;

[0039] Where, μ t The process mean estimate at time t. Let be the process variance estimate at time t, λ be the forgetting factor, and x be the value of the process variance estimate at time t. t Let μ be the observed value at time t. t-1 This is the process mean of the previous time step. This represents the process variance at the previous time step.

[0040] Preferably, when establishing a digital twin simulation model and inputting drug production data, testing data, and external data into the established digital twin simulation model to predict drug quality and perform root cause analysis on abnormal drug data, the specific steps include:

[0041] S41: Data input and model initialization. Input the drug's production data, testing data, external data, and data collected by the multimodal sensor network to build a digital twin model and bind the drug's production data to the digital twin model.

[0042] S42: Quality prediction simulation, inputting the data collected by the multimodal sensor network into the digital twin model, performing multi-drug production simulation, fine-tuning the parameters, and solving the optimal combination of process parameters in reverse;

[0043] S43: Root cause analysis of abnormal data. Based on the data results of the anomaly detection algorithm, the digital twin model is used to perform root cause analysis of abnormal data.

[0044] S44: Dynamic model calibration, which uses dynamic calibration mechanism and data feedback mechanism to optimize and calibrate the model.

[0045] Preferably, when performing root cause analysis of abnormal data using a digital twin model based on the data results of the anomaly detection algorithm, the specific steps include:

[0046] S51: Abnormal data filtering, extracting batch data for which the abnormal detection algorithm triggers an alert;

[0047] S52: Causal graph construction, defining potential causal factors, and constructing the causal graph using the CausalNLP framework;

[0048] S53: Causal inference first verifies the root cause through counterfactual analysis, and then uses Bayesian network-based calculation of the contribution of each factor to calculate the confidence level.

[0049] As a preferred method, when calculating the contribution of each factor using a Bayesian network, the underlying formula is as follows:

[0050]

[0051] Among them, C i Represents variable X i The contribution of E(Y) represents the expected value of the result Y. Represents the forced variable X i Value Represents the forced variable X i Value The absolute value of the average change in Y after that.

[0052] A drug manufacturing quality identification system, comprising:

[0053] The data management module is responsible for integrating, cleaning, and storing multi-source heterogeneous data to build a standardized data lake;

[0054] Edge computing nodes are used to deploy lightweight AI models on devices, process sensor data in real time, and perform preliminary anomaly detection.

[0055] The data analysis module is used to analyze historical and real-time data through machine learning and statistical methods to identify quality trends, pinpoint the root causes of deviations, and generate decision recommendations.

[0056] The model simulation module is used to build virtual production lines based on digital twin technology, simulate the impact of process parameters on quality attributes, predict risks, and optimize production strategies.

[0057] The beneficial effects of this invention are:

[0058] 1. Compared to traditional drug quality testing, which mainly relies on manual sampling or single-sensor monitoring, resulting in high false negative rates (approximately 5%), long response delays (several hours), and frequent false alarms (approximately 15%), leading to the risk of substandard products flowing downstream or even market recalls, this invention employs a multimodal edge intelligent sensing network and a dynamic threshold fusion decision mechanism to achieve real-time non-destructive testing throughout the entire process. Verified in actual production lines, this solution reduces the false negative rate to 0.1%, the false alarm rate to 5%, and the detection response time from 4 hours to 10 seconds. It can avoid approximately $500,000 in quality recall losses annually due to false negatives, while also reducing re-inspection labor costs by 30%.

[0059] 2. Traditional quality deviation analysis relies on manual experience and trial-and-error experiments, which is time-consuming (3-5 days) and has low accuracy (approximately 60%), severely delaying production recovery and process optimization. This solution constructs a causal-driven digital twin system, combining counterfactual intervention analysis and Bayesian contribution calculation to form a data-driven root cause reasoning chain. This reduces root cause location time to 2 hours, increases accuracy to 90%, and, based on reverse parameter optimization, increases the dissolution pass rate from 88% to 98%, reducing rework costs by $300,000 annually while reducing raw material waste by 15%.

[0060] 3. Compared to traditional decentralized data management, which is susceptible to tampering, suffers from low efficiency in multi-site collaboration, lengthy audit preparation time (2 weeks), and a pass rate of less than 70%, making it difficult to meet the stringent requirements of international drug regulatory agencies (such as the FDA and EMA). This solution utilizes a blockchain-based evidence storage chain (Hyperledger Fabric records the hashes of batch-wide lifecycle data) and federated learning for cross-site collaboration (FATE framework encrypted shared model parameters) to achieve data tamper-proof traceability and secure knowledge transfer. After implementation, audit preparation time was reduced from 2 weeks to 3 days, the audit pass rate increased to 95%, and the process transfer cycle for new production bases was reduced from 3 months to 2 weeks. One anti-cancer drug entered the European market 6 months ahead of schedule thanks to this technology, resulting in an annual revenue increase of over $20 million. Attached Figure Description

[0061] Figure 1 The diagram shows the workflow of the drug production quality identification method of the present invention.

[0062] Figure 2 The diagram shown is a schematic representation of the structure of the drug production quality identification system of the present invention. Detailed Implementation

[0063] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0064] Please see Figure 1-2 The present invention provides an embodiment of a drug production quality identification method, comprising the following steps:

[0065] S11: Data acquisition and integration. Input drug production data, testing data and external data, and clean, format and store the data to establish a full-domain data asset catalog. Among them, drug production data includes equipment parameters for drug production, process steps for drug production and real-time sensor data during drug production.

[0066] S12: Real-time monitoring and anomaly detection. Real-time data of finished drug products is collected, and edge computing technology is used to process the collected data in real time. Then, anomaly detection algorithms are used for analysis.

[0067] S13: Intelligent quality prediction and root cause analysis. Establish a digital twin simulation model and input the drug's production data, inspection data, and external data into the established digital twin simulation model to predict drug quality and perform root cause analysis on abnormal drug data.

[0068] S14: Automated processing and closed-loop control, using blockchain technology to record drug production batches and automatically process drugs with quality defects;

[0069] S15: Feedback optimization closed loop, feeding back deviation processing results and customer complaint data to the data lake to optimize the anomaly detection algorithm and digital twin simulation model.

[0070] Preferably, when collecting real-time data on the drug, processing the collected data in real-time using edge computing technology, and then analyzing it using anomaly detection algorithms, the specific steps include:

[0071] S21: Data acquisition, deployment of a multimodal sensor network, and acquisition of real-time data of finished drug products using the multimodal sensor network, including a linear CCD camera, a micro balance, a laser leak detector, and an RFID reader;

[0072] S22: Edge preprocessing, which utilizes edge computing technology to preprocess, extract features, and fuse features from the data collected by the multimodal sensor network;

[0073] S23: Multi-level anomaly detection, which uses multiple anomaly detection algorithms to detect anomalies in the data of finished drug products.

[0074] Preferably, when using edge computing technology to preprocess, extract features, and fuse features from data collected by a multimodal sensor network, the resulting feature data includes:

[0075] A11: Time-domain features, including approximate entropy, where the calculation principle formula is: ApEn(m,r)=□ m (r)-□ m+1 (r), where m is the embedding dimension and m = 2, r is the similarity threshold and r = 0.2σ, □ m (r) represents the average log probability of pattern repetition in dimension m, and σ is the standard deviation of the time series.

[0076] A12: Frequency domain features, including wavelet packet energy entropy, where the calculation principle formula is: Among them, W j (n) represents the coefficients of the nth wavelet packet in the j-th layer, E j Let E be the wavelet packet energy of the j-th layer. total H is the sum of the energies of all layers, where H is the energy entropy value and J is the total number of layers in the wavelet packet decomposition.

[0077] A13: Spatial characteristics, including the ratio of surface defect area to drug surface, where the calculation principle formula is: Among them, R defect For the surface area defect of the drug, A defect Let A be the pixel area of ​​the defect region. total This represents the total projected area of ​​the tablet;

[0078] As a preferred method, when using multiple anomaly detection algorithms to detect anomalies in the data of finished drug products, the specific methods include:

[0079] S31: Multi-source data synchronization, time alignment of multi-source data through NTP protocol, and spatial registration through mechanical coordinate and visual coordinate transformation matrix;

[0080] S32: Multi-model parallel detection, which uses three anomaly detection algorithms in parallel: an isolated forest-based anomaly detection algorithm, an autoencoder reconstruction detection algorithm, and a dynamic statistical process control method;

[0081] S33: Confidence fusion decision, which uses a weighted voting mechanism to weight and fuse the data results of the three anomaly detection algorithms;

[0082] S34: Dynamic threshold adjustment, dynamically adjusting the threshold for abnormal alarms based on the slow changes in the test results of the finished drug product.

[0083] As a preferred approach, when using three anomaly detection algorithms in parallel—isolated forest-based anomaly detection algorithm, autoencoder reconstruction detection algorithm, and dynamic statistical process control method—the underlying principle is as follows:

[0084] A21: Anomaly detection algorithm based on isolated forest:

[0085]

[0086] Where s(x,n) represents the anomaly score of sample x, E(h(x)) represents the average path length of sample x across all trees, c(n) represents the path length correction term used to standardize the results under different sample sizes, and n is the subsample size used during training. H(k) is the harmonic number, and H(k)≈ln(k)+0.5772;

[0087] A22: Autoencoder Reconstruction Detection Algorithm

[0088] Alarm condition: □>1.5Q 0.95 ;

[0089] Where x is the original input feature vector, For the autoencoder to reconstruct the output, Q 0.95 The 95th percentile of the reconstruction error of normal samples in the training set;

[0090] A23: Dynamic Statistical Process Control Methods

[0091]

[0092] Where USL is the upper limit of specifications, LSL is the lower limit of specifications, μ is the process mean, and σ is the process standard deviation.

[0093] As a preferred option, when dynamically adjusting the threshold for abnormal alarms based on the slow changes in the test results of the finished drug product, the underlying formula is as follows:

[0094] μ t =λμ t-1 +(1-λ)x t ;

[0095]

[0096] λ = 0.95;

[0097] Where, μ t The process mean estimate at time t. Let be the process variance estimate at time t, λ be the forgetting factor, and x be the value of the process variance estimate at time t. t Let μ be the observed value at time t. t-1 This is the process mean of the previous time step. This represents the process variance at the previous time step.

[0098] Preferably, when establishing a digital twin simulation model and inputting drug production data, testing data, and external data into the established digital twin simulation model to predict drug quality and perform root cause analysis on abnormal drug data, the specific steps include:

[0099] S41: Data input and model initialization. Input the drug's production data, testing data, external data, and data collected by the multimodal sensor network to build a digital twin model and bind the drug's production data to the digital twin model.

[0100] S42: Quality prediction simulation, inputting the data collected by the multimodal sensor network into the digital twin model, performing multi-drug production simulation, fine-tuning the parameters, and solving the optimal combination of process parameters in reverse;

[0101] S43: Root cause analysis of abnormal data. Based on the data results of the anomaly detection algorithm, the digital twin model is used to perform root cause analysis of abnormal data.

[0102] S44: Dynamic model calibration, which uses dynamic calibration mechanism and data feedback mechanism to optimize and calibrate the model.

[0103] Preferably, when performing root cause analysis of abnormal data using a digital twin model based on the data results of the anomaly detection algorithm, the specific steps include:

[0104] S51: Abnormal data filtering, extracting batch data for which the abnormal detection algorithm triggers an alert;

[0105] S52: Causal graph construction, defining potential causal factors, and constructing the causal graph using the CausalNLP framework;

[0106] S53: Causal inference first verifies the root cause through counterfactual analysis, and then uses Bayesian network-based calculation of the contribution of each factor to calculate the confidence level.

[0107] As a preferred method, when calculating the contribution of each factor using a Bayesian network, the underlying formula is as follows:

[0108]

[0109] Among them, C i Represents variable X i The contribution of E(Y) represents the expected value of the result Y. Represents the forced variable X i Value Represents the forced variable X i Value The absolute value of the average change in Y after that.

[0110] A drug manufacturing quality identification system, comprising:

[0111] The data management module is responsible for integrating, cleaning, and storing multi-source heterogeneous data to build a standardized data lake;

[0112] Edge computing nodes are used to deploy lightweight AI models on devices, process sensor data in real time, and perform preliminary anomaly detection.

[0113] The data analysis module is used to analyze historical and real-time data through machine learning and statistical methods to identify quality trends, pinpoint the root causes of deviations, and generate decision recommendations.

[0114] The model simulation module is used to build virtual production lines based on digital twin technology, simulate the impact of process parameters on quality attributes, predict risks, and optimize production strategies.

[0115] Example 1

[0116] A drug manufacturing quality identification method based on the above technical solution, in its specific implementation, includes:

[0117] 1. Data collection and integration

[0118] First, a multimodal sensor network is deployed, with specific equipment including: a linear CCD camera to capture images of the tablet surface and detect cracks and missing corners (resolution 0.05mm / pixel); a micro balance to weigh tablets in real time (accuracy ±0.1mg) and monitor tablet weight differences; a laser leak detector to scan the sealing of aluminum-plastic packaging (detection accuracy ±5μm); and an RFID reader to bind batch IDs to pallets and track the production process.

[0119] Next, the data lake was built: Apache NiFi was used to integrate sensor data (1000 data points per second), MES process parameters (tablet pressure 50kN, coating speed 30rpm), and LIMS detection results (dissolution rate 83%) into AWS LakeFormation, and indexes were created by batch ID, timestamp, and device ID.

[0120] 2. Real-time monitoring and anomaly detection

[0121] Edge preprocessing: First, temporal feature extraction is performed: the approximate entropy of the tablet press vibration signal (m=2, r=0.2σ) is calculated, and abnormal vibrations are detected; then, frequency domain feature extraction is performed: the sound signal from the coating pan is decomposed into 5-level wavelet packets, and the energy entropy is calculated. Among them, W j (n) represents the coefficients of the nth wavelet packet in the j-th layer, E j Let E be the wavelet packet energy of the j-th layer. total H is the sum of the energies of all layers, where H is the energy entropy value and J is the total number of layers in the wavelet packet decomposition. The system identifies unstable rotational speeds and finally extracts spatial features: the ratio of defective areas in the tablet is calculated through image segmentation, with a threshold set to... Among them, R defect For the surface area defect of the drug, A defect Let A be the pixel area of ​​the defect region. total The total projected area of ​​the tablet is R, and the threshold is set to R. defect An alarm is triggered when the value is >0.5%.

[0122] Multi-level anomaly detection

[0123] Isolation forest detection: Setting a path length threshold for patch weight data. Abnormal scoring If s(x,n)>0.65, it is considered abnormal.

[0124] Autoencoder reconstruction detection: Training an autoencoder on normal tablet images (input size 256×256, hidden layer dimension 64), reconstruction error. If □>1.5Q 0.95 (Q 0.95 If the value is 120, an alarm will be triggered.

[0125] Dynamic SPC control: Dynamic threshold adjustment is applied to dissolution data, and the update formula is as follows:

[0126] μ t =λμ t-1 +(1-λ)x t ;

[0127]

[0128] λ = 0.95;

[0129] The control limit is set to μ. t ±3σ t .

[0130] Confidence fusion:

[0131] Weighted voting weights: Isolation Forest (40%), Autoencoder (35%), Dynamic SPC (25%). If the total confidence is >70%, a CAPA ticket will be triggered.

[0132] 3. Intelligent quality prediction and root cause analysis (S13-S53)

[0133] The first step is digital twin modeling: build a tableting process model in ANSYS Twin Builder, with input parameters including pressure (50-60kN), punch speed (10-20mm / s), material flowability (Carr index ≥25), and output tablet hardness (80-120N) and dissolution rate (≥80%).

[0134] The second step is counterfactual root cause analysis: causal graph variables: tableting pressure (X1), material particle size (X2), ambient humidity (→3).

[0135] Bayesian contribution calculation: Insufficient lock-in pressure is the main cause (65% overall contribution), and it is recommended to calibrate the pressure to 55±2kN.

[0136] 4. Automated processing and feedback optimization

[0137] Blockchain Evidence Storage: Record the anomalous events (dissolution rate 72%), root cause analysis reports, and calibration operations of batch BATCH_2023-001 on Hyperledger Fabric, generating an immutable hash value (SHA-256).

[0138] Model dynamic optimization: The calibrated data (pressure 55kN → dissolution rate 85%) is fed back into the digital twin model, and the hidden layer weights are updated using incremental learning (learning rate η = 0.001).

[0139] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A method for identifying the quality of drug production, characterized in that: Includes the following steps: S11: Data acquisition and integration. Input drug production data, testing data and external data, and clean, format and store the data to establish a full-domain data asset catalog. Among them, drug production data includes equipment parameters for drug production, process steps for drug production and real-time sensor data during drug production. S12: Real-time monitoring and anomaly detection. Real-time data of finished drug products is collected, and edge computing technology is used to process the collected data in real time. Then, anomaly detection algorithms are used for analysis. S13: Intelligent quality prediction and root cause analysis. Establish a digital twin simulation model and input the drug's production data, inspection data, and external data into the established digital twin simulation model to predict drug quality and perform root cause analysis on abnormal drug data. S14: Automated processing and closed-loop control, using blockchain technology to record drug production batches and automatically process drugs with quality defects; S15: Feedback optimization closed loop, feeding back deviation processing results and customer complaint data to the data lake to optimize the anomaly detection algorithm and digital twin simulation model.

2. The method for identifying drug production quality according to claim 1, characterized in that: The process of collecting real-time drug data, processing the collected data in real-time using edge computing technology, and then analyzing the data using anomaly detection algorithms specifically includes: S21: Data acquisition, deployment of a multimodal sensor network, and acquisition of real-time data of finished drug products using the multimodal sensor network, including a linear CCD camera, a micro balance, a laser leak detector, and an RFID reader; S22: Edge preprocessing, which utilizes edge computing technology to preprocess, extract features, and fuse features from the data collected by the multimodal sensor network; S23: Multi-level anomaly detection, which uses multiple anomaly detection algorithms to detect anomalies in the data of finished drug products.

3. The method for identifying drug production quality according to claim 2, characterized in that: When using edge computing technology to preprocess, extract features, and fuse features from data collected by multimodal sensor networks, the resulting feature data includes: A11: Time-domain features, including approximate entropy, where the calculation principle formula is: ApEn(m,r)=□ m (r)-□ m+1 (r), where m is the embedding dimension and m = 2, r is the similarity threshold and r = 0.2σ, □ m (r) represents the average log probability of pattern repetition in dimension m, and σ is the standard deviation of the time series. A12: Frequency domain features, including wavelet packet energy entropy, where the calculation principle formula is: Among them, W j (n) represents the coefficients of the nth wavelet packet in the j-th layer, E j Let E be the wavelet packet energy of the j-th layer. total H is the sum of the energies of all layers, where H is the energy entropy value and J is the total number of layers in the wavelet packet decomposition. A13: Spatial characteristics, including the ratio of surface defect area to drug surface, where the calculation principle formula is: Among them, R defect For the surface area defect of the drug, A defect The pixel area of ​​the defect region. A total This represents the total projected area of ​​the tablet.

4. The method for identifying drug production quality according to claim 3, characterized in that: When using various anomaly detection algorithms to detect anomalies in finished drug product data, the specific methods include: S31: Multi-source data synchronization, time alignment of multi-source data through NTP protocol, and spatial registration through mechanical coordinate and visual coordinate transformation matrix; S32: Multi-model parallel detection, which uses three anomaly detection algorithms in parallel: an isolated forest-based anomaly detection algorithm, an autoencoder reconstruction detection algorithm, and a dynamic statistical process control method; S33: Confidence fusion decision, which uses a weighted voting mechanism to weight and fuse the data results of the three anomaly detection algorithms; S34: Dynamic threshold adjustment, dynamically adjusting the threshold for abnormal alarms based on the slow changes in the test results of the finished drug product.

5. The method for identifying drug production quality according to claim 4, characterized in that: When using three anomaly detection algorithms in parallel—isolated forest-based anomaly detection algorithm, autoencoder reconstruction detection algorithm, and dynamic statistical process control method—the underlying principle is as follows: A21: Anomaly detection algorithm based on isolated forest: Where s(x,n) represents the anomaly score of sample x, E(h(x)) represents the average path length of sample x across all trees, c(n) represents the path length correction term used to standardize the results under different sample sizes, and n is the subsample size used during training. H(k) is the harmonic number, and H(k)≈ln(k)+0.5772; A22: Autoencoder Reconstruction Detection Algorithm Alarm condition: □>1.5Q 0.95 ; Where x is the original input feature vector, For the autoencoder to reconstruct the output, Q 0.95 The 95th percentile of the reconstruction error of normal samples in the training set; A23: Dynamic Statistical Process Control Methods Where USL is the upper limit of specifications, LSL is the lower limit of specifications, μ is the process mean, and σ is the process standard deviation.

6. The method for identifying drug production quality according to claim 5, characterized in that: When dynamically adjusting the threshold for abnormal alarms based on the slow changes in the test results of the finished drug product, the underlying formula is: m t =lm t-1 +(1-λ)x t ; λ = 0.95; Where, μ t The process mean estimate at time t. Let be the process variance estimate at time t, λ be the forgetting factor, and x be the value of the process variance estimate at time t. t Let μ be the observed value at time t. t-1 This is the process mean of the previous time step. This represents the process variance at the previous time step.

7. The method for identifying drug production quality according to claim 6, characterized in that: When establishing a digital twin simulation model and inputting drug production data, testing data, and external data into the model to predict drug quality and perform root cause analysis on abnormal drug data, the specific steps include: S41: Data input and model initialization. Input the drug's production data, testing data, external data, and data collected by the multimodal sensor network to build a digital twin model and bind the drug's production data to the digital twin model. S42: Quality prediction simulation, inputting the data collected by the multimodal sensor network into the digital twin model, performing multi-drug production simulation, fine-tuning the parameters, and solving the optimal combination of process parameters in reverse; S43: Root cause analysis of abnormal data. Based on the data results of the anomaly detection algorithm, the digital twin model is used to perform root cause analysis of abnormal data. S44: Dynamic model calibration, which uses dynamic calibration mechanism and data feedback mechanism to optimize and calibrate the model.

8. The method for identifying drug production quality according to claim 7, characterized in that: When performing root cause analysis of abnormal data using a digital twin model based on the data results from the anomaly detection algorithm, the specific steps include: S51: Abnormal data filtering, extracting batch data for which the abnormal detection algorithm triggers an alert; S52: Causal graph construction, defining potential causal factors, and constructing the causal graph using the CausalNLP framework; S53: Causal inference first verifies the root cause through counterfactual analysis, and then uses Bayesian network-based calculation of the contribution of each factor to calculate the confidence level.

9. A method for identifying drug production quality according to claim 8, characterized in that: When calculating the contribution of each factor using a Bayesian network, the underlying formula is as follows: Among them, C i Represents variable X i The contribution of E(Y) represents the expected value of the result Y. Represents the forced variable X i Value Represents the forced variable X i Value The absolute value of the average change in Y after that.

10. A drug manufacturing quality identification system, characterized in that: include: The data management module is responsible for integrating, cleaning, and storing multi-source heterogeneous data to build a standardized data lake; Edge computing nodes are used to deploy lightweight AI models on devices, process sensor data in real time, and perform preliminary anomaly detection. The data analysis module is used to analyze historical and real-time data through machine learning and statistical methods to identify quality trends, pinpoint the root causes of deviations, and generate decision recommendations. The model simulation module is used to build virtual production lines based on digital twin technology, simulate the impact of process parameters on quality attributes, predict risks, and optimize production strategies.