An abnormality identification and real-time collection system for multi-source heterogeneous medical data
The system for anomaly identification and real-time acquisition of multi-source heterogeneous pharmaceutical data has solved the problem of personalized medication safety assurance in the pharmaceutical supply chain and clinical medication process, and has achieved real-time and precise control and anomaly detection of drug safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUACHEN HONGYI (BEIJING) TECHNOLOGY CO LTD
- Filing Date
- 2026-03-03
- Publication Date
- 2026-06-02
AI Technical Summary
The existing pharmaceutical supply chain and clinical drug use process lack safety guarantees for personalized medication, and cannot calculate the maximum safe superposition dose for individuals in real time. In particular, it cannot achieve continuous and accurate drug safety management in scenarios where multiple drugs coexist.
An anomaly identification and real-time acquisition system for multi-source heterogeneous pharmaceutical data is adopted. Through an event bus module, a decay function generation module, a drug property matrix mapping module, an inverse solution module, and an online regression update module, the system achieves continuous time-series processing of data without duplication or omission. Combining individual genotype and environmental factors, the system outputs the maximum safe superposition dose when multiple drugs coexist in real time, and performs anomaly decision-making and data stream playback and updates.
It achieves second-level aggregation and real-time adjustment of drug safety management, reduces the false alarm rate of dosage warning, shortens the anomaly location time, and supports precise control of personalized medication.
Smart Images

Figure CN122136030A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical information technology, and in particular to an anomaly identification and real-time acquisition system for multi-source heterogeneous medical data. Background Technology
[0002] The safety assurance of the entire pharmaceutical supply chain and clinical medication process currently relies on a dual-track approach of environmental parameter monitoring and medication dosage review. However, each track is built and maintained independently, with different hardware interfaces, sampling frequencies, and data models, which cannot support the continuous, accurate, and real-time control required for personalized medication. There is also a lack of a mechanism for solving the individualized safe dosage of multiple drugs. The current dosage review uses a fixed parameter set of the population PK equation, which neither incorporates patient genotype and physiological characteristics online nor couples environmental cumulative corrections. Therefore, it is impossible to calculate the individualized maximum safe superposition dosage in real time in the case of multiple drugs coexisting. To address this, we propose an anomaly identification and real-time acquisition system for multi-source heterogeneous pharmaceutical data. Summary of the Invention
[0003] The purpose of this invention is to solve the problems mentioned in the background art by proposing an anomaly identification and real-time acquisition system for multi-source heterogeneous medical data.
[0004] To achieve the above objectives, the present invention adopts the following technical solution: A system for anomaly identification and real-time acquisition of multi-source heterogeneous pharmaceutical data includes: The event bus module ensures a single-processing mechanism that prevents data stream duplication and omission, continuously receiving time-series data containing physicochemical properties and acquisition time and location information from drug preparation, delivery, storage and clinical use, as well as patient-specific parameters. The decay function generation module, within an independent computing unit with real-time data processing capabilities, performs time window aggregation on the continuous time-series data to generate a function that changes over time and reflects the cumulative impact of environmental factors on the drug, thereby achieving time-series standardization of multi-source heterogeneous data. The drug property matrix mapping module maps the function to the drug property correction coefficient of the interaction between drug components, obtains the drug property matrix with elements dynamically updated over time, and constructs a dynamic correlation model between environmental cumulative effect and drug property. The reverse solution module uses the pharmacological matrix as boundary conditions to perform individual-adaptive reverse solution on the differential equations describing the drug metabolism and efficacy of the population, and outputs the maximum safe superposition dose when multiple drugs coexist in real time. The abnormal decision module compares the maximum safe superposition dose with the current prescription dose, generates an abnormal prompt event graded according to the excess ratio, and writes it back to the event bus module; The online regression update module uses a complete data sequence with no adverse events and only slight environmental data fluctuations to adjust the drug property correction coefficient and the dynamic parameters of the function execution system, thereby optimizing the inverse solution module.
[0005] Compared with existing technologies, the advantages of this invention are: 1. Under the unified event bus framework, this invention aggregates, aligns, and fills in gaps in heterogeneous data from all stages of preparation, transportation, storage, and clinical use within seconds, forming a continuous time-series stream without repetition or omission. Through dynamic mapping of environmental accumulation and pharmacological correction coefficients, combined with online reverse solving of population pharmacokinetic equations based on individual genotypes, the maximum safe superposition dose of multiple drugs is output in real time. Supplemented by summary verification triggering data stream playback and periodic online regression updates, drug safety management is achieved.
[0006] 2: The attenuation function converts the cumulative amounts of temperature, humidity, light, and vibration into dynamic drug property correction coefficients in real time. The matrix elements are updated instantly with environmental exposure. For the first time, transportation heat exposure and cold storage vibration are quantified into drug interaction multiples, so that drug property assessment no longer depends on fixed empirical values and can reflect the external conditions that the drug has actually experienced at any time.
[0007] 3: The reverse solution uses environmental cumulative correction, individual genotype and physiological characteristics as boundaries to calculate the population pharmacokinetic equation online and output the maximum safe superposition dose under multiple drug coexistence; the abnormal decision is graded and warned according to the excess ratio, the summary verification triggers data stream playback, the event loss or repeated calculation can be detected and recalculated in real time, the dose warning false alarm rate is reduced and the abnormal location time is significantly shortened.
[0008] 4. The online regression module automatically retrains the response surface using data sequences with no adverse events and slight environmental fluctuations. It is updated quarterly, and the model parameters continue to evolve, preventing drift over long-term operation. The data lineage tag binds device identity, transmission integrity, and algorithm processing traces to events, supporting second-level audit rollback to achieve safe management of pharmaceuticals. Attached Figure Description
[0009] Figure 1 This is a flowchart of an anomaly identification and real-time acquisition system for multi-source heterogeneous medical data proposed in this invention. Detailed Implementation
[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0011] Reference Figure 1 An anomaly identification and real-time acquisition system for multi-source heterogeneous pharmaceutical data is proposed. This system includes a container orchestration platform that, based on a pre-defined deployment description file, sequentially launches an event bus container, a decay function generation container, a drug property matrix mapping container, a reverse solving container, an anomaly decision container, an online regression update container, and a columnar database container within a Kubernetes cluster in a data center. (The event bus module, decay function generation module, drug property matrix mapping module, reverse solving module, anomaly decision module, online regression update module, and columnar database module are deployed and run as independent containers through the container orchestration platform.) The image version number, CPU quota, memory quota, persistent volume mount point, environment variables, probe address, and network policy of each container are declared once through the description file. The image pull strategy is set to always re-pull to ensure complete consistency between the on-site deployment and the development baseline.
[0012] After the event bus container starts, it first creates five types of topics within the message queue cluster: preparation topics, delivery topics, storage topics, clinical use topics, and anomaly topics. Topic-level parameters are uniformly configured: replication factor is set to three, minimum synchronous replication is set to two, maximum message size is set to four megabytes, log segment size is set to 512 megabytes, and log retention time is set to seven days, thirty days, ninety days, one hundred and eighty days, and three hundred and sixty-five days respectively, depending on the risk level of each stage. The number of partitions for each topic is set according to the actual number of devices: the number of partitions for the preparation topic is equal to twice the total number of programmable logic controllers (PLCs) for sterilizers, PLCs for freeze dryers, and weighing modules for filling lines; the number of partitions for the delivery topic is equal to the number of nodes for vehicle-mounted cold chain recorders; the number of partitions for the storage topic is equal to the number of low-power long-range wireless gateways in the hospital cold storage; the number of partitions for the clinical use topic is equal to the total number of intravenous administration pumps and blood drug concentration monitors in the ward; and the anomaly topic is a single partition to ensure the global order of anomaly events.
[0013] The event bus container then opens five types of TCP socket listening ports to receive data from different stages: The first type of port uses the industrial Ethernet protocol to directly capture data units encapsulated in the programmable logic controllers of the equipment in the preparation stage, including the temperature probe of the sterilizer sampling value per second, the vacuum sensor of the freeze dryer sampling value every two seconds, and the weighing module of the filling line instantaneous sampling value every 0.1 seconds. Each data unit is appended with a 64-bit device serial number as a unique identifier prefix. The second type of port receives data from the transmission stage (including continuous time-series data containing physicochemical properties and acquisition time and location information) via 4G network IP subscription. This includes temperature, humidity, light intensity, and vibration acceleration arrays reported every 60 seconds by the vehicle-mounted cold chain recorder, and latitude and longitude reported every 30 seconds by the GPS module. The message body is a little-endian binary byte stream. The third type of port uses the HTTP protocol... The system receives and stores data (including continuous time-series data on physicochemical properties, acquisition time, and location information), including temperature and humidity data broadcast every five minutes by the low-power long-range wireless sensor in the hospital cold storage, warehouse door magnetic switch events, and physical impact events triggered by the rotary vibration sensor during drug handling with an acceleration greater than 0.5 times gravity and a duration greater than or equal to 200 milliseconds. The data is aggregated into a JSON array by the gateway and then pushed in batches. The fourth type of port receives abnormal event objects written back by the abnormal decision container through the internal RPC protocol. The fifth type of port receives clinical usage data through the HTTPS protocol. The data format conforms to the FHIR standard JSON and includes the flow rate sampling value per second of the intravenous administration pump, the sampling value every thirty seconds of the blood drug concentration monitor, and the sampling value every five seconds of the patient's vital signs monitor (heart rate, blood pressure, and blood oxygen saturation).
[0014] After receiving five types of data, the event bus container constructs event objects. The object fields include a unique event identifier, a unique data source identifier, a millisecond-level timestamp, a double-precision sampled value array, a cyclic redundancy check (CRC) code, an interpolation flag, and a data lineage tag (including device serial number, data transmission message CRC, interpolation or padding flag, and clinical use flag). The container performs cyclic redundancy check on the byte stream. If the check fails, the event is discarded and written to the audit log. If the check succeeds, the event object is serialized into a binary array and written to the corresponding topic. During writing, the data source identifier is used as the key to ensure the order of events from the same data source. After writing, the container generates a new event arrival signal for the decay function. The signal body only contains the event identifier and the topic partition offset to avoid redundant data transmission. In addition, the event bus container periodically retrieves individual patient parameters, including patient age, gender, weight, and liver enzyme genotype classification values, through the hospital information system's REST interface. The retrieval period is set to 300 seconds. After retrieval, the parameters are written to the clinical use topic, with the patient identifier as the key, to ensure that subsequent modules can be associated based on the patient identifier.
[0015] The decay function generation container includes a time alignment submodule and a missing data filling submodule. The time alignment submodule uses linear interpolation to uniformly adjust the continuous time series data with different sampling frequencies to a sampling frequency of one data point per second. The missing data filling submodule fills in the continuous time series data missing three or more consecutive time points using a linear extrapolation method. For data missing more than three consecutive time points, it marks the data as unfillable and triggers a first-level anomaly. After receiving a new event arrival signal, the decay function generation container pulls the event object from the topic partition offset and firstly stores the event object in the local RocksDB buffer. The buffer key is formed by concatenating the data source identifier and the partition offset, and the value is the binary number of the event object. The group has a buffer retention period of ten minutes for rapid replay in case of abnormal playback. Then, it enters the time alignment submodule, maintaining a latest point table with the data source identifier as the primary key. The table structure fields include the data source identifier, the latest millisecond timestamp, and an array of latest double-precision values. If the difference between the new event timestamp and the latest point timestamp is less than 1000 milliseconds, the latest point is directly replaced. If the difference is greater than or equal to 10000 milliseconds and less than 3000 milliseconds, an interpolation event of one point per second is generated within that interval using linear interpolation. The interpolation formula is: the interpolation point value equals the latest point value plus (new event value minus latest point value) multiplied by (interpolation point timestamp minus latest point timestamp) divided by (new event timestamp minus latest point timestamp). The interpolation event object has the same structure as the original event object, but... The interpolation flag is set to true; if the difference is greater than 3,000 milliseconds, a level 1 exception event is generated and written to the exception topic, while the interpolation is skipped and the latest point is directly replaced; after the interpolation is completed, the missing point filling submodule is entered. For consecutive missing points with a number of missing points less than or equal to three, linear extrapolation is used. The extrapolation formula is: the extrapolated point value equals the latest point value plus (latest point value minus the second-to-last point value) multiplied by (extrapolated point timestamp minus the latest point timestamp) divided by (latest point timestamp minus the second-to-last point timestamp). If the number of missing points is greater than three, a level 1 exception event is also generated; the completed event flow enters the sliding window accumulation subprocess. The window width is fixed at 60 seconds, the step size is 1 second, and the accumulation calculation rule is: the temperature accumulation is equal to the difference between the current second temperature value and 25 degrees Celsius. The maximum value of zero is accumulated. The humidity accumulation is equal to the sum of the difference between the current second relative humidity value minus 45% and the maximum value of zero. The light accumulation is equal to the sum of the difference between the current second light value minus 100 lux and the maximum value of zero. The vibration acceleration accumulation is equal to the sum of the difference between the current second absolute value of acceleration minus zero times the gravitational acceleration and the maximum value of zero. The accumulation is calculated once per second and an accumulation object is generated. The object fields include the window start millisecond timestamp, the window length in seconds, the temperature accumulation, the humidity accumulation, the light accumulation, and the vibration acceleration accumulation. The accumulation object is persisted to the ClickHouse columnar database. The partition key is the date integer obtained by dividing the window start millisecond timestamp by 86.4 million. At the same time, it is pushed to the drug property matrix mapping container.
[0016] The drug property matrix mapping container includes a four-dimensional response surface storage unit, a bilinear interpolation unit, and a dynamic offset unit. The four-dimensional response surface storage unit stores the discrete grid points of the surface model. The surface model has four dimensions: temperature, humidity, light intensity, and cumulative vibration acceleration, and the response value is the drug property correction coefficient between drug components. Each dimension of the discrete grid points is divided at a fixed interval and is generated based on fitting clinical trial data and simulation data. The bilinear interpolation unit is used to calculate the factor by which the environmental factors corresponding to the input cumulative amount cause drug property amplification based on the four adjacent discrete grid points around the input cumulative amount using a linear interpolation algorithm. The dynamic offset unit is used to multiply the drug property amplification factor with the correction coefficients corresponding to the patient-specific parameters such as weight and liver enzyme genotype to generate the final drug property correction coefficient between drug components. After receiving the cumulative quantity object, the drug property matrix mapping container first queries the local pre-set four-dimensional response surface storage unit. The surface model uses temperature cumulative quantity, humidity cumulative quantity, light cumulative quantity, and vibration acceleration cumulative quantity as independent variables, and the drug property amplification coefficient between drug components as the dependent variable. The discrete grid step size is 20 degrees Celsius per hour, 10 percent per hour, 500 lux per hour, and 0.5 gravitational acceleration per hour, respectively. The grid point values are obtained by jointly fitting clinical experimental data and physiological pharmacokinetic simulation data. The fitting process uses fourth-order polynomial regression and undergoes 10-fold cross-validation. Only those with a determination coefficient greater than 0.85 can be stored in the database. Four-dimensional linear interpolation is performed on the input cumulative quantity. During interpolation, one-dimensional linear interpolation is first performed along the temperature axis, then along the humidity axis, then along the light axis, and finally along the vibration axis. The environmental amplification factor is obtained by calculating the value of the environmental amplification factor. Then, patient-specific parameters are read, including weight and liver enzyme genotype classification. Liver enzyme genotypes are divided into four categories: weak metabolizer, intermediate metabolizer, strong metabolizer, and ultra-rapid metabolizer. Each category corresponds to a preset gene correction coefficient: 0.5 for weak metabolizer, 1 for intermediate metabolizer, 1.5 for strong metabolizer, and 2 for ultra-rapid metabolizer. The final drug efficacy correction coefficient is equal to the environmental amplification factor multiplied by the gene correction coefficient. The final drug efficacy correction coefficient is written to a matrix object. The matrix object fields include a matrix identifier, a valid start millisecond timestamp, a valid end millisecond timestamp, and a two-dimensional double-precision array. The number of rows and columns in the array equals the total number of drug components. The array element values are the drug efficacy correction coefficients between each pair of drug components. The matrix object is published to a Redis stream and persisted to a columnar database.
[0017] Upon receiving the matrix object, the inverse solver immediately loads the population pharmacokinetic differential equations. These equations take drug dose as input, blood drug concentration as output, and hepatic clearance, volume of distribution, and absorption rate constant as parameters. The equations are in the following form: Where C is the blood drug concentration, D is the dose, and k is the concentration of the drug in the blood. aLet CL be the absorption rate constant, CL be the hepatic clearance rate, and V be the volume of distribution. The hepatic clearance rate is corrected in real time by a pharmacokinetic correction coefficient provided by the matrix object. The correction formula is: individual hepatic clearance rate equals population hepatic clearance rate multiplied by the pharmacokinetic correction coefficient. Using patient physiological parameter vectors as boundary conditions, including age, sex, weight, liver enzyme genotype, and serum creatinine value, the adjoint variable method is used to solve the steady-state exposure objective function in reverse. The solution process is as follows: initialize the dose vector, obtain the blood drug concentration-time curve through a forward integral differential equation system, calculate the exposure, compare the exposure with the target value, and obtain the gradient of the objective function with respect to the dose vector through inverse integration of the adjoint equation. The dose vector is updated in the negative gradient direction, with an adaptive decay strategy for the step size. The initial step size is set to 1% of the target exposure. Every ten iterations, if the objective function decreases by less than 1 / 1000, the step size is halved. The convergence threshold is set to a relative difference in exposure of less than 1%. The maximum number of iterations is set to 200. If convergence is not achieved after reaching the upper limit, a secondary abnormal event is generated and written to the abnormal topic. After convergence, the maximum safe superimposed dose vector is obtained. The vector elements are the maximum allowable doses for each drug. The dose object fields include the drug identifier, the maximum safe dose in milligrams, and the solution time in milliseconds. The dose object is persisted to the local embedded database and simultaneously pushed to the abnormal decision container.
[0018] After receiving the dosage object, the abnormal decision container retrieves the current prescription dosage from the hospital information system view. The view fields include prescription identifier, drug identifier, and prescription dosage in milligrams. It calculates the excess percentage for each drug. The excess percentage equals the prescription dosage minus the maximum safe dosage, divided by the maximum safe dosage. If the excess percentage is less than or equal to zero, it is considered normal. If the excess percentage is greater than zero but less than or equal to 20%, it is considered a Level 1 abnormality, generating an abnormal event object and writing it to an abnormality topic. Simultaneously, a Level 1 abnormality message is published to the nurse station message agent, with the message format "prescription identifier, drug name, excess percentage, suggested dosage reduction in milligrams". If the excess percentage is greater than 20%, it is considered a Level 2 abnormality, generating an abnormal event object and writing it to an abnormality topic. Simultaneously, it calls the hospital SMS gateway interface to send a Level 2 abnormality SMS to the supervising pharmacist's mobile phone number. The SMS content includes the patient's name, drug name, excess percentage, and a suggestion to immediately stop medication. The abnormal event object fields include abnormality level, excess percentage, and other parameters. Operator ID, prescription ID, and millisecond timestamp of exception generation; after the exception decision container completes the writing of the exception event, it immediately sends a verification request to the event bus container. The request body contains the SHA-256 hash digest value of the most recent cumulative quantity object in hexadecimal string form. After receiving the request, the event bus container forwards it to the decay function generation container. The decay function generation container recalculates the SHA-256 hash digest value of the cumulative quantity object within the same time window and returns it. If the two digest values are inconsistent, the event bus container triggers the data stream replay process: the original event object within the corresponding time window is reread from the message queue persistent segment and sent back to the decay function generation container to re-execute time alignment, missing value filling, and cumulative quantity calculation, regenerate the cumulative quantity object and drive the subsequent matrix mapping, reverse solution, and exception decision process until the digest values are consistent or the maximum number of replays is reached three times. If they are still inconsistent, a level 3 exception event is generated and written to the audit log, and then transferred to the manual audit queue.
[0019] The online regression update container starts a batch task at 00:10 every day. The task first queries the columnar database for all cumulative quantity objects with an anomaly level of zero over the past 24 hours. It then filters out subsets of objects with cumulative temperature changes less than 10 degrees Celsius per hour, cumulative humidity changes less than 5% per hour, cumulative light changes less than 250 lux per hour, and cumulative vibration acceleration changes less than 0.25 gravitational acceleration per hour. These filtered cumulative quantity objects are paired with their corresponding matrix objects to form training samples. The sample features are four-dimensional cumulative quantities, and the sample labels are static drug property correction coefficients. Weighted least squares regression is used for four-dimensional polynomial regression, with a regression order of second order, a loss function of mean squared error, and a regularization coefficient of 0.01. The optimization algorithm is L-B. For FGS, the maximum number of iterations is set to 500, and the convergence threshold is set to a gradient norm of less than 1 / 1000. After regression, a new grid weight file is obtained. The file format is columnar storage, and the file contains grid point coordinates and updated dependent variable values. The new file is uploaded to the hospital database, and the response surface in the in-memory database is updated. At the same time, a model version number increment message is sent, and the old version file is retained for 90 days for traceability. At the end of each quarter, the online regression update container automatically pulls all historical zero-outlier samples and re-executes the above regression process to obtain a quarterly centralized retrained model. After retraining, a model evaluation report is generated. The report includes the coefficient of determination, mean squared error, and grid point residual distribution histogram. The report is persisted to the audit storage node for review by the Pharmacy Administration Committee.
[0020] During system operation, all event objects, cumulative quantity objects, matrix objects, dose objects, and abnormal event objects are accompanied by data lineage tags. These tags include the device serial number, cyclic redundancy check (CRC) code of the data transmission message, interpolation or padding flags, model version number, container instance identifier, and generation timestamp. The tags are written to a columnar database along with the objects. The database table structure is columnar storage, with the partition key being a date integer and the sort key being a millisecond timestamp. The compression algorithm is LZ4, with a compression ratio of approximately 5:1. The storage period is set to seven years. Traceability queries are supported via a standard SQL interface. Query examples include: locating abnormal events based on the device serial number, locating transmission errors based on the CRC code, calculating the proportion of interpolation events based on the interpolation flag, and tracing back historical model effects based on the model version number. System administrators can display the data lineage chain within any time window through a visual interface. The chain nodes include original sampling, interpolation events, cumulative quantity objects, matrix objects, dose objects, and abnormal events. Nodes are associated with each other via timestamps and identifiers, forming an end-to-end traceable audit trail.
[0021] In alternative implementation schemes, if the hospital's local computer room has insufficient computing resources, the core computing containers can be migrated to a cloud container service. During migration, the container images, environment variables, network policies, and storage mount points should remain unchanged. Only the message queue cluster should be replaced with a cloud-hosted version, the object storage with a cloud-compatible interface, and the columnar database with a cloud data warehouse. After migration, the edge acquisition nodes should be connected via a dedicated line, with network latency controlled within ten milliseconds to ensure real-time performance. In extended application scenarios, the system can be connected to clinical trial centers to monitor the environmental exposure of investigational drugs during transportation and storage in real time. Combined with individual pharmacokinetic parameters of subjects, it can provide early warning of overdose risks and reduce the incidence of adverse reactions to investigational drugs. At the same time, the system can also be connected to the data platform of drug regulatory agencies to provide technical support for drug quality traceability and risk assessment.
[0022] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A system for anomaly identification and real-time acquisition of multi-source heterogeneous pharmaceutical data, characterized in that, include: The event bus module ensures a single-processing mechanism that prevents data stream duplication and omission, continuously receiving time-series data containing physicochemical properties and acquisition time and location information from drug preparation, delivery, storage and clinical use, as well as patient-specific parameters. The decay function generation module, within an independent computing unit with real-time data processing capabilities, performs time window aggregation on the continuous time-series data to generate a function that changes over time and reflects the cumulative impact of environmental factors on the drug, thereby achieving time-series standardization of multi-source heterogeneous data. The drug property matrix mapping module maps the function to the drug property correction coefficient of the interaction between drug components, obtains the drug property matrix with elements dynamically updated over time, and constructs a dynamic correlation model between environmental cumulative effect and drug property. The reverse solution module uses the pharmacological matrix as boundary conditions to perform individual-adaptive reverse solution on the differential equations describing the drug metabolism and efficacy of the population, and outputs the maximum safe superposition dose when multiple drugs coexist in real time. The abnormal decision module compares the maximum safe superposition dose with the current prescription dose, generates an abnormal prompt event graded according to the excess ratio, and writes it back to the event bus module; The online regression update module uses a complete data sequence with no adverse events and only slight environmental data fluctuations to adjust the drug property correction coefficient and the dynamic parameters of the function execution system, thereby optimizing the inverse solution module.
2. The anomaly identification and real-time acquisition system for multi-source heterogeneous medical data according to claim 1, characterized in that, The continuous time-series data containing physicochemical properties and collection time and location information comes from the second-by-second sampling value of the sterilizer temperature probe, the two-second sampling value of the freeze dryer vacuum sensor, and the instantaneous sampling value of the filling line weighing module every 0.1 seconds during the preparation process. After the sampled values are encapsulated into data units conforming to the Industrial Ethernet protocol by the device's programmable logic controller, they are directly captured by the event bus module through a socket word byte stream, and a 64-bit device serial number is appended as a string prefix that uniquely identifies each data event.
3. The anomaly identification and real-time acquisition system for multi-source heterogeneous pharmaceutical data according to claim 1, characterized in that, The continuous time-series data containing physicochemical properties and collection time and location information comes from the temperature, humidity, light intensity, and vibration acceleration reported every sixty seconds by the vehicle-mounted cold chain recorder, and the latitude and longitude reported every thirty seconds by the global positioning system module during the transmission process. The event bus module receives the above data via a fourth-generation mobile communication subscription method, and the message body adopts a binary byte stream format.
4. The anomaly identification and real-time acquisition system for multi-source heterogeneous pharmaceutical data according to claim 1, characterized in that, The continuous time-series data containing physical and chemical properties and collection time and location information is stored in the hospital cold storage. The data is generated from the temperature and humidity broadcast every five minutes by the low-power long-range wireless sensor, the magnetic switch event of the warehouse door, and the physical impact event triggered by the rotary vibration sensor during drug handling, which is greater than 0.5 times the gravitational acceleration and lasts for more than or equal to 200 milliseconds. The aforementioned data is aggregated into an array by a low-power long-range wireless gateway and then pushed in batches to the event bus module via the Hypertext Transfer Protocol.
5. The anomaly identification and real-time acquisition system for multi-source heterogeneous pharmaceutical data according to claim 1, characterized in that, The attenuation function generation module includes: The time alignment submodule uses a linear interpolation method to uniformly adjust the continuous time-series data with different sampling frequencies to a sampling frequency of one data point per second. The missing data filling submodule fills in the missing continuous time series data for three or more consecutive time points using a linear extrapolation method; For data missing more than three consecutive time points, it is marked as an unfillable data analysis fault and a Level 1 anomaly is triggered.
6. The anomaly identification and real-time acquisition system for multi-source heterogeneous pharmaceutical data according to claim 5, characterized in that, The attenuation function generation module performs a sliding window accumulation operation on the data points per second that have been processed and filled by the time alignment submodule and the missing data filling submodule, and outputs four sets of accumulated values: temperature accumulation equals the continuous summation of the difference between temperature and 25°C within a 1-second sampling interval; humidity accumulation equals the continuous summation of the difference between relative humidity and 45% within a 1-second sampling interval; illumination accumulation equals the continuous summation of the difference between illumination and 100 lx within a 1-second sampling interval; and vibration acceleration accumulation equals the continuous summation of the difference between the absolute value of acceleration and zero-point gravitational acceleration within a 1-second sampling interval. When any of the above accumulated values is less than zero, it is forcibly set to zero.
7. The anomaly identification and real-time acquisition system for multi-source heterogeneous pharmaceutical data according to claim 1, characterized in that, The pharmacology matrix mapping module includes: The four-dimensional response surface storage unit is used to store the discrete grid points of the surface model. The surface model has four dimensions: temperature, humidity, light, and cumulative vibration acceleration, and the response value is the drug property correction coefficient between drug components. Each dimension of the discrete grid points is divided at fixed intervals and is generated based on fitting clinical trial data and simulation data; The bilinear interpolation unit is used to calculate the factor by which the environmental factors cause drug amplification based on the input cumulative amount using a linear interpolation algorithm, based on four adjacent discrete grid points around the input cumulative amount. The dynamic offset unit is used to multiply the drug amplification factor with the correction coefficients corresponding to the patient-specific parameters such as weight and liver enzyme genotype to generate the final drug component inter-drug property correction coefficients.
8. The anomaly identification and real-time acquisition system for multi-source heterogeneous pharmaceutical data according to claim 2, characterized in that, When constructing an event object, the event bus module synchronously generates a data lineage tag. Each string that uniquely identifies each data event is associated with a set of data lineage tags. The data lineage tags include the device serial number, the data transmission message cyclic redundancy check code, and interpolation or padding flags. The data lineage tag is written to the event bus module along with the event, so that the device, link or link that caused the data anomaly can be located during subsequent auditing or model rollback.
9. A system for anomaly identification and real-time acquisition of multi-source heterogeneous medical data according to claim 1, characterized in that, The online regression update module triggers a batch export task at a predetermined time every morning, writing the complete data sequence marked as having no adverse events and with slight environmental data fluctuations within the previous 24 hours into a columnar storage file and uploading it to the hospital database. The accumulated data is used to retrain the model quarterly.
10. A system for anomaly identification and real-time acquisition of multi-source heterogeneous medical data according to claim 1, characterized in that, Upon receiving the anomaly notification event written back by the anomaly decision module, the event bus module immediately sends a verification request containing the SHA-256 hash digest value of the most recent time window accumulation to the decay function generation module. The decay function generation module returns a recalculated hash digest value with the same algorithm. If the two are inconsistent, data stream replay is triggered to eliminate potential event loss or duplicate calculation. The data stream replay involves re-inputting the original data stream of the time window into the system for secondary processing.