Drug use real-time regulation and control system based on reinforcement learning and static rule base

By using a real-time medication control system based on reinforcement learning and a static rule base, protocol fingerprints and safety policy capsules are generated, solving the problem of parameter inconsistency caused by the lag in updating the static rule base. This achieves efficient automated and consistent management of infusion parameters, reducing infusion delays and the complexity of rule base governance.

CN121789892APending Publication Date: 2026-04-03THE 926TH HOSPITAL OF THE CHINESE PEOPLES LIBERATION ARMY JOINT LOGISTICS SUPPORT FORCE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In high-frequency intravenous infusion scenarios, the delayed update and activation time of the static rule base leads to inconsistent parameter semantics when the medical order system and the smart infusion pump are interconnected, resulting in automatic programming failure or the need for manual input, which increases the infusion start delay and the complexity of rule base governance.

Method used

A real-time medication control system based on reinforcement learning and a static rule base is adopted. By generating protocol fingerprints and candidate safety policy capsules, it performs device capability list handshakes and conflict classification, restricts the basic mode without drug selection, and enables controlled emergency tokens in abnormal situations to ensure parameter consistency and traceability.

Benefits of technology

It reduces the risk of matching failures without changing hard constraints, enhances the automation and consistency of input parameters, simplifies the rule base governance process, reduces input delays and alarms, and improves the traceability of execution parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789892A_ABST
    Figure CN121789892A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time medication regulation and control system based on reinforcement learning and a static rule base, relates to the technical field of medical medication control, and aims to collect information of a doctor's advice system, a pharmacy preparation system and the static rule base at a doctor's advice signing, checking or executing trigger point, unify a medicine entity and a metering system, distinguish negotiable fields and non-negotiable fields, and provide a real-time medication regulation and control system. Protocol fingerprints are generated and candidate security policy capsules are formed. Handshake and conflict grading are carried out according to the equipment capability list, and negotiable fields are converted, cut and signed to generate a final security policy capsule. The execution end enters a guardrail interconnection mode after signature verification, initial setting and titration adjustment are carried out according to capsule constraints, and a basic mode without drug selection is limited; and the controlled emergency token is enabled when the abnormity occurs. The auditing link is written in the whole process, and the governance action is generated by reinforcement learning without changing the hard limit, so that the matching failure and bypass risk are reduced, and the consistency of tracing and operation and maintenance is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical drug control technology, specifically a real-time drug control system based on reinforcement learning and a static rule base. Background Technology

[0002] In high-frequency intravenous infusion scenarios such as intensive care, surgical anesthesia, emergency resuscitation, and tumor chemotherapy, continuous infusion of vasopressors, sedatives and analgesics, insulin, and anti-infectives is often required, with titration monitoring of indicators. To reduce human programming errors and process traceability, most existing technologies utilize intelligent infusion pumps equipped with drug libraries and infusion guardrails, interconnected with medical order systems and pharmacy preparation systems: physicians select drugs, concentrations or preparation methods, dosages or rates, and target ranges in the medical order system; the pharmacy prepares the infusion according to the institution's standard concentrations and preparation specifications; through the interconnected process, drug identification, units of measurement, concentrations, and limits are transmitted to the infusion device, which performs protocol matching and limit verification based on its local drug library. The aforementioned drug libraries are all static rule bases, specifically containing information such as drug names or codes, allowed concentration sets, hard and soft limits, step sizes and minimum adjustment intervals, unit conversions, and rounding resolutions, and are updated periodically through periodic releases and terminal activation.

[0003] Due to factors such as drug shortages and substitutions, departmental differences in medication use, differences in equipment models, and changes in information system dictionaries, the standard concentrations and parameter expressions of institutions may change. There is also a time lag between terminal updates and activation, which can easily lead to differences in the identification, concentration expression, unit conversion, rounding tolerance, and limit boundaries of the same drug between different systems and different devices.

[0004] During automated programming, infusion devices need to strictly match drug identifiers and parameters. Missing entries, outdated versions, or semantic inconsistencies can lead to matching failures or interruptions, requiring nurses to use basic modes or manually input parameters. This process causes discrepancies between the guardrail verification path and the institutional rulebook, increasing differences in unit, concentration, or rate selection and requiring repetitive verification. This results in delays in infusion initiation or adjustment, increased alarms, and the inability to trace actual execution parameters and the chain of responsibility, severely impacting subsequent rulebook governance and quality control. Furthermore, the correlation between matching failures or manually entered execution records and the institutional rulebook version is weak, making it difficult to distinguish whether the failure is due to missing drug entries, inactive versions, or unit conversion differences, increasing the complexity of troubleshooting, verification, and unified configuration.

[0005] Therefore, the current technical problem is that when the medical order system is automatically programmed to connect with the intelligent infusion pump, the automatic programming parameters cannot match the pump-side drug library protocol and trigger a switch to manual or basic mode due to the lag in the update and activation time of the static rule base and the inconsistency of cross-system parameter semantics. Summary of the Invention

[0006] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a real-time medication control system based on reinforcement learning and a static rule base. It generates protocol fingerprints and forms candidate safety policy capsules. Handshakes and conflict classifications are performed based on the device capability list. Negotiable fields are converted, trimmed, and signed to generate the final safety policy capsule. After signature verification at the execution end, it enters a guardrail interconnection mode, performs initial settings and titration adjustments according to capsule constraints, and restricts the basic mode with no drug selection. In case of anomalies, a controlled emergency token is activated. Reinforcement learning is used to generate governance actions without changing hard constraints, thereby reducing matching failure and bypass risks and enhancing traceability and operational consistency; thus solving the technical problems described in the background art.

[0007] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: The real-time medication control system based on reinforcement learning and static rule base includes: at the trigger points of prescription signing, verification, or execution, collecting medication information from the prescription system, pharmacy preparation system, and static rule base, performing semantic normalization and unit conversion, distinguishing between negotiable and non-negotiable fields, generating protocol fingerprints, and forming candidate safety policy capsules; based on the candidate safety policy capsules and protocol fingerprints, obtaining the device capability list from the target execution end; converting and trimming negotiable fields according to the device capability list, determining conflicts of non-negotiable fields, and generating and signing the final safety policy capsule; After the target execution end verifies and approves the final security policy capsule, it enters the fenced interconnection mode. The final security policy capsule's constraints are set and adjusted, and the basic mode startup path without drug selection is restricted. In case of an anomaly, the controlled emergency token is activated and the cause code is recorded. The protocol fingerprint, device capability list, conflict determination result, final security policy capsule, and execution record are written into the audit link. Based on the audit link and without changing the static rule base hard constraints, reinforcement learning is used to generate governance actions and write them back.

[0008] Furthermore, a session identifier is established based on the medical order trigger event and the source version mark is recorded. Fields with the same name in the medical order system, pharmacy preparation system and static rule base are sealed. Subsequent changes are written into the audit link and associated with the session identifier. A consistency check code is generated for the sealed field and written into the candidate security policy capsule.

[0009] Furthermore, a dual-path mapping and cross-checking of the coding path and the name preparation attribute path are performed on the drug identifier to obtain a unified identifier for the drug entity. The route of administration and the preparation constraints of the pharmacy preparation system are then bound and written into the candidate safety strategy capsule. When the results of the two paths are inconsistent, the unified identifier for the drug entity is included in the non-negotiable field.

[0010] Furthermore, the dose and rate expression fields, standard concentration and weight-related fields are converted to a unified metrology system, and the values ​​before and after pruning are generated according to the rounding and resolution rules of the static rule base. Then, the key fields are spliced ​​in a fixed field order to generate a protocol fingerprint, and the hard limit, soft limit, titration step size and minimum adjustment interval are assembled into a candidate safety strategy capsule.

[0011] Furthermore, the device capability list includes unit set, minimum resolution, available infusion modes, and hardware limits, and is bound to session identifiers and protocol fingerprints; the constraints of the static rule base and the device capability list are merged by intersection with stricter boundary rules as a source of constraints for the conversion and pruning of negotiable fields.

[0012] Furthermore, a consistency check is first performed on non-negotiable fields, and if the check is not met, a non-negotiable conflict record is generated, written into the audit chain, and triggers pharmacist review. After the consistency check is passed, unit conversion, rounding, and resolution pruning are performed on negotiable fields, and pruning traces are formed and written into the final security policy capsule.

[0013] Furthermore, after the target execution terminal passes the verification, it loads the final security strategy capsule and enters the fenced interconnection mode state machine. During the initial setup, it verifies the unit and hard limit boundaries. During the adjustment, it adjusts the entry point by gating at the minimum adjustment interval. It also transfers the operation of entering the basic mode without drug selection to the controlled emergency token process and records the event.

[0014] Furthermore, the controlled emergency token is a one-time, short-term, and verifiable data object, which includes minimum bottom-line guardrails for allowed unit sets and global maximum and minimum rates; before enabling the controlled emergency token, the target execution end is required to enter a reason code and complete double-person verification of credentials, and then set and adjust the minimum bottom-line guardrail restrictions.

[0015] Furthermore, the audit event sequence is appended only, consisting of session identifier merging protocol fingerprint, device capability list, conflict record, final security policy capsule, execution record, and controlled emergency token event. The audit event sequence is then deterministically serialized according to a fixed field order, and an audit chain summary is calculated and written into the audit chain.

[0016] Furthermore, governance actions include adjusting the strategy push rhythm, setting the activation reminder intensity, and determining the review priority and work order dispatch order; reinforcement learning uses event statistics in the audit chain as input and output for governance actions, and performs hard constraint verification on governance actions before execution to ensure that governance actions do not change the hard constraints of the static rule base.

[0017] (III) Beneficial Effects This invention provides a real-time medication control system based on reinforcement learning and a static rule base, which has the following beneficial effects: Information from the medical order system, pharmacy preparation system, and static rule base is collected and semantically processed at the trigger points of medical order signing, verification, or execution. This generates protocol fingerprints and candidate security policy capsules, avoiding errors caused by homonyms or manual conversion, and ensuring unified reference objects and convenient review. Based on the candidate security policy capsules and protocol fingerprints, a device capability list is obtained and conflict classification and trimming are performed. Non-negotiable fields are blocked and reviewed, while negotiable fields are converted, resolution trimmed, and signed to generate the final security policy capsule, ensuring consistency between the distribution object and device capability boundaries. After the target execution end verifies and signs the final security policy capsule, it enters the guardrail interconnection mode. Gating is applied to initial settings, rate adjustment, titration according to hard and soft limits, titration step size, minimum adjustment interval, rounding, and resolution rules, restricting the basic mode startup path without drug selection and preventing bypass.

[0018] In cases of network, policy, signature verification, or device unavailability anomalies, controlled emergency tokens are activated, with reason codes and dual-verification credentials as activation conditions. Exceptional activation is still subject to minimum baseline constraints and is linked to session identifiers for logging. Protocol fingerprints, device capability lists, conflict determination results, final security policy capsules, execution records, and controlled emergency token events are written into the audit chain to form replayable evidence, enabling problem localization to input semantics, device capabilities, rule clauses, or execution stages to support rule base governance. Without changing the static hard constraints of the rule base, the learning-based audit chain generates governance actions and writes them back, enabling collaborative orchestration of maintenance windows, activation reminders, audit priorities, and work order dispatch. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the structure of the real-time medication control system based on reinforcement learning and static rule base of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] Please see Figure 1 This invention provides a real-time medication control system based on reinforcement learning and a static rule base, comprising: Step 1: At the medical order trigger point, converge the medication information from the medical order system, pharmacy preparation system and static rule base into the same set of standardized expressions, and generate protocol fingerprints and candidate security policy capsules, so that the next consistency handshake can directly reference them without reinterpreting the field semantics.

[0022] Interconnected programming failures are often not caused by a single system malfunction, but by deviations in the expression of medical orders, preparation, and equipment drug library within the same time slice. Therefore, it is necessary to capture the original information from the three sources at the trigger point and lock their versions and contexts, and then establish a unique mapping between the drug entities in the three sources, thereby providing stable input for subsequent unit conversion, rounding, and limit projection.

[0023] The process follows the order of first sealing, then standardizing, and then unifying: first, the original message of the trigger point is sealed and verified, then the drug entity is mapped in two paths, and finally the mapping result is written into the drug entity field and the applicable scenario field of the candidate security policy capsule.

[0024] Starting with observable actions in the clinical setting: when a physician clicks to sign a medical order, a nurse confirms and verifies it on the execution interface, or an interconnected programming request is initiated on the execution interface, the parameters are not immediately inferred. Instead, the content seen at that moment is first sealed as original evidence in a deterministic manner, thereby avoiding subsequent misinterpretations due to dictionary updates or interface refreshes. This can be compared to taking a picture of a prescription and stamping it with a timestamp, but in this solution, the sealed object is not an image, but a computable set of structured fields.

[0025] A one-time session identifier is generated based on the trigger point, and three types of data are retrieved within this session: first, drug identifier, route of administration, dosage or rate expression, and intended concentration or preparation information returned by the medical order system; second, standard concentration, solvent type, preparation volume, and allowed alternatives returned by the pharmacy preparation system; and third, hard restrictions, soft restrictions, titration step size, minimum adjustment interval, rounding and resolution rules, and applicable nursing unit strategies returned by the static rule base. To avoid the same field being overwritten multiple times, a write-only rule is adopted for field sealing: only one write is allowed from the same source within the same session, and subsequent changes to fields with the same name are recorded as subsidiary change records without overwriting the original value, thus ensuring that the original state of the trigger point can be reproduced in subsequent audits.

[0026] The process employs two sequential steps: First, it performs structural verification and version marking on each type of source data. The version marking includes the source system dictionary version number, static rule base version number, and pharmacy preparation specification version number. Second, it generates a consistency checksum for the sealed field set and writes it into the session header of the candidate security policy capsule. This checksum is used to detect whether any fields are missing or their order has changed during transmission. If an inconsistency in the checksum is found, the session enters a state where manual review is allowed only, and the process does not proceed to the next step of trimming and submission.

[0027] The session identifier is obtained by deterministically serializing the medical order identifier, nursing unit identifier, target device identifier, trigger point timestamp, and patient token using an anti-collision mapping operator, avoiding purely random values ​​that would prevent cross-system verification. The consistency check code covers three source fields (medical order field, pharmacy field, and static rule base field) and their version markers, using a fixed field order and fixed encoding to avoid false differences caused by field rearrangement. A session switching mechanism is defined for key field changes: within the same session, if a set of non-negotiable fields changes (e.g., drug entity identifier, route of administration, target concentration, hard restriction clause identifier), the system closes the current session and creates a new one. The old session is retained as an audit record and marked with the trigger point change.

[0028] The process involves sealing the three source fields of the trigger point and adding a version marker, obtaining a consistency check code which is written to the header and cannot overwrite the original value. The set of trigger point fields can still be reproduced in subsequent processing links, and semantic shifts will not occur due to changes in the dictionary. The version marker and check code can determine in subsequent steps that the difference originates from the input rather than the calculation difference, facilitating the location of the conflict mechanism. The focus is on the uniqueness of drug entity identifiers because, in interconnected scenarios, the medical order system may use the generic name or hospital code of the drug, the pharmacy preparation system may use the formulation specification or packaging code, and the static rule base may use the pump-end drug database entry name or nursing unit configuration name. If these three are not aligned, even after correct unit conversions, irreconcilable conflicts will still occur due to incorrect entry selection.

[0029] A dual-path mapping and cross-checking approach is adopted: the first path prioritizes coding, using a mapping table between hospital drug codes and pharmacy packaging codes for direct connection; the second path focuses on name and formulation attributes, constructing matching keys using a five-tuple of generic name, salt type, dosage form, strength, and route of administration, and searching for entries with the same key in a static rule base. The results of both paths are required to be consistent at the drug entity level; otherwise, they are marked as non-negotiable fields and frozen in the candidate safety strategy capsule for direct reference in the next conflict classification step. To achieve this matching, the engineering implementation can use a minimum-cost assignment solver to globally minimize candidate matching pairs, or a weighted trie search to first narrow down the candidate set before making local assignments, ensuring stable landing even when drug names have abbreviations, aliases, or space differences.

[0030] It also consists of two consecutive processing actions: First, the pharmacy preparation constraints are written into the candidate safety strategy capsule as constraints of the drug entity mapping, such as solvent type, preparation volume, allowable concentration set and prohibited concentration set; Second, the mapped drug entity identifier and administration route are used together for static rule base entry screening, so that entries with the same name but different routes will not be mistakenly selected into the candidate safety strategy capsule.

[0031] As an example, in the intensive care unit, a physician prescribes a continuous norepinephrine infusion in the medical order system, with the order interface displaying the unit as micrograms per kilogram per minute. The pharmacy preparation system returns the hospital's standard concentration for this drug as a fixed milligram per milliliter, specifying sodium chloride injection as the solvent and a fixed preparation volume. When the nurse clicks on the interconnected programming interface, they first save the medical order and pharmacy preparation expressions, then confirm the consistency of the drug entity identifier through the coding priority path, and then confirm the consistency of the salt type and specification through the name and preparation attribute path. Finally, they write the solvent type and preparation volume into the preparation constraint field of the candidate safety strategy capsule. The visible result is that in subsequent interfaces, the medical order is marked as having achieved drug entity consistency, and a traceable record is formed in the candidate safety strategy capsule, which can be directly referenced in the next step of customization without requiring the nurse to repeatedly select the concentration or solvent.

[0032] First, a unique drug entity identifier is generated through dual-path mapping. Then, pharmacy preparation constraints and routes of administration are bound and written into the candidate safety strategy capsule. This ensures that the drug entity is uniquely mapped among the three data sources, avoiding irreconcilable conflicts caused by different names. Preparation constraints are solidified in the candidate safety strategy capsule for subsequent concentration and unit conversions within the same semantic space.

[0033] Matching failures in interconnected programming often occur at the unit and resolution level: medical order systems express values ​​in weight-related units, pharmacy preparations in concentration and volume, and devices in milliliters per hour or mass per hour. If rounding and resolution are not standardized during conversion, seemingly similar values ​​may appear to be unequal according to the protocol, triggering unpredictable alarm paths. Therefore, it is necessary to project all computable fields onto a measurement system recognized by a static rule base, clearly distinguishing between negotiable and non-negotiable fields during the projection process, thereby generating a protocol fingerprint and assembling it into the main body of the candidate security policy capsule.

[0034] The main idea is to express the same thing using the same unit, but it is not just about converting the unit to a certain benchmark unit. More importantly, it is about constraining the converted result to the rounding and resolution rules recognized by the static rule base, so that subsequent consistency judgments are not affected by display precision or device step precision.

[0035] Since the weight in the medical order system may come from bedside weighing or nursing records, the sampling is not continuous. Time alignment of weight is allowed at the trigger point: when the weight record time is earlier than the trigger point and later than the handover time of the previous nursing shift, the most recent weight record is taken; when the trigger point is between two weight records and both records are valid, piecewise linear interpolation is used to calculate the weight at the trigger point to ensure that the conversion of weight-related units has a definite value at the trigger point. The selection of the interpolation endpoint and the validity judgment are written into the audit log for restoration during review.

[0036] This is achieved using a constrained projection operator: first, the original parameter set is converted into a normalized parameter set; then, the normalized parameter set is projected onto the feasible region defined by the hard constraints and rounding rules of the static rule base, thus obtaining a subsequently comparable parameter set. Specifically, this is expressed as:

[0037] Where: original parameter set : An ordered parameter set concatenated from dosage or rate expressions in the medical order system, concentration expressions in pharmacy preparation, route of administration, weight information, and static rule base entries; values ​​are limited by the data types and validity rules of each source system, including numerical fields and enumerated fields; a normalized parameter set. : An ordered set of parameters obtained after unit conversion, rounding, resolution clipping, and hard constraint truncation; the values ​​are restricted to the measurement system and hard constraints recognized by the static rule base; rule constraint set The static rule base contains a set of rules that match the drug entity identifier, route of administration, and nursing unit strategy, including at least hard constraints, soft constraints, titration step size, minimum adjustment interval, rounding rules, and resolution rules; the value range is a discrete configuration set; and the drug entity mapping operator. The input consists of a set of drug fields for medical orders, a set of drug fields for pharmacies, and a set of drug fields for static rule bases. The output is a unique drug entity identifier and mapping evidence.

[0038] The cost composition for minimum cost assignment: The cost is a weighted sum of the costs for code consistency, generic name similarity, salt type consistency, dosage form consistency, specification consistency, and route of administration consistency. If the codes are completely identical, the cost is zero and the assignment is accepted directly; otherwise, the path between name and formulation attributes is entered. When multiple candidates with equal costs exist, the candidate with both consistent route of administration and specification is selected first. Mapping failure output: When the results of the two paths are inconsistent or there are no candidates, Output a mapping failure flag and write the field to the non-negotiable field set so that step two can directly trigger a non-negotiable conflict.

[0039] Constrained projection operator : Deterministic mapping operator, specifically in the form of a composite mapping that sequentially performs unit conversion mapping, rounding mapping, resolution clipping mapping, and hard-limit truncation mapping; its output remains unique for the same input and the same input. The process includes two interconnected steps: First, unit conversion uses rational number conversion and explicitly records the conversion path, for example, mapping from mass per weight per time to mass per time via weight, and then to volume per time via concentration. Second, rounding and resolution are uniformly constrained by rounding rules in a static rule base and the minimum step size of the device. Candidate values ​​are first obtained according to the rounding rules, and then the executable values ​​are obtained by pruning according to the resolution. Both the pre-pruning and post-pruning values ​​are written into the audit field of the candidate security policy capsule, so that the next step can determine whether the difference belongs to a negotiable pruning or an unnegotiable conflict. In an equivalent implementation, weight alignment can be changed to fixed weight based on nursing shifts or weight update triggered by electronic scale events; the unit conversion path can be implemented by pre-compiled table lookup or automatically generated by the symbol converter according to the unit dimension.

[0040] The constraint projection operator generates a standardized parameter set, which is explicitly written into the conversion link. Rounding and resolution constraints are written before and after pruning. All computable fields are unified into a static rule base using a measurement system. Subsequent consistency judgments do not require the display of precision on the interface. The conversion link and the values ​​before and after pruning are entered into the audit field of the candidate security policy capsule, making it easy to distinguish whether the difference mechanism belongs to the unit link difference or the resolution pruning difference.

[0041] Furthermore, the normalized parameter set is transformed into a comparable, traceable, and transferable protocol fingerprint, and the same normalization result is assembled into the text of a candidate security policy capsule, enabling the next step to simultaneously complete conflict classification and policy tailoring based on the same facts. Unlike simply recording a few strings, protocol fingerprints must be sensitive to field order, units, rounding rules, and applicable scenarios; otherwise, audit gaps will arise where fingerprints are identical but protocols differ.

[0042] A combination of deterministic serialization operators and collision-resistant mapping operators is used: first, the normalized parameter group is serialized into a byte sequence according to a fixed field order, and then the protocol fingerprint is calculated from the byte sequence. This can be expressed as:

[0043] In the formula: protocol fingerprint : A fixed-length digest output by the collision-resistant mapping operator; its value range is a finite bit string space; its function is to serve as one of the primary keys for subsequent consistency judgment, audit association, and candidate security policy capsule index; Serialization operator : A deterministic operator that maps byte sequences according to predefined field order, field type encoding, length prefix and delimiter rules; the value range is a finite set of byte sequences; its function is to eliminate differences in field arrangement and encoding, so that different execution ends produce the same byte sequence when they receive the same input.

[0044] Collision-resistant mapping operator This operator maps byte sequences to fixed-length digests, satisfying collision resistance requirements and enabling integrity verification; its value range is the fixed-length digest space; its function is to generate protocol fingerprints and support subsequent verification and traceability. (Normalized parameter set) The meaning is the same as described above; its function is to serve as the sole source of factual information for the protocol fingerprint.

[0045] The process involves two interconnected steps: First, the candidate security policy capsule is assembled using a header-body-audit tail structure. The header contains the session identifier, source version marker, and consistency check code; the body contains hard limits, soft limits, titration step size, minimum adjustment interval, rounding and resolution rules, and configuration constraints; and the audit tail contains the conversion link and values ​​before and after pruning. Second, after assembly, the protocol fingerprint is written to the header of the candidate security policy capsule, and the list of non-negotiable fields is written to the field partition section of the body, enabling the next step to directly reference non-negotiable fields without re-parsing the three source fields.

[0046] The process involves serializing and normalizing parameter groups in a fixed field order to generate a protocol fingerprint. Candidate security policy capsules are then assembled and written to field partitions, following a header-body-audit tail sequence. The protocol fingerprint relies on field order, units, rounding, and applicable scenarios to support subsequent consistency checks and audit relationships. Candidate security policy capsules can store both the normalized body and the audit tail, allowing for further trimming and submission on the same object.

[0047] Step 2: Using the protocol fingerprint and candidate security policy capsules generated in Step 1 as the sole proposal input, and combining them with the capability list of the target execution end, complete the conflict classification and policy tailoring, and generate a final security policy capsule that can be verified for direct reference in Step 3.

[0048] Although candidate security policy capsules have completed the metrological system normalization and hard constraint projection according to the static rule base, they may still encounter situations such as insufficient device resolution, limited available infusion modes, lagging device-side rule base versions, or being in an inactive state when implemented in interconnected systems. If security policy capsules are directly trimmed and issued before the device capability boundaries are clearly defined, it will be difficult to distinguish between the two mechanisms: the policy does not meet the device capability and the medical order or preparation expression is inconsistent. Therefore, it is necessary to first form an auditable constraint set of the capability status of the target execution end at the trigger point, then bind this constraint set with the session encapsulation information, and then proceed to the next step of conflict classification.

[0049] Starting with the first executability gate in the execution chain: After receiving a candidate security policy capsule, the system does not immediately generate the final security policy capsule. Instead, it first sends a capability list request to the target execution end, requesting the target execution end to return its supported unit of measurement, minimum resolution, minimum step size, available input mode set, hardware limit boundaries, and a summary of the current device-side rule base version. The capability list request and response use a defined field order and field type, ensuring that it can enter the audit chain like the sealed fields in step one, avoiding disputes arising from different meanings of the same field due to device firmware upgrades or configuration changes. The device's minimum resolution and minimum step size directly affect rate pruning. Without this information, negotiable differences may be misjudged as non-negotiable conflicts, triggering unnecessary blocking and manual review.

[0050] The capability list undergoes a two-stage normalization process: first, the unit set is mapped to the unified metrology system used in step one, and the available infusion patterns are mapped to the path and pattern enumeration in the static rule base; then, the hardware limit boundaries and minimum resolution are written into the equipment capability constraint clauses, which are expressed as a quadruple of field name-allowed set-boundary-rounding rule, facilitating reference to each clause during subsequent trimming. In the engineering implementation path, the capability list can be obtained directly through the infusion pump's interconnection interface or forwarded through the infusion pump's centralized management system; when forwarded through the centralized management system, the system adds a forwarding link identifier to the returned capability list to distinguish between direct equipment connection and forwarded acquisition during auditing.

[0051] First, a capability list request is sent to the target execution end, and the return values ​​are saved sequentially according to the field order. Then, the unit set, input mode, and hardware boundary are written into the session record as device capability constraints, and the resolution, step size, and mode limit are solidified as verifiable constraints. After obtaining the capability list, the device state at the trigger point is remembered when inconsistencies occur.

[0052] Furthermore, attention should be paid to the concept of the same time slice, because in a clinical setting, a nurse may initiate an interconnect programming request on the execution interface before connecting to an infusion pump, or change the infusion pump during the connection process. If the candidate safety policy capsule and the capability list come from different devices or different time slices, the subsequent tailoring results will lack interpretability. Step three relies on the final safety policy capsule to control the initial parameters and adjustment behavior; therefore, it must be ensured that the final safety policy capsule is obtained by closing the gap between the candidate safety policy capsule, the protocol fingerprint, and the device capability list within the same session.

[0053] Using the session identifier stored in step one as the primary key, the capability list and candidate security policy capsules are bidirectionally bound: on the one hand, the capability list summary is written to the audit tail of the candidate security policy capsule, so that the candidate security policy capsule carries the reference of the device capability at that time in subsequent transmission; on the other hand, the protocol fingerprint of the candidate security policy capsule is written to the capability list record, so that the capability list clearly serves which proposal in the audit link.

[0054] After binding is complete, the system performs a proposal consistency check: It verifies that the consistency check code in the candidate security policy capsule header matches the consistency check code archived in the current session; it verifies that the list of non-negotiable fields in the candidate security policy capsule has not been rewritten; and it verifies that there are no contradictions between the drug entity identifier and the route of administration in the infusion mode mapping results between the candidate security policy capsule and the capability list. If any check fails, the system marks the session as allowing only manual review and stops proceeding to the next step. Specifically, the capability list and candidate security policy capsule are first bound using the session identifier. Proposal consistency checks are then performed on the consistency check code, the list of non-negotiable fields, and the infusion mode mapping results. The candidate security policy capsule and the capability list are set on the same time slice to avoid errors such as using old capabilities to trim new proposals or using the wrong device to trim proposals. Consistency check failures are intercepted before trimming, and subsequent conflict classification focuses on actual differences rather than transmission or binding errors.

[0055] Differences in interconnected scenarios do not all carry the same risk level: inconsistencies in drug entity identification constitute non-negotiable conflicts, while unit differences can often be resolved through conversion and rounding. Indiscriminate blocking without grading would lead to numerous unnecessary interruptions; conversely, indiscriminate allowing could introduce critical conflicts into the fenced interconnection mode of step three. Therefore, a deterministic link needs to be established, from candidate safety policy capsules to device capability constraints, static rule base clauses, conflict grading, submission, and audit writing, ensuring that each difference has a verifiable mechanism and outcome.

[0056] The system organizes its decisions by first determining non-negotiable fields and then negotiable fields, avoiding misleading decisions made before confirming consistency of drug entities. The list of non-negotiable fields is already written into the candidate safety strategy capsule in step one. If this list is not prioritized in this step, it weakens the significance of the field partitioning in step one and makes it difficult to pinpoint responsibility in subsequent audits. The system first performs a consistency check on the non-negotiable fields: the drug entity identifier must be consistent, the route of administration must fall within the set of infusion modes allowed by the device capability constraints, and the set of allowed concentrations in the pharmacy preparation constraints must cover the target concentration expression in the candidate safety strategy capsule text. If any condition is not met, the system generates a non-negotiable conflict record, specifying the conflicting field, candidate value, allowed set or boundary of device capabilities, and a clause reference from the static rule base, and then forwards the session to the pharmacist or prescription-side reviewer.

[0057] After all non-negotiable fields have passed, the system proceeds to determine negotiable differences: For unit differences, the system uses the unit conversion link from step one to map the medical order expression and the device's executable units to a unified measurement system; for resolution differences, the system compares the pre-pruning and post-pruning values ​​in the audit tail of the candidate safety strategy capsule and further prunes them again according to the device's minimum step size; for soft constraint differences, the system allows the soft constraint prompt text and threshold to be folded and written according to the nursing unit strategy without breaking the hard constraints. Negotiable difference determination can be performed either by using a deterministic rule engine to judge each condition individually, or by using a constraint satisfaction solver to model the selection of negotiable field values ​​as a constraint satisfaction problem and obtain a set of solutions that satisfy all boundaries. The solver can be a mixed-integer linear programming solver or a backtracking search solver, and constraint relaxation during the solution process must be written into the audit log.

[0058] If the unit set or minimum step size is missing: This session is prohibited from entering the guardrail interconnection mode, directly enters the non-negotiable conflict path, and requires manual review. If the rule base version summary is missing: Continue is allowed, but the device version is marked as unknown at the end of the audit, and its priority is increased when generating the maintenance task order in step four. If the available infusion mode is missing: The administration route is set as a non-negotiable field, blocking the process and pushing it up for review.

[0059] The process begins by reviewing the list of non-negotiable fields as non-negotiable conflict records. Then, negotiable differences in units, resolution, and soft constraints are assessed, and constraint references and relaxation traces are recorded. Non-negotiable conflicts are blocked before pruning, preventing critical contradictions from entering the execution chain and forming the basis for safeguards. Negotiable differences are controlled within hard constraint boundaries, ensuring that subsequent pruning has interpretable boundary conditions rather than arbitrary pruning. Negotiable differences are decomposed into executable parameter sets, which are then assembled into a final security policy capsule for direct use in step three.

[0060] The final security policy capsule must not only include hard limits and titration constraints, but also carry session-related security commit fields to ensure its one-time nature and traceability in the event of network retries, device switching, or operator misoperation. Static rule base clauses and device capability constraint clauses are merged into a composite rule constraint set. Under this composite rule constraint set, the normalized parameter set output from step one is subjected to a second constraint projection to obtain an executable parameter set. This projection uses the constraint projection operator defined in step one to maintain consistency between notation and mechanism, expressed as:

[0061] Where: Executable parameter set : An ordered set of parameters output by the constraint projection operator under a composite rule constraint set, which includes at least the target concentration expression, target unit expression, executable start rate, hard limit boundary, soft limit threshold, titration step size and minimum adjustment interval; the value range is jointly limited by the static rule base and the device capability; its function is to serve as the sole source of numerical facts for the final security strategy capsule text.

[0062] Constrained projection operator : Deterministic composite mapping operator, specifically in the form of sequentially performing unit conversion mapping, rounding mapping, resolution clipping mapping, and hard constraint truncation mapping, and referencing the corresponding clauses in the composite rule constraint set at each step; the value range is the mapping set from the normalized parameter set to the executable parameter set; For any numerical field whose projection needs to be constrained (e.g., rate, dose, concentration), its original value is defined as the original numerical value. Its target unit is a numerical target unit. The original unit is the numerical original unit. The unit conversion factor is the conversion factor. The lower bound given by a static rule base or a composite rule constraint set is the lower bound. The upper boundary is the upper boundary. The device step size is the step size. The rounding scale is the rounding scale. The constraint projection is then performed in the following deterministic order:

[0063] Where: original numerical value : Values ​​input from medical orders or execution terminals; value range defined by the input interface or upstream system; function as the projection starting point; conversion factor. : Positive rational number; range of values ​​is Its function is to convert the original unit into the target unit. Numerical target unit Numerical units : Unit enumeration value; the value range is the set of units in the unified measurement system; its function is to determine the conversion dimension.

[0064] To avoid inconsistencies caused by different values ​​for the same result, it is recommended to provide at least one feasible rounding method. For example, rounding to the nearest value by the scale increment, and rounding down by half a scale increment:

[0065] In the formula: rounding scale : Positive number; range of values ​​is Its function is to determine the rounding granularity (e.g., ); : Round down; a well-known mathematical operation; its function is to achieve deterministic rounding. Resolution cropping (projecting onto the executable grid according to the device step size, defaulting to the largest grid point not exceeding the rounding value):

[0066] In the formula: step size : Positive number; range of values ​​is ; Derived from the capability list; Its function is to project numerical values ​​onto device-executable grid points. Hard constraint truncation (clamping numerical values ​​within rule boundaries):

[0067] In the formula: lower bound : Numerical value; the range of values ​​is defined by a static rule base or a composite rule constraint set; its function is to hard limit the lower bound. Upper bound : Numerical value; the range of values ​​is defined by the static rule base or composite rule constraint set; its function is to hard limit the upper limit. , Commonly known mathematical operations are used to achieve truncation.

[0068] Processing of parameter sets: Each numerical component of the normalized parameter set is obtained according to the above process. Enumerated fields (e.g., route of administration, infusion mode) only undergo set inclusion checks, while set fields (e.g., allowed unit sets) are written to the object after deterministic sorting. When using composite rule constraint sets, , , , All clauses are taken from the corresponding clauses of the composite rule constraint set; the rest of the form remains unchanged, thus ensuring that the same operator definition is used in step two (trimming), step three (adjustment), and step four (replay).

[0069] Composite rule constraint set The rule set is a combination of static rule base clauses and device capability constraint clauses, and includes at least hard limits, soft limits, titration step size, minimum adjustment interval, rounding rules, resolution rules, available infusion mode set and hardware limit boundaries; the value range is a discrete clause set; its function is to define the executability boundary and support the determinism of pruning.

[0070] Normalized parameter set The normalized parameter set output from step one; its value range is restricted by the static rule base clauses; its function is to uniformly express the input before pruning.

[0071] After obtaining the executable parameter set, the system assembles the final security policy capsule according to the structure of header-body-audit tail: the header contains the protocol fingerprint, session identifier, target device identifier, source version mark, effective and expiration time, one-time random number, maximum number of uses, and revocation version number; to avoid drift in the determination of effective and expiration time due to the local clock deviation of the execution end, the system writes the trigger point timestamp when the session is sealed, and writes the trigger point timestamp and expiration timestamp in the header of the final security policy capsule at the same time; when verifying the signature, the execution end prioritizes the use of the trigger point timestamp and its own monotonic counter to jointly determine whether the object is available, thereby reducing misjudgment caused by time synchronization switching or power failure reset.

[0072] The main body of the audit file contains hard constraints, soft constraints, and titration constraints determined by the executable parameter set. The audit tail contains conflict classification results, negotiable difference pruning traces, proof of non-negotiable field pass, and a capability list summary. The system then digitally signs the final security policy capsule and writes the signature and summary into the audit link. After successful signature verification at the execution end, a confirmation receipt is returned. This receipt must contain the protocol fingerprint and session identifier, enabling the system to correlate the receipt with the final security policy capsule. In an alternative implementation, the digital signature can be generated by a hardware cryptographic device or a protected key store; the revocation version number can be provided by a centralized revocation list or updated by the nursing unit policy during shift changes.

[0073] For numerical fields that need to be stepped (such as rate), first obtain the rounded value according to the rounding rules, and then perform step projection according to the device's minimum step size; when the rounded value is not on the step grid, the default selection is to not exceed the maximum executable step point of the original input, so as to avoid exceeding the hard limit due to rounding up; if this selection causes the value to be lower than the soft limit lower bound, a second confirmation is prompted and recorded.

[0074] Specifically, the normalized parameter set is constrained and projected under a composite rule constraint set to obtain an executable parameter set. This executable parameter set is then assembled into the final security policy capsule, and the signature and receipt binding are completed and written to the audit link. The executable parameter set is uniquely determined at the boundary between the static rule base and the device capabilities, ensuring that the final security policy capsule is executable and verifiable. The final security policy capsule carries one-time and time-limited fields and is protected by signature, preventing it from being reused during transmission and retries.

[0075] Step 3: The final safety strategy capsule is verified, loaded, and used as a reference for initial setup and subsequent adjustments of the safety barriers at the target execution end. In abnormal situations, a controlled emergency token maintains the minimum safety barrier and a traceable chain of evidence. The final safety strategy capsule has already been trimmed and signed in Step 2. If consistency confirmation and state switching are not completed at the execution end, nursing staff will still enter the basic mode without medication selection, rendering the trimming results of Step 2 ineffective. Therefore, it is necessary to first complete verification, session binding, and state initiation at the execution end through a deterministic process, and then load hard constraints, soft constraints, and titration constraints as executable verification rules to open the infusion initiation and adjustment entry points.

[0076] The prerequisite of object-oriented trust is elaborated as follows: After receiving the final security policy capsule, the execution end does not directly enter the running interface, but first completes the digital signature verification in the local signature verifier and processes the protocol fingerprint in the capsule header. The session identifier, target device identifier, and revocation version number are checked item by item to ensure that the capsule was indeed submitted with the signature in step two and has not been replaced. Infusion devices operate in environments subject to power outages, movement, and network fluctuations. Allowing the same capsule to be reused or reused across patients would cause the guardrail to become disconnected from the medical order object. Therefore, the one-time random number and maximum number of uses must be implemented as hard constraints on the execution end, rather than remaining merely in field descriptions.

[0077] The execution end first verifies the signature using the public key material stored in the hardware security storage, and simultaneously performs a local revocation table match on the revocation version number; then it reads a one-time random number and constructs a usage trajectory record together with a local monotonic counter. Once the same one-time random number is detected to appear a second time on the same device, the capsule is marked as unusable and written to the audit event.

[0078] To reduce the failure judgment bias caused by local clock discontinuity, the failure time determination adopts a dual-condition approach: trigger point timestamp + monotonic counter. The trigger point timestamp is written from the header of step two, and the monotonic counter is maintained by the execution end when it is powered off. If either condition is not met, the capsule is determined to be unusable. In terms of engineering implementation, the signature verification device can be built into the infusion pump firmware or deployed in the interconnect adapter. Hardware secure storage can use a protected key area or an external cryptographic device, and the monotonic counter can use a rollback-proof counter register or a persistent counter file. When the execution end lacks a monotonic counter, the interconnect adapter maintains a session counter and writes the session counter value into the acknowledgment and audit event sequence. The execution end must send back the session counter each time it receives a capsule to prevent reuse.

[0079] When there is neither a monotonic counter nor an interconnect adapter, the system performs deduplication on the server side using the maximum number of uses + a one-time random number + a session identifier. The execution end only serves for display and gating. Specifically, digital signature verification and protocol fingerprint checking are first performed on the execution end. The system uses a session identifier and revocation version number, then combines a one-time random number with a local monotonic counter to create a usage trajectory and reject duplicate capsules. The final security policy ensures that capsules have verifiable origins and are non-replaceable at the execution end, maintaining consistency between the guardrail criteria and the submission object in step two. The one-time random number and monotonic counter together restrict reuse behavior, preventing capsules from repeatedly entering the execution chain across patients or sessions.

[0080] For guardrail-oriented systems, the following implementation condition applies: After verification, the execution end needs to load the hard limits, soft limits, titration step size, minimum adjustment interval, rounding, and resolution rules from the final safety strategy capsule as runtime verification rules, and set the entry point for gating the initial parameters accordingly. In clinical settings, a common initiation action is for the nurse to first input or confirm the initial rate, and then adjust it sequentially based on the patient's response; if the initial entry point is not gated, subsequent adjustments, even with restrictions, cannot avoid errors from the initial input.

[0081] The execution end establishes a state machine for the guardrail interconnected mode, which includes at least six states: pending loading, pending startup, running, adjusting, paused, and terminated. In the pending startup state, the execution end performs three checks on the initial settings: first, it checks if the unit falls within the capsule's allowed unit set; second, it checks if the starting rate falls within the hard constraint boundary; and third, it checks if the starting rate is compatible with the titration step size and resolution rules. If incompatible, the starting rate is projected to the nearest executable step point and the nurse is prompted for confirmation. In the running and adjusting states, the execution end adds a minimum adjustment interval gating to each adjustment: before the interval is reached, the adjustment entry remains unavailable and the attempt event is recorded for later review, thus ensuring the titration rhythm aligns with capsule constraints. For basic mode entries without drug selection, no direct entry is provided by default in the guardrail interconnected mode; when a nurse attempts to enter this entry, the execution end switches to the emergency entry judgment branch and requires a controlled emergency token or dual-verification credentials to prevent starting without guardrails from becoming the default path. In terms of engineering implementation, the state machine can be implemented in the pump-end firmware, or it can be maintained by the interconnect adapter and constrained by the pump-end input interface through interconnect commands; the step point projection can be implemented by looking up a table or by using integer step operators; the double-person verification of credentials can be implemented by combining employee ID card swiping and independent confirmation key.

[0082] The process begins by loading the capsule text rules into the state machine of the guardrail interconnected mode and verifying the execution unit, hard limits, and step compatibility of the initial settings. Then, a minimum adjustment interval gating is applied to the adjustment entry, and the entry in the basic mode without drug selection is restricted. The initial settings are gating before entering the run and aligned with the capsule hard limits and resolution rules to reduce the likelihood of initial input deviating from the guardrail. The adjustment entry is opened under time interval and step constraints to ensure that the titration rhythm is consistent with the capsule titration constraints and provides traceable records.

[0083] During the execution process, network unavailability, equipment inaccessibility, expired capsules, or failed verification may occur. Directly blocking the infusion start in such situations would interrupt nursing operations; allowing basic mode start directly would disable the safety barriers and reduce audit traceability. Therefore, during the operation and adjustment phases, it is necessary to execute the same constraint projection logic as in step two for each parameter change, and to enable controlled emergency tokens when an anomaly is triggered, so that the minimum bottom-line safety barriers still participate in the execution. The reason code, token trajectory, and manual verification action should be written into the audit link for subsequent governance in step four.

[0084] To ensure each change is subject to the same boundary constraints: In the fence interconnection mode, when a nurse increases, decreases, or pauses and restarts the rate, the execution end does not directly write the input value into the running parameters. Instead, it combines the input value with the session archive field, capsule text constraints, and device resolution to form a new normalized parameter set. Then, it calls the constraint projection operator to obtain the executable parameter set, thus maintaining a calculation path consistent with the trimming submission in step two. Clinical titration often occurs after changes in patient monitoring indicators. The frequency and magnitude of adjustments cannot be pre-enumerated. If only the initial settings are validated once, subsequent adjustments may still exceed hard limits or violate the minimum adjustment interval, rendering the fence ineffective.

[0085] The execution end reuses the metrology system normalization rule from step one during each adjustment: converting the input rate expression to a unified metrology system and forming a standardized parameter set according to the capsule rounding rule. Subsequently, constraint projection is performed under the composite rule constraint set to obtain the executable parameter set. The runtime parameters are replaced with the executable rate specified in the formula. This reuse process can be expressed as:

[0086] Where: normalized parameter set This is an ordered parameter group formed by concatenating the input and session archive fields with the capsule text rules and completing unit conversion and rounding; the value range is limited by the allowed range of the input interface and the validity rules of the session fields; its function is to convert the nurse's input into a calculable and unified expression.

[0087] Executable parameter group : An ordered set of parameters output by the constraint projection operator under a composite rule constraint set, which includes at least the executable rate, hard limit boundaries, and step information; the value range is limited by both the static rule base and the device capability; its function is to serve as the sole basis for replacing runtime parameters.

[0088] Composite rule constraint set : A set of rules obtained by merging static rule base clauses and device capability constraint clauses; the value range is a discrete set of clauses; its function is to define the feasible domain for each adjustment.

[0089] Constrained projection operator : A deterministic composite mapping operator, specifically, performs unit conversion mapping, rounding mapping, resolution clipping mapping and hard constraint truncation mapping in sequence; the value range is the set of mappings from to; the function is to converge any adjusted input to within the executable boundary.

[0090] For ease of understanding, this can be understood as the nurse proposing an adjustment value, which is then converted according to predetermined boundaries into a value that the device can execute without exceeding the limits. When the projection causes the executable rate to differ from the input rate, the execution terminal displays both the input and executable rates on the interface and requires the nurse to confirm again; this confirmation action is recorded in the audit event. When the projection triggers a hard limit truncation, the execution terminal marks this event as a hard limit trigger and locks any further inputs until the nurse chooses to reduce the rate or end the infusion. In engineering implementation, unit conversion and rounding can use pre-compiled conversion tables and a fixed decimal strategy; the projection operator can be implemented using integer arithmetic to avoid floating-point errors; and secondary confirmation can be achieved through a combination of a confirmation button and card swiping.

[0091] Each adjustment of the input involves a normalized set of parameters, which is then used to generate an executable set of parameters using the constraint projection operator. Replace the running parameters with the executable rate, and confirm and write the audit event when a difference occurs; each adjustment uses the same boundary and calculation path as in step two, and the guardrail constraint is applied to the entire process from startup to titration; the input rate and executable rate are displayed in parallel for secondary confirmation, and the pruning behavior can be seen and traced by the operator.

[0092] The controlled emergency token consists of a session identifier, target device identifier, trigger point timestamp, one-time random number, short validity period, and minimum bottom line guard fields, and generates a digital signature using the same signature mechanism as in step two. The minimum bottom line guard field includes at least: allowed unit set, global maximum rate boundary, global minimum rate boundary, mandatory two-person verification flag, and mandatory reason code flag. Before entering the emergency path, the executor verifies the signature and checks the consistency between the session identifier and the target device identifier, then requires the operator to enter a structured reason code and complete the two-person verification credential input before opening the initial settings entry; after opening, the minimum bottom line guard is still used to verify the unit and truncate the boundary of the input value, and all truncation events are written to the audit events. The emergency token can be issued by printing a QR code at the nurse station workstation and having the nurse scan it for import, or by writing it through the near-field channel of the interconnect adapter, or by distributing it within the ward through a protected mobile terminal; the two-person verification credential can be obtained by two nurses swiping their cards and pressing independent confirmation keys, or by a combination of one nurse swiping their card and another nurse entering an independent password.

[0093] First, a controlled emergency token is generated and signed by binding the session identifier and the target device identifier, and the signature is verified at the execution end. Then, after the reason code and the two-person verification credentials are verified, a restricted entry point is opened, and verification and boundary truncation are performed using the minimum bottom-line guard unit. In abnormal situations, the minimum bottom-line guard unit still participates in the execution, ensuring that the basic mode is not used as the default fallback and that the bypass path has boundaries. The reason code and the two-person verification credentials are written into the audit event, so that the emergency activation mechanism and chain of responsibility can be reviewed and can be continued and governed in step four.

[0094] Step 4: Using the protocol fingerprint as an index, close and preserve the execution facts and rule boundaries within the session, and generate governance task sheets and operation and maintenance action sequences accordingly. This ensures that subsequent sessions operate along the same composite rule constraint set in the links from Step 1 to Step 3. In interconnected scenarios, key differences are often scattered across multiple system logs. If these differences are not first merged into the same session sequence and boundary references are preserved, governance personnel will find it difficult to determine whether the deviation stems from input semantics, device capabilities, or differences in execution end implementation. Therefore, a single chain of session merging, summary solidification, replay verification, and difference localization is adopted, ensuring that each conclusion can point back to the protocol fingerprint and executable parameter set. .

[0095] The discussion revolves around documenting bedside actions as verifiable sequential facts. Nurses confirm the infusion pump's starting rate, trigger prompts, interruptions, enter reason codes, complete double-checks, or use controlled emergency tokens. These actions are visible on-site, but their internal representation is fragmented and inconsistently coded, easily leading to different meanings for the same-named fields during debriefing.

[0096] If only scattered logs are stored without a unified append write sequence, it is difficult to prove whether a certain adjustment occurred within the validity period of the final security policy capsule, and it is also difficult to prove whether the emergency token was reused.

[0097] Using the session identifier as the boundary, append the following elements to the audit event sequence in a fixed field order. The steps are: Step 1: Source version tag and consistency check code; Step 2: Capability list summary, conflict classification record, and final security policy capsule signature information; Step 3: Initial setup event, adjustment event, hard limit trigger event, emergency token event, and receipt confirmation event. Each event is recorded with event type, sequence number, and protocol fingerprint. The target device identifier and payload digest are used. The payload digest is generated by a deterministic serialization operator and a collision-resistant mapping operator to eliminate coding differences between different systems.

[0098] The audit chain summary is calculated using the following formula and written to the beginning and end of the session as an immutable facade identifier:

[0099] In the formula: audit chain summary Fixed-length digest, with values ​​within a finite bit string space; collision-resistant mapping operator. Deterministic hash operators whose value range is a fixed-length digest space; for example, using a hash algorithm with a fixed output length; and requiring that the hash output length is fixed and that the same implementation version is used on the execution side. Deterministic serialization operator Deterministic encoding operators, whose value range is a finite set of byte sequences, encode input objects into unique byte sequences. They specify the field order, encoding method, and set field sorting rules. For example, numeric fields use fixed-point integer encoding (fixed decimal places or numerator / denominator encoding), string fields use a unified character set encoding with a length prefix, and set fields (allowing unit sets and allowed pattern sets) are serialized after being sorted lexicographically. Protocol fingerprint. Fixed-length digest, in a finite bit string space; executable parameter set. : An ordered set of parameters with constrained values; a composite set of rule constraints. The set of clauses, with a value range being a discrete set of clauses; the merging rules for the composite rule constraint set are as follows: intersection of unit sets is allowed; if the intersection is empty, there is no negotiable conflict; the maximum boundary is the minimum of the two, and the minimum boundary is the maximum of the two; if the minimum boundary exceeds the maximum boundary, there is no negotiable conflict; a coarser resolution (the one with the larger step size) is used to ensure executable consistency; rounding rules are based on device rules, and the device rounding version identifier is recorded at the end of the audit. Audit event sequence. : An ordered sequence of events, whose value range is a finite-length sequence space.

[0100] Specifically, append writes can occur in sequential write files or log storage with immutable policies; audit chain digests can be calculated either by hash chain or by Merkle tree segmentation, both based on determinism. The encoding version number of the deterministic serialization operator is written into the header of the candidate security policy capsule to prevent subsequent upgrades from generating different byte sequences for the same field.

[0101] First, the session identifiers are merged and appended to the audit event sequence. In the process, a deterministic serialization operator and a collision-resistant mapping operator are used to calculate the audit chain digest, which is then written to the beginning and end of the session. This unifies cross-system actions to the same order and the same protocol fingerprint, facilitating session-by-session replay. The audit chain digest cannot be inserted, deleted, or rearranged, giving the evidence chain a verifiable appearance. It also carries references to ensure that facts and boundaries are closed on a single chain. Intent: To fix scattered logs as verifiable session facts.

[0102] The discussion revolves around ensuring consistent implementation of the same constraint projection operator across different execution endpoints. If there are subtle differences in unit conversion, rounding, or resolution clipping at the execution endpoints, the on-site manifestation might be discrepancies in the prompt path or clipping results, which are difficult to pinpoint through manual review alone. To pinpoint deviations to specific components and clauses, governance actions can be transformed from abstract debates into verifiable clause revisions or implementation corrections. The input load for each adjustment is extracted from the audit event sequence, and the normalized parameter set is reconstructed according to the conversion link in step one. Then, the calculation is performed on the same playback. The maximum deviation between the component and the playback component was then recorded using the playback difference measure:

[0103] Where: playback difference : Non-negative real number, taking values Component Index Integer index, with a value range of 100%. Its function is to locate the component quantities of the parameter, where is the total number of components; and is the component value. : The first component of the executable parameter group, with a restricted value range, whose function is to represent the boundary value evidence used by the execution end.

[0104] Playback components The first component of the playback projection output has a restricted value; normalized parameter set. : Replay and reconstruct the input, with the value range limited by the input load and rounding rules; constrained projection operator : A deterministic composite mapping operator, whose value range is the set of mappings from to , and whose function is to give a unique projection result under ; a composite rule constraint set : A set of clauses, where the value range is a discrete set of clauses. When the playback difference is zero, a consistency flag is appended; when it is non-zero, a difference location record is appended. The difference location record includes the component name, clause reference, and related event sequence number. In the project implementation path, the playback device can replay events one by one using event sourcing, and the difference location can be fixed by graph query to solidify the three-way association between component, clause, and event.

[0105] First, playback calculations are performed. Generate difference location records and then append them to the audit event sequence. The replay verification transforms discrepancies into calculable quantities, and the review conclusions are repeatable. Discrepancy location records are attributed to component names and clause references, and governance actions directly connect to clause review to achieve corrections.

[0106] The evidence chain needs to be further transformed into dispatchable tasks and executable actions to reduce repetitive conflicts and emergency paths in the next round of sessions. Therefore, a deterministic task formation-constrained action selection-result write-back sequence is adopted, so that reinforcement learning only participates in the selection of operational actions and does not participate in the rewriting of hard constraint clauses.

[0107] The process revolves around integrating conflict and emergency response paths into actionable tasks. Maintenance personnel can only conduct reviews within the approval boundary after seeing session identifiers, protocol fingerprints, care unit policy identifiers, root cause labels, and clause references. If the root cause is specified within the stable label set, the governance task form will drift with changes in log definitions, preventing the review from being closed-loop.

[0108] The system extracts non-negotiable conflict records, emergency token events, and difference location records from the audit event sequence and generates root cause labels according to deterministic rules: inconsistent drug entity identifiers generate drug entity mapping conflicts; non-equivalent formulation concentrations generate formulation concentration conflicts; inconsistent resolution clipping generates step clipping conflicts; non-zero playback difference with components falling within the titration step size domain generates titration constraint realization differences; network unavailability as the cause code generates network unavailability; and device unreachability as the cause code generates device unreachability. Root cause labels and clause references are assembled into a governance task order and dispatched to the work order system, which uses session identifiers and protocol fingerprints as retrieval keys. In the engineering implementation path, root cause rules can be executed by the rule engine, and clause references can be represented by clause numbers or clause summary fingerprints, avoiding the need to expand the full clause content in the task order.

[0109] First, conflict, emergency, and difference location records are extracted from the audit event sequence and root cause labels are generated. Then, the root cause labels and protocol fingerprints are combined. Nursing unit strategy identifiers and clause references are assembled into governance task sheets and dispatched. Original differences are compressed into stable root cause labels, ensuring consistent task dispatch and facilitating review. The task sheet carries nursing unit strategy identifiers and clause references, ensuring that review falls on specific clause fragments in the static rule base.

[0110] The discussion revolves around how to select actions during the governance window. In this solution, reinforcement learning only outputs a sequence of operational actions, which are limited to four types: reminding devices to load the rule base version, adjusting the final security policy capsule distribution window, increasing the review priority of specific root cause tags, and triggering a two-person verification reminder. Furthermore, none of these actions may rewrite the hard constraint boundaries of the composite rule constraint set.

[0111] It takes time for a governance task order to go from being dispatched to being closed in time. New sessions will still occur during the window period. If the order of reaching and reviewing is fixed for a long time, emergency paths may be repeatedly triggered in some nursing units and review resources may be consumed.

[0112] The audit event sequence and governance task status are converted into an operation and maintenance (O&M) status vector, which drives action selection. Before an action is implemented, it is verified by a hard-restriction firewall. The hard-restriction firewall only checks whether the action involves rewriting hard-restriction clauses; if so, it is rejected and manual approval is required. In the engineering implementation path, action selection can adopt a contextualized multi-arm selection method or a policy gradient method. Offline training can use historical replay sequences and importance sampling to correct historical policy biases. Gradient calculation can adopt an automatic differentiation framework, and the gradient only applies to the action selection parameters. In an equivalent implementation, if the organization does not use reinforcement learning, the upper confidence bound method or Thompson sampling method can be used to select within the same action set.

[0113] Operation and maintenance state vector Extracted from the session evidence chain within a time window, it includes at least counting and identifying features such as the number of non-negotiable conflicts, the number of emergency tokens, the number of device unreachable reason codes, the number of network unavailability reason codes, the number of times the same care unit is triggered, the number of times the rule base version is unknown, and the number of devices to be activated. All features can be obtained by aggregation by time window; maintenance action set. :Limited to actions that do not trigger hard restrictions, such as scheduling equipment maintenance windows, sending rule base activation reminders, adjusting security policy capsule distribution times, adjusting review queue sorting, and triggering dual-person verification reminders; Reward function : It is obtained by aggregating audit events in the next time window. For example, negative rewards are formed by weighting emergency tokens, non-negotiable conflicts, and repeated reminders. The weights are configured by the organization and written into the operation and maintenance terms area of ​​the static rule base.

[0114] Action selection and update: Contextualized multi-arm selection with exponential weighted update is employed to maintain the weight parameters of each action under the state features, and soft maximum distribution is used to select actions. The in-line strategy is expressed as: in the state, a score is calculated for each action. , and according to

[0115] Perform the sampling selection action. Then press:

[0116] Update, where is the learning rate; is the operational state vector. The value range is a finite-dimensional real vector space, and its function is to express the aggregation state of the audit evidence chain within a time window; action set. The value range is a finite set of actions, and its function is to limit the executable operation and maintenance actions; action , The range of values ​​is Its function is to represent candidate actions and selected actions; scoring function The value range is for real numbers, and its function is to express the relative preference of the action in the current state; its specific implementation can be linear scoring or table lookup scoring, whichever is more suitable for those skilled in the art; policy distribution The range of values ​​is and the summation of is . Its function is to provide the probability of action selection; Learning rate The range of values ​​is Its function is to control the rate of score updates; rewards The value range is real numbers, and its function is to reflect the changing trend of subsequent audit events caused by the action. It is specifically obtained by aggregation. First, reinforcement learning outputs a sequence of maintenance actions based on the task order status. Then, after verification by a hard-restriction firewall, it performs maintenance window orchestration, reminder delivery, and review queue sorting, and writes the action execution results back to the audit event sequence. Reinforcement learning selects only from the set of actions and is constrained by hard-limit firewalls, preventing it from violating hard-limit clauses.

[0117] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0118] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0119] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0120] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0121] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A real-time medication control system based on reinforcement learning and a static rule base, characterized in that: include, At the trigger points of signing, verifying or executing medical orders, medication information is collected from the medical order system, pharmacy preparation system and static rule base, and semantically normalized and unit converted. Negotiable and non-negotiable fields are distinguished, and protocol fingerprints are generated to form candidate security strategy capsules. Based on candidate security policy capsules and protocol fingerprints, obtain the device capability list from the target execution end; Based on the equipment capability list, the negotiable fields are converted and trimmed, conflicts of non-negotiable fields are determined, and the final security policy capsule is generated and signed. After the target execution end verifies and approves the final safety strategy capsule, it enters the fenced interconnection mode. The final safety strategy capsule constraint settings and adjustments are made, and the basic mode start path without drug selection is restricted. In case of an anomaly, the controlled emergency token is activated and the reason code is recorded. Write the protocol fingerprint, device capability list, conflict determination results, final security policy capsule, and execution records into the audit chain; Based on the audit chain and without changing the hard constraints of the static rule base, reinforcement learning is used to generate governance actions and write them back.

2. The real-time medication control system according to claim 1, characterized in that: A session identifier is established based on the medical order trigger event, and the source version mark is recorded. Fields with the same name in the medical order system, pharmacy preparation system and static rule base are sealed. Subsequent changes are written into the audit link and associated with the session identifier. A consistency check code is generated for the sealed field and written into the candidate security policy capsule.

3. The real-time medication control system according to claim 2, characterized in that: The drug identifier is mapped and cross-checked using both the coding path and the name-preparation attribute path to obtain a unified identifier for the drug entity. The route of administration is then bound to the preparation constraints of the pharmacy preparation system and written into the candidate safety strategy capsule. When the results of the two paths are inconsistent, the unified identifier for the drug entity is included in the non-negotiable field.

4. The real-time medication control system according to claim 3, characterized in that: The dose and rate expression fields, standard concentration and weight-related fields are converted to a unified metrology system, and the values ​​before and after pruning are generated according to the rounding and resolution rules of the static rule base. Then, the key fields are spliced ​​in a fixed field order to generate a protocol fingerprint, and the hard limit, soft limit, titration step size and minimum adjustment interval are assembled into a candidate safety strategy capsule.

5. The real-time medication control system according to claim 4, characterized in that: The device capability list includes unit set, minimum resolution, available injection modes and hardware limits, and is bound to session identifiers and protocol fingerprints; the constraints of the static rule base and the device capability list are merged by intersection with stricter boundary rules as a source of constraints for the conversion and pruning of negotiable fields.

6. The real-time medication control system according to claim 5, characterized in that: First, perform a consistency check on non-negotiable fields and generate a non-negotiable conflict record if the check is not met. This record is then written into the audit chain and triggers pharmacist review. After the consistency check is passed, perform unit conversion, rounding, and resolution pruning on negotiable fields and create pruning traces, which are then written into the final security policy capsule.

7. The real-time medication control system according to claim 6, characterized in that: After the target execution terminal passes the verification, it loads the final security strategy capsule and enters the fenced interconnected mode state machine. During the initial setup, it verifies the unit and hard limit boundaries. During the adjustment, it adjusts the entry point by gating at the minimum adjustment interval. The operation of entering the basic mode without drug selection is transferred to the controlled emergency token process and the event is recorded.

8. The real-time medication control system according to claim 7, characterized in that: The controlled emergency token is a one-time, short-term, and verifiable data object, which includes minimum bottom-line guards for allowed unit sets and global maximum and minimum rates. Before enabling the controlled emergency token, the target execution end is required to enter the reason code and complete the double-verification of credentials, and then set and adjust the minimum bottom-line guard limit.

9. The real-time medication control system according to claim 8, characterized in that: The audit event sequence is appended only, consisting of session identifier merging protocol fingerprint, device capability list, conflict record, final security policy capsule, execution record, and controlled emergency token event. The audit event sequence is deterministically serialized according to a fixed field order, and the audit chain summary is calculated and written into the audit chain.

10. The real-time medication control system according to claim 9, characterized in that: Governance actions include adjusting the strategy push rhythm, setting the activation reminder intensity, and determining the review priority and work order dispatch order. Reinforcement learning uses event statistics in the audit chain as input and output for governance actions, and performs hard constraint verification on governance actions before execution to ensure that governance actions do not change the hard constraints of the static rule base.