A multi-modal guided emergency supplies positioning and medical order recording system

The multimodal guided emergency supplies location and medical order recording system utilizes voice recognition and LED indicator lights to achieve precise drug location and timely recording of medical orders. This solves the problems of low efficiency, error-proneness, and information silos in existing technologies, and achieves efficient supplies management and medical order recording.

CN122417331APending Publication Date: 2026-07-17SHENZHEN LONGGANG DISTRICT MATERUITY & CHILD HEALTHCARE HOSPITAL

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN LONGGANG DISTRICT MATERUITY & CHILD HEALTHCARE HOSPITAL
Filing Date
2026-04-20
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In pre-hospital emergency care and in-hospital emergency rescue scenarios, existing technologies rely on manual searching for medicines and equipment, which is inefficient and prone to errors. They lack efficient intelligent material management and medical order recording, resulting in information silos and heavy cognitive burdens. Records are inaccurate and delayed, and a closed-loop data flow cannot be formed.

Method used

The emergency supplies location and medical order recording system adopts multimodal guidance. It collects voice commands through a directional microphone array, combines a central processing unit and a visually guided medicine cabinet, uses LED indicator lights for precise location guidance, and realizes timely recording and display of medical order information through voice recognition and structured record generation modules.

Benefits of technology

It improved the accuracy of emergency supplies guidance and the efficiency of medical order recording, realized the timely recording and traceability reliability of medical order information, reduced cognitive burden and error occurrence, and formed a closed-loop data flow of instruction-execution-recording.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122417331A_ABST
    Figure CN122417331A_ABST
Patent Text Reader

Abstract

This application discloses a multimodal guided emergency supplies location and medical order recording system, relating to the field of medical equipment technology. The system includes: a central processing unit, a mobile interactive terminal, and a visually guided medicine cabinet. The mobile interactive terminal has a built-in directional microphone array, used to collect on-site voice through the directional microphone array and transmit the voice to the central processing unit. The central processing unit issues LED control commands to the visually guided medicine cabinet based on the on-site voice and records the medical order information identified from the on-site voice. The visually guided medicine cabinet includes multiple partitions, each corresponding to an independently addressable LED indicator. The visually guided medicine cabinet has a built-in medicine cabinet controller, used to control the corresponding LED indicator to light up according to the LED control commands. This application improves the accuracy of emergency supplies guidance and the efficiency of medical order recording.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical device technology, and in particular to a multimodal guided emergency supplies positioning and medical order recording system. Background Technology

[0002] In pre-hospital emergency care (such as inside an ambulance) and in-hospital emergency rescue scenarios, medical staff need to quickly and accurately retrieve various medicines and equipment from medicine cabinets in tense, noisy, bumpy, and space-constrained environments. Currently, this generally relies on traditional manual methods.

[0003] (1) Artificial memory and visual search: Medical staff rely on memory and visual sight to find target items in the medicine cabinet, which is inefficient and prone to errors under pressure.

[0004] (2) Verbal orders versus manual records: Doctors give verbal orders, nurses execute them manually, and all emergency measures need to be recorded in paper or electronic medical records afterward based on memory. This process has problems such as omissions in records, inaccurate timing, and fragmented information.

[0005] (3) Shortcomings of existing intelligent attempts: Some intelligent medicine cabinets adopt the "voice command + mechanical pop-up" solution, which is complex in structure and high in cost. In mobile emergency scenarios, the reliability is poor, the response is slow, and there is a risk of failure. Independent voice recording devices can only convert voice into text and are not integrated with the material retrieval process, so they cannot form a closed-loop data flow of "command-execution-recording". Existing solutions have not solved the medical staff's need for immediate access to information such as drug usage, dosage, and contraindications, and have failed to reduce their cognitive load.

[0006] In summary, the shortcomings of existing technologies are as follows.

[0007] (1) Low and uncertain retrieval efficiency: It relies on manual search, which is time-consuming and easily affected by the environment and personal state.

[0008] (2) Significant safety risks: There is a lack of effective error prevention and verification mechanisms for the use of high-risk drugs.

[0009] (3) Severe information silos: Material management, medical order issuance, and record generation are disconnected, and data cannot be automatically linked.

[0010] (4) Heavy cognitive burden: Medical staff need to memorize a large amount of drug information, and memory deviation is likely to occur in a stressful environment.

[0011] (5) Inaccurate and delayed records: The quality of the emergency records made after the fact is poor and cannot provide a basis for medical quality analysis and legal tracing. Summary of the Invention

[0012] The purpose of this application is to provide a multimodal guided emergency supplies location and medical order recording system, which improves the accuracy of emergency supplies guidance and the efficiency of medical order recording.

[0013] To achieve the above objectives, this application provides the following solution: In the first aspect, this application provides a multimodal guided emergency supplies location and medical order recording system, including: a central processing unit, a mobile interactive terminal, and a visually guided medicine cabinet; The mobile interactive terminal has a built-in directional microphone array, which is used to collect on-site voice through the directional microphone array and transmit the on-site voice to the central processing unit. The central processing unit sends control commands to the light-emitting diodes (LEDs) of the visually guided medicine cabinet based on the on-site voice, and records the medical order information identified from the on-site voice. The visually guided medicine cabinet includes multiple partitions, each partition corresponding to an independently addressable LED indicator; the visually guided medicine cabinet has a built-in medicine cabinet controller, which is used to control the corresponding LED indicator to light up according to the LED control command.

[0014] Optionally, the LED control instructions include a medicine cabinet ID, a zone ID, and a lighting mode, with each zone corresponding to a zone ID; there are multiple lighting modes, each representing a type of medicine / item and its urgency level, and each lighting mode is displayed through a combination of LED color and flashing pattern.

[0015] Optionally, each partition is equipped with a weight sensor or a radio frequency identification (RFID) reader, which is used to feed back the item retrieval status of the corresponding partition to the central processing unit in real time.

[0016] Optionally, the central processing unit includes a speech recognition engine, which employs an end-to-end deep learning architecture; The speech recognition engine includes a noise reduction module, an acoustic feature extraction module, an acoustic model, a language model, and a key information extraction module; The noise reduction module employs a multi-channel adaptive beamforming algorithm to enhance the target speech direction signal in the on-site speech through the directional microphone array, thereby obtaining the noise-reduced speech. The acoustic feature extraction module is used to convert the noise-reduced speech into Mel frequency cepstral coefficients or filter bank features to obtain speech domain features. The acoustic model is used to determine the phoneme state sequence based on the speech domain features; The language model is used to determine a text sequence based on the phoneme state sequence; the text sequence constitutes the medical order information. The key information extraction module is used to extract structured information from the text sequence; the structured information includes drug / item name, dosage, dosage form, route of administration, and operation actions.

[0017] Optionally, the acoustic model is obtained by training a time-delay neural network or a convolutional augmented Transformer model using a corpus specific to emergency rescue scenarios. The language model adopts a Transformer-based statistical language model, which decodes the phoneme state sequence with the entries in the preset emergency medical lexicon as output constraints to obtain the text sequence; the Transformer-based statistical language model is trained on a training set composed of the preset emergency medical lexicon. The key information extraction module uses a medical-specific named entity recognition model.

[0018] Optionally, the central processing unit further includes an instruction parsing and linkage module; The instruction parsing and linkage module is used to retrieve the partition ID from the drug-location mapping table based on the structured information, and drive the LED indicator corresponding to the retrieved partition ID to light up; the drug-location mapping table is used to store the partition ID corresponding to each drug / item.

[0019] Optionally, when the visually guided medicine cabinet is replenishing medicines / items, the mobile interactive terminal is used to obtain medicine / item information by scanning the barcode of the medicine / item, and when the user lights up the LED indicator and confirms the partition ID corresponding to the lit LED indicator on the display interface of the mobile interactive terminal, the currently replenished medicine / item is bound to the partition ID corresponding to the lit LED indicator and stored in the medicine-location mapping table.

[0020] Optionally, the central processing unit further includes a structured record generation module; The structured record generation module adopts an event-driven architecture. The event-driven architecture predefines multiple event types. When each event type is triggered, the event occurrence time, event type, associated data, and operator ID are encapsulated into a structured data entry in JSON format and stored as an operation log. The event types include voice recognition completion, LED indicator light illuminating, and medicine / item being taken away; The associated data includes the names of drugs / items.

[0021] Optionally, the display interface of the mobile interactive terminal includes a real-time voice-to-text area, a dynamic drug guidance card area, and a rescue timeline log area; The real-time speech-to-text area is used to display the recognized medical order information in real time. The dynamic drug guidance card area is used to pop up a floating card when the drug / item name is recognized. The floating card includes the standard name and specifications of the drug, recommended usage, key contraindications and warning information, as well as relevant information about the current patient. The relevant information about the current patient includes age and weight. The rescue timeline log area is used to display the operation log in reverse chronological order. The real-time speech-to-text area is located at the top of the display interface, the dynamic drug guidance card area is located in the center of the display interface, and the rescue timeline log area is located on the right side or bottom of the display interface.

[0022] Optionally, the system also includes a digital display screen integrated into the visually guided medicine cabinet, which dynamically displays the identified medical order information in real time and provides voice broadcast.

[0023] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a multimodal guided emergency supplies positioning and medical order recording system. By recognizing the collected on-site voice, LED control commands are obtained to automatically generate guidance information. Precise guidance is provided through independently addressable LED indicator lights, improving the accuracy of emergency supplies guidance. Medical order information recognized from on-site voice is recorded in a timely manner, improving the efficiency of medical order recording and enhancing the traceability reliability of medical order information. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a schematic diagram of a multimodal guided emergency supplies location and medical order recording system provided in one embodiment of this application.

[0026] Figure 2 This is a schematic diagram of the rectification process and multimodal guidance provided in an embodiment of this application.

[0027] Figure 3 This is a schematic diagram of the display interface of a mobile interactive terminal provided in an embodiment of this application.

[0028] Figure 4 This is a schematic diagram illustrating the technical principle of speech recognition and information extraction provided in an embodiment of this application.

[0029] Figure label: 101-Visually guided medicine cabinet, 102-Central processing unit, 103-Mobile interactive terminal, 1011-Medicine cabinet controller, 1012-LED indicator light. Detailed Implementation

[0030] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0031] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0032] In one exemplary embodiment, this application provides a multimodal guided emergency supplies location and medical order recording system, such as... Figure 1 As shown, a multimodal guided emergency supplies location and medical order recording system includes: a central processing unit 102, a mobile interactive terminal 103, and a visually guided medicine cabinet 101.

[0033] The mobile interactive terminal 103 has a built-in directional microphone array, which is used to collect on-site voice through the directional microphone array and transmit the on-site voice to the central processing unit 102.

[0034] The central processing unit 102 issues LED control commands to the visually guided medicine cabinet 101 based on the on-site voice, and records the medical order information identified from the on-site voice.

[0035] The visually guided medicine cabinet 101 includes multiple partitions, each corresponding to an independently addressable LED indicator 1012. The visually guided medicine cabinet 101 has a built-in medicine cabinet controller 1011, which controls the corresponding LED indicator 1012 to light up according to the LED control commands. The independently controllable LED indicator 1012 achieves pixel-level precise positioning guidance for medicines / items, which is more accurate and faster than lighting up the entire drawer or cabinet door.

[0036] This application recognizes collected on-site voice recordings to obtain LED control commands, thereby automating the generation of guidance information. Precise guidance is provided through independently addressable LED indicator lights 1012, improving the accuracy of emergency supplies guidance. Timely recording of medical orders identified from on-site voice recordings improves the efficiency of medical order recording and enhances the traceability reliability of medical order information.

[0037] In one exemplary embodiment, the directional microphone array is a high-sensitivity directional microphone array. The directional microphone array is connected to the central processing unit 102 via USB or Bluetooth to form a data acquisition link, which is responsible for collecting voice commands at the emergency scene and transmitting them to the central processing unit 102.

[0038] The central processing unit 102 serves as the core processing unit, employing an embedded industrial computer. It communicates bidirectionally with the mobile interactive terminal 103 via industrial-grade Wi-Fi / wired Ethernet to receive voice data and issue LED control commands. Simultaneously, it connects to the intelligent vision-guided medicine cabinet (vision-guided medicine cabinet 101) via a controller area network (CAN) bus or an RS485 industrial bus to ensure stable and reliable transmission of control commands in bumpy environments.

[0039] In an exemplary embodiment, the LED control instructions include a medicine cabinet ID, a zone ID, and a lighting mode, with each zone corresponding to a zone ID; there are multiple lighting modes, each representing a type of medicine / item and its urgency level, and each lighting mode is displayed through a combination of LED color and flashing pattern.

[0040] The visually guided medicine cabinet 101 includes multiple open drawers, each of which is divided into several independent storage units by partitions, with each independent storage unit serving as a partition. LED indicator lights 1012 are directly integrated into the front of the partitions or the inside of the drawer panels, ensuring that nurses can see the illuminated position at a glance the moment they open the drawer.

[0041] In one exemplary embodiment, each partition is equipped with a weight sensor or an RFID reader, which is used to provide real-time feedback on the drug / item retrieval status of the corresponding partition to the central processing unit. When a drug or item is taken, the weight sensor or RFID reader is triggered to generate an execution record, which is then fed back to the central processing unit.

[0042] The functions of the central processing unit 102 can be partially integrated into the mobile interactive terminal 103 or the medicine cabinet controller 1011 to form a distributed or integrated architecture.

[0043] The execution feedback link of this application specifically includes: the built-in medicine cabinet controller 1011 of the visually guided medicine cabinet 101 receives LED control commands from the central processing unit 102 and drives the corresponding LED indicator 1012 to light up; at the same time, the optional weight sensor or RFID reader in the medicine cabinet partition feeds back the item retrieval status to the central processing unit 102 in real time through the same bus.

[0044] The medicine cabinet controller 1011 is a control module with a microprocessor, which is responsible for communication relay, drive execution and status acquisition.

[0045] Communication relay: Receives control commands (including target medicine cabinet ID, zone ID, and lighting mode commands) from the central processing unit 102 via the industrial bus.

[0046] Drive execution: Parse the instruction into a specific level signal or PWM signal, and drive the LED indicator 1012 of the corresponding storage unit to light up according to the specified mode.

[0047] Status acquisition: Collect sensor data from each zone and transmit the status of the medicine cabinet (such as cabinet door opening and closing, item retrieval) back to the central processing unit 102.

[0048] In one exemplary embodiment, the mobile interactive terminal 103 is a shockproof, sterilizable industrial tablet PC that can be fixed with a stand or held by hand. It has a built-in high-sensitivity microphone. Its dedicated application interface (such as...) Figure 3 The design (as shown) consists of three core areas: a real-time voice-to-text area, a dynamically pop-up drug / equipment guidance card area (dynamic drug guidance card area), and a real-time updated rescue timeline log area.

[0049] In one exemplary embodiment, the mobile interactive terminal 103 can be a smart head-mounted device (augmented reality (AR) glasses) or integrated with a vehicle-mounted screen. In the AR glasses solution, guidance information can be directly overlaid on the field of view of the actual medicine cabinet, in which case the light guidance may be simplified to ordinary labels.

[0050] In one exemplary embodiment, the central processing unit 102, serving as the system brain, may be an embedded industrial computer.

[0051] The software core of the central processing unit 102 includes a speech recognition engine, an instruction parsing and linkage module, and a structured record generation module.

[0052] Speech recognition engine: It adopts an acoustic model trained with a massive amount of emergency rescue scene audio and professional terminology database, and integrates noise reduction algorithm to ensure recognition rate in the noisy environment of ambulance.

[0053] Command parsing and linkage module: Maps the identified item name (such as "adrenaline") and specification (such as "1mg") to a unique "medicine cabinet-layer-partition" code, and generates corresponding LED control commands and screen display commands.

[0054] Structured record generation module: Monitors all voice events and system state changes, and automatically generates JSON or similar structured data entries in the "time-event-subject" format.

[0055] The speech recognition engine adopts an end-to-end deep learning architecture, and its operating principle is as follows: Figure 4 As shown.

[0056] The speech recognition engine includes a noise reduction module, an acoustic feature extraction module, an acoustic model, a language model, and a key information extraction module.

[0057] The noise reduction module employs a multi-channel adaptive beamforming algorithm to enhance the target speech direction signal in the on-site speech through the directional microphone array, thereby obtaining the noise-reduced speech.

[0058] The noise reduction module is mainly designed to address the low-frequency engine noise and high-frequency alarm sounds unique to ambulances. It employs a multi-channel adaptive beamforming algorithm to enhance the target speech direction signal in the on-site speech through the directional microphone array, while suppressing noise from non-target directions, thus obtaining noise-reduced speech.

[0059] The acoustic feature extraction module is used to convert the denoised speech into Mel-Frequency Cepstral Coefficients (MFCC) or filter bank features (Fbank) to obtain speech domain features.

[0060] The acoustic model is used to determine the phoneme state sequence based on the speech domain features.

[0061] The language model is used to determine the text sequence based on the phoneme state sequence; the text sequence constitutes the medical order information and the emergency log, avoiding the need for medical staff to manually record information during on-site emergency rescue.

[0062] The key information extraction module is used to extract structured information from the text sequence; the structured information includes drug / item name, dosage, dosage form, route of administration, and operation actions.

[0063] In one exemplary embodiment, the acoustic model employs a Time Delay Neural Network (TDNN) or a Convolutional Enhancement Transformer (Conformer) model, trained (fine-tuned) using corpus specific to emergency medical scenarios. The acoustic model exhibits extremely high recognition sensitivity for emergency medical terminology (such as "amiodarone" and "adrenaline").

[0064] The language model employs a Transformer-based statistical language model, combined with a pre-defined emergency medical terminology (containing drug names, operational instructions, and symptom descriptions) to decode the candidate sequences (phoneme state sequences) output by the acoustic model, obtaining the text sequence that best fits the context. Specifically, the language model uses a Transformer-based statistical language model, with entries from the pre-defined emergency medical terminology as output constraints, to decode the phoneme state sequences and obtain the text sequence; the Transformer-based statistical language model is trained on a training set composed of the pre-defined emergency medical terminology.

[0065] In Transformer-based statistical language models, incorporating an emergency medical services (EMS) lexicon is not simply about feeding the model a vocabulary list. Instead, it involves injecting domain knowledge from the lexicon into the model's training, inference, and decoding processes using various techniques. The specific implementation includes: 1. Preprocessing and vectorization of the EMS lexicon; 2. Fine-tuning the Transformer-based statistical language model using the preprocessed and vectorized EMS lexicon; 3. Implementing lexicon-guided constraint decoding at the language model decoding level.

[0066] Therefore, the language model in this application is obtained by fine-tuning the Transformer-based statistical language model based on a pre-set emergency medical terminology database. When the language model outputs a text sequence based on the phoneme state sequence, it uses entries from the emergency medical terminology database as constraints. An emergency medical terminology database index is integrated into the output of the language model to ensure that any non-medical terms not in the database are effectively suppressed.

[0067] Preprocessing and vectorization of the emergency medical terminology database involves converting the terms in the database into a form usable by the model. (1) Standardization of terms: Organize drug names (adrenaline, amiodarone), operation instructions (intravenous push, endotracheal intubation), symptom descriptions (chest pain, dyspnea), etc. That is, convert each term into the corresponding preset name to obtain standardized terms.

[0068] (2) Construct a phoneme-vocabulary mapping table: label each standardized word with a standard pinyin or phoneme sequence for subsequent acoustic alignment.

[0069] (3) Generate word vector embeddings: The standardized words in the vocabulary are initialized by pre-trained Word2Vec or BERT to obtain vector representations of fixed dimensions.

[0070] Fine-tuning the Transformer-based statistical language model is to shift the model parameters toward the corpus of emergency medical services.

[0071] Starting from the preprocessed and vectorized emergency medical lexicon, two types of training samples are constructed: (1) Masked Language Modeling: Randomly mask parts of the vocabulary in a sentence containing technical terms, and let the model predict the masked words. For example: Input: "[MASK] is an antiarrhythmic drug used to treat ventricular fibrillation", where [MASK] is the mask.

[0072] Target: Amiodarone.

[0073] (2) Terminology insertion / replacement: Replace common words in general sentences with emergency medical terminology, so that the model can learn the rationality of terms in the context.

[0074] The basic Transformer model was retrained using a first aid terminology database (two types of training samples) to give it strong prior probabilities of first aid terminology.

[0075] During the inference phase of the language model (i.e., when decoding the phoneme state sequence output by the acoustic model), the output space of the model is pruned and biased using a first aid lexicon. (1) Trie constraint The emergency medical terminology database is constructed as a Trie tree, where each node represents a prefix of a word.

[0076] During decoding, the Transformer outputs the probability distribution word by word, allowing only the current output word to be a valid continuation of a branch in the Trie tree.

[0077] If the word predicted by the model is not in the Trie vocabulary, then the probability of that word is forced to be set to zero.

[0078] [Example] The current output is "Ready".

[0079] Legitimate words following "preparation" in a Trie tree: {"1 mg", "adrenaline", "defibrillator"...}.

[0080] The language model predicts "apple" → the probability is set to zero.

[0081] The language model predicts the retention probability of "adrenaline".

[0082] (2) Dynamic lexicon bias Set a higher bias gain for entries in the emergency medical terminology database than for general terms. (Usually taken as 1.2 to 1.5) Modify the probability distribution of the Softmax output. : .

[0083] in, It is the bias value of term w (words in the dictionary > general words). For the value of term w, For the value of term v, This is the bias value for term v.

[0084] This way, even if the model's original output score for a certain first aid term is not high, it can still be selected after biasing.

[0085] (3) Phoneme-vocabulary alignment forced The phoneme sequence output by the acoustic model is dynamically aligned with the phoneme sequence of each word in the vocabulary.

[0086] Only words whose phoneme sequence matching degree exceeds the threshold are included in the candidate set.

[0087] For example: acoustic model output → Only words with high phoneme similarity, such as "amiodarone", are allowed as candidates.

[0088] The text sequence output by the speech model not only correctly identified the drug names but also provided reliable entity information for subsequent instruction parsing and linkage modules. This "lexicon-enhanced" Transformer language model enables this system to achieve significantly higher speech recognition accuracy than general solutions in noisy emergency environments.

[0089] The key information extraction module uses a medical-specific named entity recognition (NER) model (such as BioBERT or ClinicalBERT) to input the text sequence recognized by the language model into the medical-specific named entity recognition model and output structured information such as drug name, dosage, dosage form, route of administration, and operation actions.

[0090] In an exemplary embodiment, the instruction parsing and linkage module is used to retrieve the partition ID from the drug-location mapping table according to the structured information, and drive the LED indicator 1012 corresponding to the retrieved partition ID to light up; the drug-location mapping table is used to store the partition ID corresponding to each drug / item.

[0091] The central processing unit 102 maintains a drug-location mapping table, which is stored in a local database or in memory. This mapping table records the precise physical location code of each drug (distinguished by name and specification) within the medicine cabinet. The process of creating the mapping table is as follows: When the visually guided medicine cabinet 101 is installed for the first time or when medicines / items are replenished, the mobile interactive terminal 103 is used to obtain medicine / item information by scanning the barcode of the medicine / item. When the user (medicine cabinet administrator) lights up the LED indicator 1012 and confirms the partition ID corresponding to the lit LED indicator 1012 on the display interface of the mobile interactive terminal 103, the currently replenished medicine / item is bound to the partition ID corresponding to the lit LED indicator 1012 and stored in the medicine-location mapping table.

[0092] Once the voice command is parsed (e.g., recognizing the drug name "adrenaline" and the specification "1mg"), the central processing unit 102 executes the following steps: Step 1: Based on the parsing results, retrieve the corresponding physical location code from the drug-location mapping table. For example: CabinetID:01, Drawer:02, Cell:05, this location code is represented as 01-02-05. Here, the first digit, CabinetID, represents the cabinet ID; the second digit, Drawer, represents the drawer ID; and the third digit, Cell, represents the partition ID.

[0093] Step 2: Generate standard format LED control instructions based on the parsing results. The instruction format is: [Medicine Cabinet ID][Zone ID][Lighting Mode], for example, 01-02-05-RED-FLASH.

[0094] Step 3: Send the instruction to the medicine cabinet controller 1011 of the intelligent vision-guided medicine cabinet 101 via CAN bus / RS485.

[0095] Step 4: The medicine cabinet controller 1011 parses the instructions, determines the target zone, and closes the driving circuit of the corresponding LED to make it light up according to the specified mode.

[0096] Table 1 shows some information from the drug-location mapping table.

[0097] Table 1. Partial Information from the Drug-Location Map

[0098] To achieve a "ready-to-understand" guidance effect, this invention establishes indicator light coding rules, and the specific correspondence is shown in Table 2.

[0099] Table 2 Indicator Light Coding Rules

[0100] Using the coding method shown in Table 2, medical staff can not only know "where the thing is" the moment the light comes on, but also simultaneously perceive "what type of item it is" and "its urgency level", further accelerating decision-making.

[0101] The coding rule for the indicator lights in this application is to determine the color and flashing mode of the LED indicator light 1012, that is, to determine the light mode. Each light mode enables medical staff to subconsciously obtain information about the type of medicine and the degree of urgency while confirming the location, thereby further shortening the cognitive reaction time.

[0102] This application sets predefined colors and flashing patterns for LED indicator lights 1012 corresponding to different drugs based on drug type or urgency level, so as to achieve one-time visual communication of complex information.

[0103] In one exemplary embodiment, small digital displays are provided in each partition to display the abbreviated name of the drug.

[0104] Extended triggering methods: In addition to voice, the system can support gesture recognition (making specific gestures in front of the camera) or physical button panels (setting shortcut buttons for commonly used medicines) to trigger guidance.

[0105] The guidance information is not only displayed on the mobile interactive terminal 103, but can also be provided as a secondary prompt through the small digital display screen integrated into the visual guidance medicine cabinet 101 or through voice broadcast.

[0106] In one exemplary embodiment, the structured record generation module employs an event-driven architecture. This architecture predefines multiple event types. When each event type is triggered, the event occurrence time, event type, associated data, and operator ID are encapsulated into a structured data entry in JSON (JavaScript Object Notation) format and stored as an operation log. The event occurrence time is accurate to milliseconds.

[0107] Operator ID is determined through voiceprint recognition, facial recognition, or fingerprint recognition. Operator identification ensures the uniqueness of the operator's identity and further enhances the security level.

[0108] The event types include voice recognition completed (EVENT_VOICE_RECOGNIZED), LED indicator lit (EVENT_LED_ACTIVATED), medicine / item taken away (EVENT_ITEM_TAKEN), and user confirmation (EVENT_USER_CONFIRM).

[0109] The associated data includes the names of drugs / items.

[0110] This application records operation logs through event-driven logging. The generation of operation logs is completely decoupled from the operation process, achieving truly "seamless" recording. Each record is also identified by the operator, ensuring traceability of responsibility.

[0111] Operation logs can be stored in various ways, including local caching, cloud synchronization, and redundant backup.

[0112] Local caching: Structured data entries are first stored in the local time-series database (such as SQLite or InfluxDB) of the central processing unit 102 to ensure that no records are lost in a network-off environment.

[0113] Cloud synchronization: When the network is restored, the system will encrypt and upload locally cached data in batches to the hospital information system or cloud emergency data platform via MQTT or HTTPS protocol.

[0114] Redundant backup: Critical operation logs are simultaneously written to the local storage of the mobile interactive terminal 103 to achieve multiple backups.

[0115] The functions of the various storage methods for operation logs are as follows.

[0116] Accountability: Generate an immutable and precise timeline as evidence of medical procedures.

[0117] Medical debriefing: After the rescue is completed, the entire rescue process can be reviewed based on a timeline to analyze decision-making points and execution efficiency.

[0118] Quality control: Statistical analysis of indicators such as drug dispensing time and operational standardization is used to improve the quality of emergency care.

[0119] Scientific analysis: Providing high-precision, structured raw data for emergency medicine research.

[0120] Automatic billing and consumables management: Records can be directly integrated with the hospital's ERP system to achieve automatic drug billing and automatic inventory deduction.

[0121] In one exemplary embodiment, the display interface of the mobile interactive terminal 103 includes a real-time voice-to-text area, a dynamic drug guidance card area, and a rescue timeline log area.

[0122] The real-time speech-to-text area is used to display the recognized speech command text (medical order information) in real time, and to highlight key information such as the name and dosage of the successfully parsed drug.

[0123] The dynamic drug instruction card area is used to pop up a floating card when the drug / item name is recognized. The floating card includes the standard name and specifications of the drug, recommended usage (dosage, diluent, infusion rate), key contraindications and warnings (marked in red), and current patient information. The current patient information includes age and weight, which are used to facilitate dosage calculation. The floating card is semi-transparent. This application automatically pops up the floating card, transforming the static drug instruction manual into a contextualized, dynamic information push, directly supporting clinical decision-making and reducing hesitation.

[0124] The rescue timeline log area is used to display automatically generated operation logs in reverse chronological order. Each operation log contains a precise timestamp, operation description, and operator information. Real-time generation and display of the timeline log makes the rescue process visible to all members, facilitating team collaboration and command. Furthermore, the records are tamper-proof and of high value.

[0125] The real-time speech-to-text area is located at the top of the display interface, the dynamic drug guidance card area is located in the center of the display interface, and the rescue timeline log area is located on the right side or bottom of the display interface.

[0126] In one exemplary embodiment, such as Figure 2 As shown, the workflow of the multimodal guided emergency supplies positioning and medical order recording system of this application is as follows.

[0127] Step 1: Voice input: The doctor or nurse speaks the command into the mobile interactive terminal 103, such as "Prepare amiodarone 150mg for intravenous injection".

[0128] Step 2: Recognition and Analysis: The central processing unit 102 recognizes the speech and extracts key entities ("Amiodarone", "150mg", "Vein").

[0129] Step 3: Multimodal bootstrap triggering (parallel execution): 1) Light guidance: The central unit sends a command to the medicine cabinet to light up the LED lights of the specific compartments where the 150mg amiodarone ampoules are stored (e.g., the second drawer and the third compartment with flashing red lights).

[0130] 2) Screen display guidance: The central unit sends a command to the mobile interactive terminal 103, and the translated text "Medical order: Amiodarone 150mg IV" is displayed in the center of the screen. An amiodarone guidance card is automatically popped up, which displays the configuration method, infusion rate, precautions, etc. in detail.

[0131] Step 4: Staff Operation: The nurse proceeds directly to the corresponding drawer based on the flashing light, and retrieves the medication from the precisely illuminated compartment. On-screen instructions can be used as a reference during the operation.

[0132] Step 5: Synchronous Recording: After Step 2 is completed, the system generates a record entry "[Timestamp] Voice command: 'Prepare amiodarone...'". When the light is activated, add "[Location] Medicine cabinet X zone Y grid activated". All records are pushed to the terminal's timeline log area in real time, forming a complete rescue narrative sorted by time.

[0133] In one exemplary embodiment, laser projection is used to project the outline or name of medicines onto the surface of a medicine cabinet or inside a drawer, instead of LED indicator lights.

[0134] The large amount of structured operation logs accumulated in this application can be combined with machine learning algorithms to analyze the optimal medication path and operation mode under different conditions and scenarios. In the future, it can be upgraded into an intelligent decision support system that can proactively suggest the best solution when doctors issue instructions.

[0135] This application has the following technical advantages.

[0136] 1. Non-mechanical precision guidance mechanism: The smallest storage unit in the medicine cabinet is visually encoded and guided by an independently controllable LED indicator network, replacing the traditional mechanical pop-up solution.

[0137] 2. Contextualized information push: When a voice command triggers the location of an item, the standardized usage guide information for that item is automatically associated with and pops up on the mobile interactive terminal, achieving "instant access and instant knowledge".

[0138] 3. Automatic closed-loop generation of process data: Automatically associates voice commands, light guidance events, and time information to generate a structured and visualized timeline log of the rescue process in real time, without any manual input.

[0139] 4. Refined visual coding guidance mechanism: Standardized mapping rules have been established between drug type, urgency level and LED color and flashing mode, upgrading the single physical location light guidance to a composite visual guidance with semantic information.

[0140] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0141] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A multimodal guided emergency supplies location and medical order recording system, characterized in that, include: Central processing unit, mobile interactive terminal and visually guided medicine cabinet; The mobile interactive terminal has a built-in directional microphone array, which is used to collect on-site voice through the directional microphone array and transmit the on-site voice to the central processing unit. The central processing unit issues LED control commands to the visually guided medicine cabinet based on the on-site voice, and records the medical order information identified from the on-site voice. The visually guided medicine cabinet includes multiple partitions, each partition corresponding to an independently addressable LED indicator; the visually guided medicine cabinet has a built-in medicine cabinet controller, which is used to control the corresponding LED indicator to light up according to the LED control command.

2. The multimodal guided emergency supplies positioning and medical order recording system according to claim 1, characterized in that, The LED control commands include medicine cabinet ID, zone ID, and lighting mode. Each zone corresponds to a zone ID. There are multiple lighting modes, each representing a type of medicine / item and its urgency level. Each lighting mode is displayed through a combination of LED color and flashing pattern.

3. The multimodal guided emergency supplies positioning and medical order recording system according to claim 1, characterized in that, Each partition is equipped with a weight sensor or an RFID reader, which is used to feed back the item retrieval status of the corresponding partition to the central processing unit in real time.

4. The multimodal guided emergency supplies positioning and medical order recording system according to claim 1, characterized in that, The central processing unit includes a speech recognition engine, which employs an end-to-end deep learning architecture. The speech recognition engine includes a noise reduction module, an acoustic feature extraction module, an acoustic model, a language model, and a key information extraction module; The noise reduction module employs a multi-channel adaptive beamforming algorithm to enhance the target speech direction signal in the on-site speech through the directional microphone array, thereby obtaining the noise-reduced speech. The acoustic feature extraction module is used to convert the noise-reduced speech into Mel frequency cepstral coefficients or filter bank features to obtain speech domain features. The acoustic model is used to determine the phoneme state sequence based on the speech domain features; The language model is used to determine a text sequence based on the phoneme state sequence; the text sequence constitutes the medical order information. The key information extraction module is used to extract structured information from the text sequence; the structured information includes drug / item name, dosage, dosage form, route of administration, and operation actions.

5. The multimodal guided emergency supplies positioning and medical order recording system according to claim 4, characterized in that, The acoustic model is obtained by training a time-delay neural network or a convolutional augmented Transformer model using a corpus specific to emergency rescue scenarios. The language model adopts a Transformer-based statistical language model, which decodes the phoneme state sequence with the entries in the preset emergency medical lexicon as output constraints to obtain the text sequence; the Transformer-based statistical language model is trained on a training set composed of the preset emergency medical lexicon. The key information extraction module uses a medical-specific named entity recognition model.

6. The multimodal guided emergency supplies positioning and medical order recording system according to claim 4, characterized in that, The central processing unit also includes an instruction parsing and linkage module; The instruction parsing and linkage module is used to retrieve the partition ID from the drug-location mapping table based on the structured information, and drive the LED indicator corresponding to the retrieved partition ID to light up; the drug-location mapping table is used to store the partition ID corresponding to each drug / item.

7. The multimodal guided emergency supplies positioning and medical order recording system according to claim 4, characterized in that, When the visually guided medicine cabinet is replenishing medicines / items, the mobile interactive terminal is used to obtain medicine / item information by scanning the barcode of the medicine / item. When the user lights up the LED indicator and confirms the partition ID corresponding to the lit LED indicator on the display interface of the mobile interactive terminal, the currently replenished medicine / item is bound to the partition ID corresponding to the lit LED indicator and stored in the medicine-location mapping table.

8. The multimodal guided emergency supplies positioning and medical order recording system according to claim 4, characterized in that, The central processing unit also includes a structured record generation module; The structured record generation module adopts an event-driven architecture. The event-driven architecture predefines multiple event types. When each event type is triggered, the event occurrence time, event type, associated data, and operator ID are encapsulated into a structured data entry in JSON format and stored as an operation log. The event types include voice recognition completion, LED indicator light illuminating, and medicine / item being taken away; The associated data includes the names of drugs / items.

9. The multimodal guided emergency supplies positioning and medical order recording system according to claim 4, characterized in that, The display interface of the mobile interactive terminal includes a real-time voice-to-text area, a dynamic drug guidance card area, and a rescue timeline log area. The real-time speech-to-text area is used to display the recognized medical order information in real time. The dynamic drug guidance card area is used to pop up a floating card when the drug / item name is recognized. The floating card includes the standard name and specifications of the drug, recommended usage, key contraindications and warning information, as well as relevant information about the current patient. The relevant information about the current patient includes age and weight. The rescue timeline log area is used to display the operation log in reverse chronological order. The real-time speech-to-text area is located at the top of the display interface, the dynamic drug guidance card area is located in the center of the display interface, and the rescue timeline log area is located on the right side or bottom of the display interface.

10. The multimodal guided emergency supplies positioning and medical order recording system according to claim 9, characterized in that, The system also includes a digital display screen integrated into the visually guided medicine cabinet, which is used to dynamically display the identified medical order information in real time and to provide voice broadcast.