Multi-modal medical data processing method and system based on wearable device
By employing a multimodal data processing method based on wearable devices, we have achieved full-dimensional perception and compliant collection of medical data. This solves the problems of low efficiency, data fragmentation, poor real-time performance, and weak privacy protection in existing technologies, improves the accuracy and security of medical records, and provides real-time alerts and operation capture.
Patent Information
- Application Number
- CN202610021517.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-08
- Publication Date
- 2026-02-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing methods of recording medical information are inefficient, fragmented, lack real-time performance, have insufficient intelligence, and weak privacy protection, which affects the quality and safety of diagnosis and treatment.
By employing a wearable device-based multimodal data processing method, multimodal data streams are synchronously collected using a unified time reference. Local preprocessing and privacy policy fuzzing are performed using a deep reinforcement learning agent model. Combined with edge computing and cloud platforms, cross-modal temporal alignment and logical verification are performed to automatically generate structured medical records.
It achieves full-dimensional perception and compliant data collection, improves the integrity and security of medical data, ensures low-latency transmission of key information, enhances the accuracy and credibility of records, frees medical staff from paperwork, and provides a real-time early warning mechanism and precise capture of operational actions.
Smart Images

Figure CN121483472A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical technology, and in particular to a method and system for processing multimodal medical data based on wearable devices. Background Technology
[0002] Currently, medical information recording and processing mainly rely on electronic medical record (EMR) systems. Healthcare staff need to spend a lot of time manually entering text and checking forms. This method is not only inefficient, consuming more than 30% of healthcare staff's valuable work time, but it is also prone to errors or omissions due to human negligence, affecting the quality of diagnosis and treatment.
[0003] To address the efficiency issues of manual input, some medical institutions have introduced voice transcription devices. However, these devices typically only support single-modal speech-to-text functionality. They cannot integrate other crucial information generated during diagnosis and treatment, such as vital sign monitoring data, medical images, or surgical videos, leading to severe data fragmentation. For example, a doctor's verbal statement "the patient's wound is healing well" is isolated from abnormal heart rate data detected by their smart bracelet; the system cannot correlate the two for comprehensive analysis and early warning.
[0004] Secondly, the actual diagnosis and treatment process involves multi-dimensional information, including the doctor's verbal description (audio), the appearance of the affected area (video), vital signs (physiological parameters), and medical procedures (gestures). Current technologies typically collect and store this data independently from different devices: videos from surveillance cameras, vital signs recorded by monitors, and audio recordings of the doctor's verbal descriptions are often not interconnected. The system cannot accurately align and cross-reference the "doctor's verbal medication instructions" with the "injection actions captured by the wristband" on a timeline. This data fragmentation prevents the system from identifying potential medical risks (e.g., the doctor verbally prescribed medication but did not detect the actual action; or the patient's abnormal vital signs do not match the doctor's subjective description), making it difficult to form a complete chain of evidence.
[0005] Furthermore, existing technologies also face bottlenecks in terms of real-time data transmission and processing. Current mainstream 4G / 5G network communication latency is typically above 10 milliseconds, and bandwidth is limited, making it difficult to support real-time transmission of high-definition surgical videos and instant collaborative analysis with AI models. This latency prevents the implementation of immediate risk warnings and decision support in scenarios with extremely high real-time requirements, such as surgery.
[0006] More importantly, existing solutions have significant shortcomings in protecting patient privacy. The centralized cloud-based model for collecting and processing audio and video data lacks effective privacy control mechanisms, especially at the hardware level, posing a risk of sensitive information leakage. For example, it cannot automatically mask the faces of irrelevant people in the background while recording wounds, nor can it automatically identify and obscure medical records of non-patients. This makes the compliant deployment of wearable devices in hospitals face significant legal obstacles.
[0007] Finally, the existing systems generally have a low level of intelligence. Even if data collection is completed, they cannot automatically generate structured medical record reports that conform to medical standards. Medical staff still need to manually sort and review them a second time, which fails to fundamentally liberate medical staff from heavy paperwork.
[0008] Therefore, there is an urgent need in this field for an intelligent recording system and method that can achieve seamless, real-time fusion of multimodal data, ultra-low latency communication capabilities, ensure data privacy and security, and automatically generate structured medical reports, in order to solve many pain points in the existing technologies mentioned above, such as inefficiency, data fragmentation, real-time performance, privacy, and insufficient intelligence. Summary of the Invention
[0009] Therefore, the purpose of this invention is to provide a multimodal medical data processing method and system based on wearable devices, so as to fundamentally solve the problems of low efficiency, data fragmentation, poor real-time performance, insufficient intelligence and weak privacy protection in the existing medical record process.
[0010] A method for processing multimodal medical data based on a wearable device according to an embodiment of the present invention includes: By utilizing wearable device sets deployed on medical staff and vital sign monitoring devices deployed on patients via wireless connection, a multimodal data stream containing first-person perspective video, dialogue audio, vital signs from patients, and medical operation gestures from medical staff is synchronously collected based on a unified time reference. The multimodal data stream is then preprocessed locally using terminal devices within the wearable device set. The local preprocessing includes at least unified timestamp marking and sensitive area blurring based on a privacy policy. A multi-dimensional state vector containing data content features, visual scene features, and network environment features is constructed using the terminal devices in the wearable device group. The multi-dimensional state vector is then inferred using a deep reinforcement learning agent model deployed on the terminal devices. This generates network slice selection instructions and resource allocation strategies. The preprocessed multimodal data stream is then transmitted to the edge computing node through the established corresponding logical transmission channel. The edge computing node receives the multimodal data stream and, through a cross-modal temporal alignment algorithm based on a dual-stream asynchronous attention mechanism, maps the semantic features of the dialogue audio in the multimodal data stream to the spatiotemporal features of the first-view video within a preset dynamic backtracking window for matching, generating a fusion medical event containing a multidimensional evidence chain, and uploading the fusion medical event to the cloud platform. The cloud platform receives the fused medical events and, through its built-in medical-specific multimodal big data model, performs logical verification and deep reasoning based on the fused medical events and related knowledge retrieved from the medical knowledge graph, automatically generates a structured medical record report.
[0011] In addition, the multimodal medical data processing method based on a wearable device according to the above embodiments of the present invention may also have the following additional technical features: Furthermore, in the step of synchronously acquiring a multimodal data stream based on a unified time reference, including first-person perspective video, dialogue audio, vital signs from the patient, and medical operation gestures from medical staff, the step of acquiring dialogue audio includes: The system utilizes the built-in microphone array on the smart name tag in the wearable device group deployed among medical staff to collect multi-channel raw dialogue audio, calculate the arrival time difference or phase difference between each channel, and estimate the direction of sound source. An adaptive beamforming algorithm is used to point the main pickup beam directly in front of the patient area, while creating nulls in the side and rear areas to suppress environmental noise and conversations with others. Combining voice endpoint detection and segmented clustering algorithms, the system extracts targets using pre-stored doctor voiceprint features, separates audio from the wearer's direction that matches the doctor's voiceprint features into the doctor channel, and separates the remaining audio from the front beam into the patient channel, outputting a dual-channel independent dialogue audio stream.
[0012] Furthermore, in the step of synchronously acquiring a multimodal data stream based on a unified time reference, including first-person perspective video, dialogue audio, vital signs from the patient, and medical staff's gestures, the step of acquiring medical staff's gestures includes: The high-frequency inertial measurement unit built into the medical bracelet worn on the wrist of the wearable device group is used to collect triaxial acceleration and triaxial angular velocity data of wrist movement; The gravitational component is separated using a complementary filtering algorithm to extract linear acceleration, and an overlapping motion data window is generated using a time series slicing algorithm. The action data window is input into a deep convolutional long short-term memory network model to identify the corresponding medical operation gestures. The deep convolutional long short-term memory network model uses a multi-layer one-dimensional convolutional neural network to extract local waveform features generated by the operation action along the time axis, and uses the long short-term memory network to receive local waveform features and capture the long-term dependency relationship of the action data window, and outputs the posterior probability of each medical operation gesture category.
[0013] Furthermore, the step of blurring sensitive regions based on privacy policies in the local preprocessing of the multimodal data stream using terminal devices within the wearable device group includes: The terminal devices within the wearable device group use a target detection algorithm to scan the first-view video frame by frame to identify key medical entities and non-medical background areas in the video frames. The semantic region of interest is calculated by combining the head posture data of medical staff collected by the smart glasses in the wearable device group with the coordinates of the key medical entity, and a binary soft mask matrix is generated based on the semantic region of interest. Perform Gaussian blurring or pixelation on the background area of the original video frame, excluding the area covered by the binarized soft mask matrix. By using pixel-level weighted calculations, a clear semantic region of interest is synthesized with a blurred background region to generate a desensitized first-person perspective video.
[0014] Furthermore, the step of constructing a multi-dimensional state vector containing data content features, visual scene features, and network environment features using terminal devices within the wearable device group includes: The terminal device uses a built-in voice keyword detection algorithm to scan the dialogue audio stream within the current preset time window and extract voice keyword feature vectors that represent emergency or surgical instructions. Read vital sign data uploaded by the connected vital sign monitoring device, calculate the deviation of the current heart rate, blood oxygen or blood pressure values from the normal physiological range, and generate vital sign abnormality index features. The voice keyword feature vector is combined with the vital sign abnormality index feature to form data content features; The target detection model built into the terminal device is used to identify key frames of the first-person perspective video, detect whether there are preset key medical devices or specific anatomical parts, and output scene semantic labels containing the detected object categories, confidence levels and quantities as visual scene features. The wireless channel quality at the current camp location is measured in real time using the communication baseband module of the terminal device, and channel quality data including signal-to-noise ratio, reference signal received power and block error rate are obtained. Statistics on the current transmission queue buffer occupancy rate and round-trip latency of the backhaul link; The channel quality data and transmission statistics data are combined to form network environment characteristics; The data content features, visual scene features, and network environment features are numerically encoded and normalized, and then concatenated into a single one-dimensional tensor as a multi-dimensional state vector.
[0015] Furthermore, the step of using a deep reinforcement learning agent model deployed on the terminal device to reason about the multidimensional state vector, outputting network slice selection instructions and resource allocation strategies, and transmitting the preprocessed multimodal data stream to the edge computing node through the established corresponding logical transmission channel includes: The data content features and visual scene features in the multidimensional state vector are input into a pre-trained lightweight classification model to output the business priority level of the current scene. The business priority level includes at least the critical business level corresponding to remote surgery or emergency resuscitation, and the ordinary business level corresponding to routine care or log uploading. The service priority level and the network environment characteristics are input into a pre-trained deep reinforcement learning agent model using a deep Q network or a near-end policy optimization algorithm. The deep reinforcement learning agent model searches based on a reward function that aims to maximize the transmission success rate and minimize the latency of critical services, and outputs the optimal action instruction including the target slice identifier, reserved bandwidth value, routing path hop count and scheduling weight parameters. According to the optimal action instruction, a resource reconfiguration request is initiated to the software-defined network controller on the network side to establish a logical transmission channel, and the preprocessed multimodal data stream is transmitted to the edge computing node through the established corresponding logical transmission channel.
[0016] Furthermore, the step of generating a fused medical event containing a multidimensional chain of evidence by mapping the semantic features of the dialogue audio in the multimodal data stream to the spatiotemporal features of the first-person video in a joint embedding space for matching within a preset dynamic backtracking window using a cross-modal temporal alignment algorithm based on a dual-stream asynchronous attention mechanism includes: Maintain a multimodal circular buffer pool to continuously retain multimodal data streams for a preset duration; When a key medical instruction is detected based on the audio of the conversation, a dynamic backtracking window is activated in the circular buffer pool; A cross-modal dual-tower encoder structure is used to extract the textual semantic feature vector of key medical instructions and the spatiotemporal feature vector of video frame sequence in the first-person perspective video within the dynamic backtracking window, respectively. Map the text semantic feature vector and the spatiotemporal feature vector to the same joint embedding space, and calculate the multi-head attention weight or cosine similarity matrix of the video frame sequence pointed to by the text semantic feature vector; The video frame timestamp corresponding to the peak value in the multi-head attention weight or cosine similarity matrix is selected as the ground truth moment of the action, and the video frame, voice command and corresponding medical operation gesture corresponding to the ground truth moment are packaged into the fused medical event.
[0017] Furthermore, the step of automatically generating a structured medical record report by using a built-in medical-specific multimodal large model to perform logical verification and deep reasoning based on the fused medical events and related knowledge retrieved from the medical knowledge graph includes: Extract the current diagnosis and treatment entity, drug entity, and dosage entity from the fused medical event; Based on the unique identifier of the current patient, the stored historical medical record data is retrieved, and contraindication and drug interaction knowledge related to the drug entity are retrieved from the medical knowledge graph. The medical-specific multimodal large model is used to determine whether the diagnosis and treatment behavior entity conforms to the clinical pathway specification, and to infer whether the drug entity and dosage entity have medical logical conflicts with the historical medical record data, contraindication knowledge or drug interaction knowledge. If no conflict is detected, the fused medical event is converted into natural language text and filled into the corresponding chapter according to the preset electronic medical record template to generate the medical record report; If a conflict is detected, the corresponding diagnosis and treatment behavior will be highlighted as a risk in the generated medical record report, along with the reasoning basis and correction suggestions for the medical logic conflict. The generated feedback instruction containing risk warnings and correction suggestions will be sent to the terminal device through the logic transmission channel.
[0018] Furthermore, the method also includes: If the cloud platform or edge computing node identifies a critical value in the vital signs of the current patient or a contraindication risk in medical procedures during the generation of fusion medical events or deep reasoning, it generates a real-time warning instruction. The real-time warning command is sent to the smart glasses in the wearable device group via the high-priority downlink of the logical transmission channel; The smart glasses receive the real-time warning command and use eye-tracking technology to obtain the coordinates of the medical staff's current visual gaze point. The display layout is dynamically planned based on the coordinates of the visual gaze point, and the real-time warning command is presented in the unobstructed area of the augmented reality display area in the form of highlighted text or graphics until confirmation feedback is received from medical staff.
[0019] Another embodiment of the present invention aims to provide a multimodal medical data processing system based on wearable devices, the system comprising: The wearable device group is used to wirelessly connect to the vital sign monitoring equipment deployed on the patient. Based on a unified time reference, it synchronously collects multimodal data streams including first-person view video, dialogue audio, vital signs from the patient, and medical operation gestures from medical staff. The wearable device group uses terminal devices to perform local preprocessing on the multimodal data streams. The local preprocessing includes at least unified timestamp marking and sensitive area blurring based on privacy policies. It also constructs a multidimensional state vector containing data content features, visual scene features, and network environment features. The multidimensional state vector is inferred using a deep reinforcement learning agent model to output network slice selection instructions and resource allocation strategies. The preprocessed multimodal data stream is then transmitted to edge computing nodes through the established corresponding logical transmission channels. Edge computing nodes are used to receive the multimodal data stream and, through a cross-modal temporal alignment algorithm based on a dual-stream asynchronous attention mechanism, map the semantic features of the dialogue audio in the multimodal data stream and the spatiotemporal features of the first-view video to a joint embedding space for matching within a preset dynamic backtracking window, generating a fusion medical event containing a multidimensional evidence chain, and uploading the fusion medical event to the cloud platform. The cloud platform is used to receive the fused medical events and automatically generate structured medical record reports by performing logical verification and deep reasoning based on the fused medical events and related knowledge retrieved from the medical knowledge graph through a built-in medical-specific multimodal big model.
[0020] The multimodal medical data processing method based on wearable devices provided in this invention utilizes a group of wearable devices deployed on medical staff and a vital sign monitoring device deployed on the patient via wireless connection. Based on a unified time reference, it synchronously collects a multimodal data stream including first-person perspective video, dialogue audio, vital signs from the patient, and medical staff's gestures. The terminal device performs local preprocessing and privacy anonymization of the multimodal data stream, achieving full-dimensional perception and compliant collection of medical data from the source. This solves the information dimension loss caused by traditional single recording methods (such as pure voice or pure text), and the contradiction between comprehensive video recording and patient privacy protection. This significantly improves the integrity and security of raw medical data. Furthermore, by constructing a multi-dimensional state vector encompassing data content features, visual scene features, and network environment features using terminal devices, and employing a deep reinforcement learning agent model for inference to output network slice selection instructions and resource allocation strategies, an intelligent routing mechanism that drives network behavior based on data content is achieved. This ensures that critical medical information such as emergency instructions and critical vital signs can be transmitted through independent logical channels with millisecond-level latency, solving the problem that existing network transmission modes cannot meet the deterministic guarantee of high reliability and low latency communication in critical medical scenarios such as acute and critical illnesses and remote surgery. Moreover, by utilizing dual-stream asynchronous... The attention-based cross-modal temporal alignment algorithm matches the semantic features of dialogue audio with the spatiotemporal features of first-person video within a dynamic backtracking window, generating fused medical events containing multi-dimensional evidence chains. This enables precise timestamp correction and factual verification for asynchronous diagnostic and treatment behaviors such as "doing first, then speaking," solving the technical challenge of traditional recording methods where the lack of effective mutual verification mechanisms between voice and operation, and subjective description and objective facts, makes it difficult to verify the accuracy and temporality of recordings, significantly enhancing the credibility of medical data. Furthermore, by utilizing a built-in medical-specific multimodal large model on a cloud platform, combined with related knowledge retrieved from a medical knowledge graph, the fused medical events undergo logical verification and deep integration. The reasoning and automatic generation of structured medical record reports not only completely liberates medical staff from heavy paperwork, solving the deep-seated problems of low efficiency and error-proneness of manual data entry, as well as the difficulty in reusing unstructured data generated by clinical decision support systems, but also significantly improves diagnostic and treatment efficiency and data value. Furthermore, by identifying critical vital signs or operational contraindications at the edge or in the cloud, real-time warning instructions are generated and visualized in unobstructed areas using smart glasses based on eye-tracking technology. This establishes a real-time closed-loop intervention mechanism that does not interfere with normal operations, solving the technical pain points of traditional post-event quality control models being lagging behind and unable to effectively prevent medical errors at the moment they occur.By using an inertial measurement unit to collect high-frequency motion data on a medical wristband and inputting it into a deep convolutional long short-term memory network model to recognize specific medical hand gestures, precise and objective capture of key operations such as injection and pressure is achieved. This overcomes the technical limitations of relying solely on visual perception, which is easily obstructed, or verbal descriptions, making it difficult to capture subtle and rapid movements. It provides a new and reliable data dimension for the quantitative analysis and quality control of operations. This addresses the problems of inefficiency, data fragmentation, poor real-time performance, insufficient intelligence, and weak privacy protection in existing medical record processes. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating the multimodal medical data processing method based on wearable devices according to the first embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of the multimodal medical data processing system based on wearable devices according to the second embodiment of the present invention; The following detailed description of the embodiments will further illustrate the present invention in conjunction with the above-described accompanying drawings. Detailed Implementation
[0022] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.
[0023] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0025] Example 1 Please see Figure 1The diagram illustrates a multimodal medical data processing method based on wearable devices according to a first embodiment of the present invention. For ease of explanation, only the parts relevant to the embodiments of the present invention are shown. The multimodal medical data processing method based on wearable devices provided by the embodiments of the present invention includes: Step S10: Using a wearable device group deployed on medical staff and a vital sign monitoring device deployed on the patient via wireless connection, multimodal data streams including first-person perspective video, dialogue audio, vital signs from the patient and medical operation gestures from medical staff are collected synchronously based on a unified time reference. The multimodal data streams are preprocessed locally using terminal devices within the wearable device group. Local preprocessing includes at least unified timestamp marking and blurring of sensitive areas based on privacy policies. In one embodiment of the present invention, the method is primarily applied to a multimodal medical data processing system based on wearable devices. The system architecture mainly includes a wearable device group, edge computing nodes, and a cloud platform. The wearable device group includes, but is not limited to, smart glasses, smart badges, and medical wristbands. The edge computing nodes are typically MEC (Mobile Edge Computing) servers deployed in hospital data centers, directly connected to base stations (e.g., 6G base stations) via fiber optic cables, and possess GPU acceleration capabilities. The cloud platform is deployed in a private cloud or a trusted public cloud, running large-scale medical-specific multimodal models and maintaining a large-scale medical knowledge graph, responsible for complex logical reasoning, medical record generation, and long-term data mining. Furthermore, the smart glasses integrate at least a high-definition camera, an inertial measurement unit (IMU), an eye-tracking module, and an AR display; the smart badge integrates at least a microphone array; and the medical wristband integrates at least a PPG / ECG sensor and a 6-axis or 9-axis inertial measurement unit (IMU). It should also be noted that the wearable device group is connected to vital sign monitoring devices deployed on the patient via short-range wireless communication protocols (such as Bluetooth / NFC / WIFI). These vital sign monitoring devices include, but are not limited to, vital sign monitoring wristbands / patches, bedside monitors, or medical bracelets as described above. Worn by the patient, these devices continuously monitor vital sign data such as PPG / ECG / blood oxygen levels and broadcast the monitored vital sign data to the wearable device group (i.e., the subsequent terminal device) via wireless protocols such as Bluetooth Low Energy (BLE). Furthermore, in one embodiment of the invention, the wearable device group includes at least one terminal device responsible for data aggregation, local computation (preprocessing), short-range communication networking, and communication with external networks (specifically, for example, 6G). Since smart glasses are typically large, have relatively large batteries, and require video processing, and integrate high-performance chips and communication modules, in this specific embodiment, the smart glasses serve as the core of the terminal device, integrating a computing platform for data processing and a 6G communication module. It should be noted that in other embodiments of the present invention, the terminal device in the wearable device group can also be a smart badge (if the smart badge has sufficient computing power), or a separate computing box worn on the waist of medical staff (i.e., doctors), or even a terminal device in the doctor's pocket (such as a smartphone, smart tablet, etc.). The terminal device in the wearable device group can be set according to actual usage needs, and no specific limitation is made here.
[0026] Furthermore, before starting work, medical staff correctly wear a wearable device group consisting of smart glasses, smart name tags, and medical wristbands, ensuring all devices are powered on and ready to use. At this time, the smart glasses, acting as the computing and control core (i.e., the terminal device) within the wearable device group, automatically connect to the hospital's high-precision time synchronization server. The terminal device receives the standard time signal from the server and calibrates its local clock accordingly. Simultaneously, the terminal device establishes a local connection with the smart name tag, medical wristband, and vital sign monitoring device worn by the patient via a low-power wireless communication protocol, broadcasting a time synchronization signal as the master clock source. The smart name tag, medical wristband, and vital sign monitoring device, acting as slave devices, receive this time synchronization signal and calibrate their respective local clocks, ensuring that the timing start point and frequency of all devices are completely consistent, thus establishing a unified time reference covering all devices. Subsequently, throughout the work period, the devices periodically resynchronize to compensate for any minor clock drift, thereby maintaining a globally unified time reference. Therefore, the terminal device synchronizes its clock with the edge computing node via PTP (Precise Time Protocol) and transmits time to the smart badge, medical bracelet and vital sign monitoring device via low-latency Bluetooth broadcast, ensuring that all collected data frames (video frames, audio sampling points, IMU data packets, vital sign data and gesture data) are stamped with a unified timestamp and session ID.
[0027] The smart glasses deployed in the wearable device group for medical staff are configured as a visual acquisition and auxiliary command channel. They utilize the smart glasses' microphone to collect near-field voice commands and the smart glasses' digital signal processor (DSP) to monitor these commands in real time. When a specific voice command is detected, a hardware interrupt signal is sent to wake up the smart glasses' main processor and camera module to initiate video recording, thereby capturing first-person perspective video. Specifically, the smart glasses employ a tiered wake-up mechanism to balance power consumption. By default, the high-power main application processor and camera module are in deep sleep mode, while only the microwatt-level DSP remains in real-time monitoring mode. The DSP runs a lightweight keyword detection model for real-time monitoring of near-field voice commands. When medical staff are conducting ward rounds, consultations, or surgeries, and utter specific voice commands (such as "start recording" or "record wound"), the DSP sends a GPIO interrupt signal, waking up the smart glasses' main application processor and camera module within a preset time (e.g., 200ms) to initiate first-person perspective video recording. Simultaneously, the smart glasses send activation commands to the smart badge and medical wristband via low-latency Bluetooth broadcast. At this point, the devices begin to work in parallel to collect multimodal data streams. Specifically, the smart glasses activate their built-in high-definition camera, aligning its optical axis directly in front of the medical staff's line of sight, and continuously capture a first-person perspective video stream at a fixed frame rate. This video stream clearly records the patient's affected area, monitor readings, and surgical field of view within the medical staff's field of vision. At the same time, the smart glasses activate their built-in eye-tracking sensor to capture the direction of the medical staff's gaze in real time. Meanwhile, the smart badge activates its built-in microphone array to continuously collect sound wave signals from the environment and uses beamforming technology to directionally amplify the voice signals from the direction of the conversation between the medical staff and the patient, generating a clear audio stream of the conversation. Simultaneously, the medical bracelet worn on the wrist of the medical staff activates its photoplethysmography (PPG) sensor and electrode pads at the bottom, adhering to the skin to collect the medical staff's pulse waves and electrocardiogram (ECG) signals in real time, generating a vital signs data stream for the doctor. At the same time, the high-frequency inertial measurement unit inside the medical bracelet also begins to work, recording the acceleration and angular velocity changes of the wrist in three-dimensional space at a high-frequency sampling rate, generating a raw motion data stream reflecting the medical staff's gestures. Simultaneously, the vital signs monitoring device worn on the patient collects the patient's vital signs data (such as heart rate, blood oxygen, blood pressure, etc.) in real time. At the instant of collection, the processing units within each device immediately, according to the aforementioned unified time reference, assign a precise, unified timestamp to each frame of video image, each audio sample, each vital signs reading, and each gesture data, ensuring that all subsequent data are strictly aligned on the timeline.It should be noted that the data synchronously collected based on a unified time reference mainly includes first-person perspective video, dialogue audio, vital signs from the patient, and medical staff's gestures, enabling precise correlation between the medical staff's treatment actions and the patient's physiological responses, as well as the audio and video information collected on-site, on the same timeline, thus forming a complete diagnosis and treatment context. For example, if the medical staff's wristband detects an "adrenaline injection," then immediately after a preset time following the detected action, the patient's vital sign monitoring device can detect a "heart rate increase." This design ensures that the system can accurately capture the medical staff's operational behavior and monitor the patient's physiological response to the operation in real time (e.g., heart rate changes after drug injection), thereby constructing a complete chain of diagnostic and treatment evidence. Of course, in other embodiments of the invention, if it is also necessary to synchronously collect the medical staff's vital signs to monitor their physiological state and determine whether a critical state such as overwork has occurred, the medical staff's vital sign data can also be added to the synchronously collected multimodal data stream.
[0028] The smart badges deployed among wearable devices for healthcare workers are configured as the main channel for dialogue acquisition. They utilize a built-in microphone array to capture multi-channel raw audio and employ an adaptive beamforming algorithm to direct the main pickup beam towards the patient's area directly in front, while simultaneously creating nulls in the rear-side area to suppress ambient noise, thus obtaining high signal-to-noise ratio (SNR) audio for doctor-patient dialogue. Specifically, the pickup position determines the SNR and usability. While smart glasses' microphones are closer to the mouth and nose, suitable for doctors' dictation and instructions, patient voices in doctor-patient dialogues often come from directly in front or beside the bed. A fixed-direction array facing the chest is more likely to achieve stable "doctor-patient dual-channel dialogue" quality and avoids acoustic instability caused by rapid head movements. Furthermore, the form factor of smart glasses typically makes it difficult to accommodate large arrays and maintain high-performance noise reduction over extended periods; smart badges, on the other hand, offer more space and battery flexibility, accommodating multiple microphone arrays and continuously running acoustic algorithms. Therefore, the smart badges are configured as the main channel for dialogue acquisition, facilitating on-site compliance prompts and audits; and, when necessary, can also be used for desensitization in scenarios where only voice is captured without video. The smart glasses are configured as a dialogue acquisition auxiliary channel, mainly used for visual acquisition and as an auxiliary channel when doctors issue instructions.
[0029] Furthermore, in the above steps of synchronously acquiring multimodal data streams based on a unified time reference, including first-person perspective video, dialogue audio, vital signs from patients, and medical staff's gestures, the steps for acquiring dialogue audio include: The system utilizes the built-in microphone array on the smart name tag in the wearable device group deployed among medical staff to collect multi-channel raw dialogue audio, calculate the arrival time difference or phase difference between each channel, and estimate the direction of sound source. An adaptive beamforming algorithm is used to point the main pickup beam directly in front of the patient area, while creating nulls in the side and rear areas to suppress environmental noise and conversations with others. Combining voice endpoint detection and segmented clustering algorithms, the system extracts targets using pre-stored doctor voiceprint features, separates audio from the wearer's direction that matches the doctor's voiceprint features into the doctor channel, and separates the remaining audio from the front beam into the patient channel, outputting a dual-channel independent dialogue audio stream.
[0030] Specifically, the built-in microphone array of the smart work badge consists of multiple (e.g., four or more) omnidirectional microphones arranged in a preset geometric pattern (e.g., linear or circular), enabling it to receive sound wave signals from the surrounding environment from all directions. During the diagnosis and treatment process, the smart work badge collects raw environmental audio signals in real time, including the voices of medical staff, the voices of patients, and ambient noise. Then, the audio processing unit inside the smart work badge receives these multiple raw environmental audio signals and calculates the arrival time difference or phase difference between signals received from the same sound source by different microphones, estimating the angle of arrival of each independent sound source in the raw environmental audio relative to the smart work badge in real time. Specifically, this embodiment of the invention uses the Generalized Cross-Correlation-Phase Transform (GCC-PHAT) algorithm to calculate the arrival time difference between each microphone pair, thereby estimating the sound source angle. The specific process is as follows: For any two microphone signals, their cross-power spectrum is calculated. To eliminate the influence of reverberation and noise on the amplitude, only the phase information (i.e., PHAT weighting) is retained, and an inverse Fourier transform is performed to obtain the generalized cross-correlation function. Then, the peak value is found in the generalized cross-correlation function; the time delay corresponding to this peak value is the arrival time difference between the two microphones. Then, based on the geometric topology of the microphone array (with known microphone coordinates), a spatial steering vector is constructed. Using geometric trigonometry or maximum likelihood estimation, the observed time delay sets are mapped to the spatial azimuth (horizontal angle) and pitch angle of the sound source. In this way, the audio processing unit can construct a spatial distribution map of the current sound field, identifying which sounds come from directly in front, and which come from the side or rear. Then, to extract sounds from specific directions and suppress interference, the audio processing unit uses an MVDR (Minimum Variance Distortion-Free Response) beamforming algorithm to dynamically adjust the weights of each microphone. At this point, the main pickup beam is precisely pointed directly in front (i.e., the patient area), amplifying the signal in that direction; simultaneously, nulls are formed in the side and rear (the direction of ambient noise and conversations with bystanders) to physically suppress interference. Specifically, the audio processing unit presets the area directly in front (e.g., azimuth angle)... The area to the left and right is designated as the "patient pickup area," while the area to the right and left (the other angles) is designated as the "interference suppression area." The core of the algorithm is to solve for a set of optimal filter weight vectors. The optimization objective is to minimize the total noise power of the array output while ensuring that the signal gain from the direction of the "patient pickup area" is 1 (no distortion). Through this optimization, the MVDR algorithm automatically forms "null traps" (i.e., spatial notches with extremely low gain) in the direction of strong noise sources (such as the direction of the alarm sound from the side of the ECG monitor), thus physically shielding environmental noise. Although the beamforming audio suppresses environmental noise, it still mixes the voices of the doctor and the patient. At this point, the audio is separated into independent channels through the following steps: First, a VAD algorithm based on energy and zero-crossing rate is used to detect speech activity segments, dividing the continuous audio into short speech segments. Each speech segment is input into a lightweight voiceprint recognition model (such as TDNN or ResNet-34 architecture). The voiceprint recognition model maps variable-length speech segments to fixed-dimensional embedding vectors (such as 128-dimensional d-vectors or x-vectors), which represent the speaker's identity characteristics. Then, the cosine similarity between the embedding vector of the current segment and the pre-stored "doctor's registered voiceprint vector" is calculated. If the cosine similarity is greater than a preset threshold (e.g., 0.75) and the sound source direction estimation indicates that the sound source comes from the wearer's direction, the segment is determined to be the doctor's speech, and it is routed to channel A. If the similarity is lower than the threshold and the sound source direction estimation indicates that the sound source comes from directly in front (patient area), the segment is determined to be the patient's speech, and it is routed to channel B. If the similarity is low and the sound source comes from the side or rear, it is determined to be the speech of someone else, and it is either discarded or retained as background noise at a low volume. This ultimately outputs a clear dual-channel independent audio stream, enabling the extraction of a clear and separate doctor-patient dialogue stream from a noisy environment without human intervention, providing a high-quality input source for subsequent Automatic Speech Recognition (ASR) transcription.
[0031] Furthermore, in the aforementioned steps of synchronously acquiring multimodal data streams based on a unified time reference, including first-person perspective video, dialogue audio, vital signs from patients, and medical staff's gestures, the steps for acquiring medical staff's gestures include: The high-frequency inertial measurement unit built into the medical bracelet worn on the wrist of medical staff in the wearable device group was used to collect triaxial acceleration and triaxial angular velocity data of wrist movement; The gravitational component is separated using a complementary filtering algorithm to extract linear acceleration, and an overlapping motion data window is generated using a time series slicing algorithm. The motion data window is input into a deep convolutional long short-term memory network model to identify the corresponding medical and nursing operation gestures. The deep convolutional long short-term memory network model uses a multi-layer one-dimensional convolutional neural network to extract local waveform features generated by the operation action along the time axis, and uses the long short-term memory network to receive local waveform features and capture the long-term dependency relationship of the motion data window, and outputs the posterior probability of each medical operation gesture category.
[0032] Specifically, the medical wristband worn by healthcare workers is activated, and its built-in high-frequency inertial measurement unit (six-axis or nine-axis IMU) begins operation. This IMU includes at least a three-axis accelerometer and a three-axis gyroscope, capable of sensing minute wrist movements in three-dimensional space. The wristband's processing unit controls the IMU to continuously read data at a preset high-frequency sampling rate (e.g., 100Hz or higher), thereby acquiring a raw motion data stream reflecting changes in wrist acceleration and angular velocity in real time, ensuring the capture of rapid and precise details of healthcare operations. Since the raw acceleration signal contains gravitational components and linear acceleration components caused by motion, direct use would be affected by wrist posture. The system employs a complementary filtering algorithm or a Kalman filtering algorithm to fuse accelerometer and gyroscope data to estimate the wristband's posture (quaternions or Euler angles), then calculates the component of gravity in the sensor coordinate system, and subtracts this gravitational component from the raw acceleration to extract a pure linear acceleration signal. Next, in order to analyze and identify continuous actions from the continuous data stream, the processing unit of the medical bracelet uses a sliding window algorithm (e.g., window length of 2 seconds, step size of 1 second) to slice the data stream, thereby generating a series of overlapping action data windows.
[0033] Furthermore, to accurately capture specific operations performed by medical staff (such as injection, incision, and suturing), this system deploys a lightweight Deep Convolutional Long Short-Term Memory (DeepConvLSTM) model on the medical wristband. This DeepConvLSTM model includes a feature extraction layer, a temporal dependency layer, and a classification output layer. The feature extraction layer consists of four stacked 1D-CNN layers. The convolutional kernel slides along the time axis on the input matrix, automatically extracting local waveform features generated by specific operational actions (such as the micro-tremors when pushing the syringe or the abruptness when cutting the skin). These features include specific vibration frequencies when pushing the lever and abrupt acceleration slope when making a cut. After each convolutional operation, a batch normalization layer is applied to accelerate convergence, and a ReLU activation function is used to introduce non-linearity. This stage maps the original time-series data into a high-dimensional feature map. The feature map extracted by the feature extraction layer is then input into the temporal dependency layer (i.e., a 2-layer Long Short-Term Memory (LSTM) network). LSTM units utilize their internal gating mechanisms (forget gate, input gate, output gate) to maintain a memory state. They can capture long-range temporal dependencies in action sequences. For example, a complete "injection" action involves an irreversible temporal sequence: "positioning (stillness) – injection (smooth, continuous movement) – withdrawal (rapid reverse movement)." The temporal dependency layer can recognize this temporal pattern, rather than simply identifying instantaneous motion states. The output vector of the last time step of the temporal dependency layer is fed into the fully connected layer of the classification output layer, and finally normalized using the Softmax function to output the posterior probability of the current window belonging to each predefined medical gesture category (injection, pressure, disinfection, etc.).
[0034] At this point, the medical bracelet inputs the preprocessed motion data window into a lightweight deep convolutional long short-term memory network model for real-time inference. To avoid false alarms, the system sets a judgment threshold (e.g., 0.85). Only when the posterior probability of a certain type of gesture is greater than the judgment threshold and the state remains consistent within a consecutive preset number of sliding windows, does the medical bracelet finally determine and recognize the medical staff's gesture, and output the corresponding action label and timestamp.
[0035] Furthermore, before the data is transmitted or stored, the terminal device (i.e., the processing unit within the smart glasses) performs real-time local preprocessing on the converged multimodal data stream to perform sensitive area blurring based on a privacy policy. In one embodiment of the invention, the step of blurring sensitive areas based on a privacy policy during the local preprocessing of the multimodal data stream by the terminal device within the wearable device group includes: The terminal devices within the wearable device group use target detection algorithms to scan the first-person perspective video frame by frame to identify key medical entities and non-medical background areas in the video frames. The semantic region of interest is calculated by combining the head pose data of medical staff collected by the smart glasses in the wearable device group with the coordinates of key medical entities, and a binary soft mask matrix is generated based on the semantic region of interest. Perform Gaussian blurring or pixelation on the background area of the original video frame, excluding the area covered by the binarized soft mask matrix. By using pixel-level weighted calculations, a clear semantic region of interest is synthesized with a blurred background region to generate a desensitized first-person perspective video.
[0036] Specifically, the terminal device within the wearable device group (such as the processing unit of smart glasses) receives a first-person view video stream captured by a high-definition camera. The processing unit loads a pre-trained lightweight object detection model (such as SSDLite or YOLO-Nano architecture based on the MobileNetV3 backbone) to process the newly acquired first-person view video stream frame by frame. This object detection model is trained on a specific medical dataset and can identify key medical entities such as surgical wounds, medical instruments (such as scalpels and hemostatic forceps), anatomical locations, and sensitive privacy objects such as "faces" and "identity tags." At this point, the object detection model infers from each frame of the first-person view video image captured by the camera and outputs a set of bounding boxes, each containing a category label, confidence score, and coordinates.
[0037] To ensure that blurring doesn't obscure the doctor's focus, the system relies not only on visual detection but also on head pose data from the smart glasses' inertial measurement unit (IMU) to calculate semantic regions of interest (ROIs). Specifically, it first filters out bounding boxes labeled "key medical entities" with a confidence level greater than a threshold. Using the smart glasses' built-in IMU, it obtains the pitch and yaw angles of the medical staff's head. Combined with the camera's field of view, it calculates the projected coordinates of the medical staff's current gaze center on the image plane and generates a Gaussian weighted window centered on these coordinates. Then, it performs an intersection operation between all bounding box regions and this Gaussian weighted window, or assigns higher retention priority to medical entities near the gaze center, thus determining the final high-confidence ROIs. This ensures that even with multiple instruments in the frame, only the area the doctor is currently operating on or looking at maintains the highest clarity. Further, a binary soft mask matrix is generated based on the ROIs. Specifically, a single-channel matrix with the same resolution as the original video frame is created, initially all zeros. Then, the pixel regions corresponding to the ROIs are assigned a value of 1 in the matrix. To avoid abrupt visual transitions between blurred and sharp areas, a Gaussian blur operation is performed on the matrix. This smoothly transitions the mask values from 1 at the center of the semantically interesting region to 0 in the background region, forming a soft mask matrix. Further, the background region (the region with a mask value of 0) of the original video frame is subjected to Gaussian blur or pixelation. Then, using the soft mask matrix as the alpha channel (weight map), a pixel-level weighted calculation formula is used to synthesize the sharp focus region with the blurred background layer. The final product is a compliant video stream that retains only key medical details, background, and completely desensitized irrelevant personnel.
[0038] It's important to note that after the individual devices and vital sign monitoring devices within the wearable device group complete the collection of raw data, the data aggregation phase begins. Specifically, the medical wristband worn on the wrist of medical staff, the smart badge worn on the chest, and the vital sign monitoring device worn by the patient, acting as data acquisition slave devices, utilize their built-in low-power short-range wireless communication modules (such as Bluetooth Low Energy (BLE) modules or near-field magnetic induction communication modules) to establish a stable local data link with the smart glasses (i.e., the terminal device), which acts as the master control device. Each slave device locally caches the raw environmental audio data frames, raw motion data frames, and vital sign data packets it has just collected, and packages them according to a preset transmission protocol. To ensure that the timing of the data is not disrupted during transmission, the slave devices assign a precise transmission timestamp to each data packet before sending it, based on a unified time reference previously calibrated with the smart glasses. Subsequently, the medical wristband, smart badge, and vital sign monitoring device stream these timestamped data packets to the smart glasses in real time through the established local data link. The communication receiving module of the smart glasses continuously listens for and receives data streams from these slave devices. Upon receiving the data, the processor inside the smart glasses performs data unpacking and buffering operations. Based on the transmission timestamp in the data packet and the local timestamp of the first-person perspective video it captures, the processor aligns and merges data streams from different devices and modalities on the same timeline. Specifically, the processor uses interpolation or buffering queue mechanisms to compensate for the slight delays introduced by wireless transmission, ensuring that voice, gestures, vital signs, and video images occurring at the same moment are logically strictly synchronized. Only after completing the above data reception, unpacking, timeline alignment, and merging operations does the smart glasses treat it as a complete, unprocessed multimodal data stream and perform the aforementioned local preprocessing steps (such as secondary confirmation marking with a unified timestamp and blurring of sensitive areas based on privacy policies).
[0039] It should be specifically noted that the collection, storage, processing, and transmission of medical data (including but not limited to voice, video, vital signs, and medical gestures) involved in this embodiment of the invention strictly comply with the laws and regulations of relevant countries (such as the Personal Information Protection Law of the People's Republic of China). Before the actual deployment and operation of the system, explicit permission and authorization from relevant users (including medical personnel and patients) have been obtained through written informed consent forms, pop-up prompts, or biometric authorization. The system design incorporates strict privacy protection mechanisms (such as real-time data anonymization on the client side) to fully protect users' privacy rights and data security.
[0040] Step S20: Construct a multi-dimensional state vector containing data content features, visual scene features and network environment features using terminal devices in the wearable device group; use a deep reinforcement learning agent model deployed on the terminal devices to reason about the multi-dimensional state vector; output network slice selection instructions and resource allocation strategies; and transmit the preprocessed multimodal data stream to the edge computing node through the established corresponding logical transmission channel. In one embodiment of the present invention, after the multimodal data stream has completed local preprocessing on the terminal device, in order to ensure the real-time performance and reliability of critical medical data under limited network resources, the terminal device does not simply transmit data, but runs an intelligent routing proxy. This proxy dynamically controls 6G network slice resources by constructing a multidimensional state vector and using a deep reinforcement learning proxy model for inference. Specifically, the steps of constructing a multidimensional state vector containing data content features, visual scene features, and network environment features using terminal devices within the wearable device group include: The terminal device uses a built-in voice keyword detection algorithm to scan the dialogue audio stream within the current preset time window and extract voice keyword feature vectors that represent emergency or surgical instructions. Read vital sign data uploaded by the connected vital sign monitoring device, calculate the deviation of the current heart rate, blood oxygen or blood pressure values from the normal physiological range, and generate vital sign abnormality index features. The voice keyword feature vector is combined with the vital sign abnormality index feature to form the data content feature; The target detection model built into the terminal device is used to identify key frames of the first-person perspective video, detect whether there are preset key medical devices or specific anatomical parts, and output scene semantic labels containing the detected object categories, confidence levels and quantities as visual scene features. The wireless channel quality at the current camp location is measured in real time using the communication baseband module of the terminal device, and channel quality data including signal-to-noise ratio, reference signal received power and block error rate are obtained. Statistics on the current transmission queue buffer occupancy rate and round-trip latency of the backhaul link; Channel quality data and transmission statistics are combined to form network environment characteristics; The data content features, visual scene features, and network environment features are numerically encoded and normalized, and then concatenated into a single one-dimensional tensor as a multi-dimensional state vector.
[0041] Furthermore, the steps described above, which utilize a deep reinforcement learning agent model deployed on the terminal device to reason about the multidimensional state vector, output network slice selection instructions and resource allocation strategies, and transmit the preprocessed multimodal data stream to the edge computing node through the established corresponding logical transmission channel, include: The data content features and visual scene features in the multidimensional state vector are input into a pre-trained lightweight classification model to output the business priority level of the current scene. The business priority level includes at least the critical business level corresponding to remote surgery or emergency resuscitation, and the ordinary business level corresponding to routine care or log uploading. The service priority level and network environment characteristics are input into a pre-trained deep reinforcement learning agent model that uses a deep Q network or a near-end policy optimization algorithm. The deep reinforcement learning agent model searches based on a reward function that aims to maximize the transmission success rate and minimize the latency of critical services. The output includes the target slice identifier, reserved bandwidth value, routing path hop count and scheduling weight parameters as the optimal action instruction. Based on the optimal action instruction, a resource reconfiguration request is initiated to the software-defined network controller on the network side to establish a logical transmission channel, and the preprocessed multimodal data stream is transmitted to the edge computing node through the established corresponding logical transmission channel.
[0042] Specifically, the terminal device's processor aggregates data from the application layer and physical layer in real time to construct a multi-dimensional state vector. The terminal device uses its built-in digital signal processor to run a lightweight keyword detection algorithm, performing a sliding window scan of the dialogue audio stream within the most recent preset time (e.g., 3 seconds). If high-risk keywords such as defibrillation, intubation, or massive hemorrhage are detected, the corresponding one-hot encoded vector is set, generating a voice keyword feature vector. The terminal device reads heart rate (HR), blood oxygen (SpO2), and blood pressure (BP) data reported by the vital signs monitoring device via Bluetooth, calculates the Euclidean distance (deviation) between the current value and the preset normal range benchmark value, and generates a normalized vital signs abnormality index feature. For example, if the heart rate is 150 bpm (benchmark 100 bpm), the deviation is 0.5. The voice keyword feature vector and the vital signs abnormality index feature are then combined to form the data content feature. Furthermore, based on the output from the aforementioned lightweight object detection model, the terminal device identifies objects in the current video keyframe, outputs a list containing object categories (such as "scalpel" or "wound") and corresponding confidence scores, and encodes this list into a fixed-length scene semantic label vector. Further, the terminal device reads real-time network status via the baseband module API, including the signal-to-noise ratio (SNR), reference signal received power (RSRP), and block error rate (BLER) of the current slice, and reads the buffer status of the network protocol stack to obtain the current buffer occupancy and the round-trip time (RTT) of the most recent heartbeat packet, then combines these data to generate network environment features. Finally, the data content features, visual scene features, and network environment features are numerically encoded and normalized (mapped to the [0,1] interval), concatenated into a single one-dimensional tensor as a multi-dimensional state vector, serving as the state input of the deep reinforcement learning agent model at the corresponding time step.
[0043] Before inputting data into the deep reinforcement learning agent model, to reduce the search space and improve decision-making speed, the terminal device first needs to qualitatively classify the current business. This involves extracting "data content features" (such as keyword vectors for "defibrillation" and "rescue" identified by ASR, as well as normalized blood oxygen and heart rate deviation values) and "visual scene features" (such as confidence scores for "scalpel" and "wound" output by the object detection model) from multimodal data. These feature vectors are then input into a pre-trained lightweight classification model (such as a random forest or shallow multilayer perceptron MLP). This model, trained on a large amount of labeled data, can identify the current medical scenario type. The model outputs a discrete business priority level label. For example, detecting a "cardiopulmonary resuscitation" gesture or "defibrillation" command, or blood oxygen <90%, corresponds to a critical business level of remote surgery or emergency resuscitation. Detecting abnormal vital signs without emergency procedures corresponds to an important business level of expert consultation. Detecting routine ward round recordings or electronic medical record log uploads corresponds to a normal business level of routine nursing care or log uploads.
[0044] Furthermore, this embodiment employs a Deep Q-Network (DQN) or Proximal Policy Optimization (PPO) algorithm to construct a deep reinforcement learning agent model. This deep reinforcement learning agent model architecture includes an input layer (receiving the aforementioned service priority level and network environment characteristics), three fully connected hidden layers (128 neurons per layer, using ReLU activation), and an output layer. To guide the model in learning the optimal policy, a reward function is defined. for: in, It is an indicator function for successful transmission (1 for success, 0 for failure). This is the current transmission delay. It is the bandwidth cost of the selected slice. This is a weighting coefficient. For critical business levels, The weight of this factor will be significantly increased, forcing the model to seek the path with the lowest latency regardless of cost. That is, to minimize the latency of critical services while ensuring transmission success rate and taking bandwidth costs into account.
[0045] At this point, the deep reinforcement learning agent model searches for an optimal action instruction based on the input state (i.e., service priority level and network environment characteristics), aiming to maximize transmission success rate and minimize critical service latency. This optimal action instruction includes a specific combination of network parameters: target slice identifier (S-NSSAI), reserved bandwidth value (GBR), routing control, and scheduling weight (ARP). The target slice identifier indicates whether to switch to a URLLC (Ultra-Reliable Low Latency) slice or an eMBB (Extremely High-Bandwidth) slice. The reserved bandwidth value indicates the minimum guaranteed bit rate requested from the network (e.g., 50Mbps). Routing control indicates the hop limit of the routing path. The scheduling weight indicates the preemption priority during network congestion.
[0046] Furthermore, the terminal device encapsulates the optimal action instructions output by the deep reinforcement learning agent model into control signaling compliant with 3GPP standards. This signaling is sent to the SDN controller on the network side. The SDN controller parses the instructions and reconfigures the resources of the Radio Access Network (RAN) and Core Network (CN) within milliseconds. For example, it may preempt radio resource blocks (RBs) of low-priority users, allocate dedicated time-frequency resources for the current terminal, and establish a low-latency logical transmission channel directly to the edge computing node. Once the channel is established and confirmed, the terminal device immediately sends the preprocessed multimodal data stream through this dedicated channel, thereby ensuring deterministic transmission of critical medical data even in complex network environments.
[0047] In step S30, the edge computing node receives the multimodal data stream and, through a cross-modal temporal alignment algorithm based on a dual-stream asynchronous attention mechanism, maps the semantic features of the dialogue audio in the multimodal data stream to the spatiotemporal features of the first-person video in the preset dynamic backtracking window for matching, generating a fusion medical event containing a multidimensional evidence chain, and uploading the fusion medical event to the cloud platform. In one embodiment of the present invention, the step of generating a fused medical event containing a multidimensional chain of evidence by mapping the semantic features of dialogue audio in a multimodal data stream to the spatiotemporal features of a first-view video into a joint embedding space for matching within a preset dynamic backtracking window using a cross-modal temporal alignment algorithm based on a dual-stream asynchronous attention mechanism includes: Maintain a multimodal circular buffer pool to continuously retain multimodal data streams for a preset duration; When a key medical instruction is detected based on the audio of the conversation, a dynamic backtracking window is activated in the circular buffer pool; A cross-modal dual-tower encoder structure is used to extract the textual semantic feature vector of key medical instructions and the spatiotemporal feature vector of video frame sequence in the first-person perspective video within the dynamic backtracking window, respectively. Map the text semantic feature vector and the spatiotemporal feature vector to the same joint embedding space, and calculate the multi-head attention weight or cosine similarity matrix of the video frame sequence pointed to by the text semantic feature vector. The video frame timestamps corresponding to the peak values in the multi-head attention weight or cosine similarity matrix are selected as the ground truth moments of the actions, and the video frames, voice commands, and corresponding medical operation gestures corresponding to the ground truth moments are packaged into fused medical events.
[0048] Specifically, when multimodal data arrives at edge computing nodes, due to network jitter and device startup delays, there is often a time discrepancy of seconds between voice, video, and actions (e.g., a doctor performs the action before uttering the words). Edge computing nodes utilize a cross-modal temporal alignment algorithm based on a two-stream asynchronous attention mechanism for logical synchronization.
[0049] The edge computing nodes maintain a high-frequency, multimodal circular buffer that continuously receives and stores video streams, audio streams, vital signs data, and gesture data streams from terminal devices. The buffer employs a first-in, first-out (FIFO) strategy, always retaining historical data from the most recent preset time (e.g., 30 seconds), with older data automatically overwritten by newer data. When the speech recognition engine (ASR) detects a preset key medical instruction (such as "start suturing" or "inject adrenaline") from the real-time audio stream, it marks it as a semantic anchor. Then, using the timestamp of this semantic anchor as a reference, a dynamic asymmetric backtracking window is activated in the buffer. For example, it captures all video frame sequences and gesture data within a time range of 5 seconds before and 2 seconds after the instruction. This asymmetric design fully considers the operational habits of doctors who typically perform actions first and then provide verbal confirmation. To match speech and video and calculate the similarity between heterogeneous data, a cross-modal dual-tower encoder structure is employed to map the heterogeneous data to the same feature space. This structure includes a text tower and a video tower. The text tower uses a pre-trained BERT or BioBERT model as the encoder. The input is the key instruction text transcribed by the Automatic Speech Recognition (ASR) engine (e.g., "Start stitching"), and the output is a fixed-dimensional text semantic feature vector containing high-level semantic information of the instruction. The video tower uses a 3D convolutional neural network (e.g., C3D, I3D) or Video VisionTransformer as the encoder. The input is a sequence of video frames within a backtracking window, and the output is a sequence of spatiotemporal feature vectors for each frame. Furthermore, a fully connected layer maps the text semantic feature vector and the spatiotemporal feature vector sequence to the same joint embedding space, ensuring their comparability in the same dimension. The alignment relationship between the text and video is calculated using a multi-head attention mechanism or simple cosine similarity. Specifically, the dot product or cosine distance between the text semantic feature vector and each vector in the spatiotemporal feature vector sequence is calculated to obtain a set of similarity scores. This sequence actually constitutes a semantic response curve that changes over time. The peak of the semantic response curve represents the moment when the "stitching action" feature in the video is most significant. The timestamp corresponding to this peak is then determined as the physical truth moment when the action occurred, indicating that at this moment, the action features in the video frame and the semantic features of the voice command have the highest matching degree. Based on this, the audio and video streams are aligned, and the video clip, voice command, and corresponding medical operation gestures at this moment are packaged to generate a fused medical event, providing credible evidence for subsequent medical record generation.
[0050] Furthermore, in other embodiments of the present invention, the method further includes: When a deep convolutional long short-term memory network model recognizes a specific medical operation gesture and the confidence level exceeds a preset threshold, a trigger signal is sent to the smart glasses via a low-latency wireless protocol. After receiving a trigger signal, the smart glasses wake up the camera module to acquire short-term images of the current field of view, and use the target detection algorithm to identify the drug label or medical device specifications in the center of the field of view. The identification results are semantically associated with specific medical operation gestures to update the evidence chain of the fused medical event.
[0051] Specifically, the deep convolutional long short-term memory (LSTM) network model within the medical bracelet continuously infers from the action data window. When the LTM network model recognizes a specific medical operation gesture and the confidence level exceeds a preset threshold, the bracelet's Bluetooth communication module immediately generates a trigger signal containing the action type and timestamp, which is then broadcast or unicast to the paired smart glasses via Bluetooth Low Latency (BLE). The smart glasses' digital signal processor (DSP) is in Bluetooth listening mode. Upon receiving the trigger signal, the DSP immediately sends a hardware interrupt request (IRQ) to the main application processor (AP), waking up the AP and camera module from sleep mode within milliseconds. The awakened camera module immediately performs a short-term image acquisition of the current field of view (e.g., continuously capturing 5 high-resolution images or recording 1 second of video). Since doctors typically focus their gaze on the object being operated on when performing an operation (such as administering medication), the images acquired at this time are highly likely to contain the critical medical object. At this point, the smart glasses utilize a built-in lightweight object detection algorithm to quickly identify the captured images, detecting and recognizing text on drug labels, syringe specifications, or types of surgical instruments. Then, the visually recognized object information (such as "cefotaxime sodium" or "5ml syringe") is semantically correlated with the action information sent by the wristband ("injection action"). A structured record containing <timestamp, action: injection, object: ceftriaxone sodium, evidence: image file ID> is created and updated in the currently constructed evidence chain of the fused medical event. This mechanism ensures that even if the doctor does not verbally state the drug name, key information can be automatically completed through action and vision.
[0052] Finally, the edge computing nodes upload these structured, fused medical events to the cloud platform via a secure interface, allowing subsequent large-scale models to perform deep reasoning and report generation.
[0053] Step S40: The cloud platform receives the fused medical event and uses the built-in medical-specific multimodal big model to perform logical verification and deep reasoning based on the fused medical event and related knowledge retrieved from the medical knowledge graph, and automatically generates a structured medical record report. In one embodiment of the present invention, the step of automatically generating a structured medical record report by using a built-in medical-specific multimodal large model to perform logical verification and deep reasoning based on fused medical events and related knowledge retrieved from a medical knowledge graph includes: Extract the current diagnosis and treatment entity, drug entity, and dosage entity from the integrated medical event; Based on the unique identifier of the current patient, the stored historical medical record data is retrieved, and contraindication and drug interaction knowledge related to the drug entity are retrieved from the medical knowledge graph. Using a medical-specific multimodal large model, we can determine whether the entities involved in diagnosis and treatment conform to the clinical pathway guidelines, and infer whether the entities involved in drugs and dosages have any medical logical conflicts with historical medical records, contraindication knowledge, or drug interaction knowledge. If no conflict is detected, the integrated medical event will be converted into natural language text and filled into the corresponding chapter according to the preset electronic medical record template to generate a medical record report. If a conflict is detected, the corresponding diagnosis and treatment behavior will be highlighted as a risk in the generated medical record report, along with the reasoning basis for the medical logic conflict and correction suggestions. The generated feedback instruction containing risk warnings and correction suggestions will be sent to the terminal device through the logic transmission channel.
[0054] Specifically, the cloud platform receives fused medical events uploaded from edge computing nodes via a secure data interface. This fused medical event is a structured data packet containing a multidimensional chain of evidence. The cloud platform immediately activates its built-in medical-specific multimodal large-scale model. This model is not a general-purpose language model, but a specially constructed multimodal model specific to the medical vertical domain. This medical-specific multimodal large-scale model employs a composite architecture of "Transformer base + multimodal adapter + retrieval enhancement (RAG)". The base model uses a Transformer decoder architecture with hundreds of billions of parameters (such as a deeply customized version based on LLaMA or PaLM), and undergoes further pre-training with massive amounts of medical literature, treatment guidelines, and de-identified medical record data, enabling it to understand complex medical terminology and clinical logic. To understand video and vital sign data, the medical-specific multimodal large-scale model introduces a visual encoder (such as ViT) and a temporal signal encoder. By using a linear projection layer or Q-Former structure, the video feature vectors and vital sign waveform features uploaded by edge computing nodes are mapped to the same semantic embedding space as text, enabling medical-specific multimodal large models to read electrocardiograms and surgical videos as if they were text.
[0055] Meanwhile, the reasoning process of the medical-specific multimodal large model not only relies on the memory of model parameters, but also combines external medical knowledge graphs for real-time verification. The medical-specific multimodal large model first scans the text descriptions and structured tags in the fused medical events, and uses named entity recognition (NER) technology to extract diagnostic and treatment behavior entities that represent medical operation actions from the fused events. Specifically, these include diagnostic and treatment behavior entities that represent medical staff operations (e.g., "intravenous drip"), drug entities that represent the use of drugs (e.g., "cefotaxime sodium"), and dosage entities that represent the amount of drugs used (e.g., "two grams").
[0056] Next, the cloud platform enters the information retrieval and association phase. The cloud platform reads the unique identifier of the currently attending patient (such as hospital number or ID number) contained in the integrated medical event, and based on this, quickly retrieves the patient's historical medical record data from the cloud-stored medical database. This historical medical record data includes, but is not limited to: allergy history (e.g., whether allergic to penicillin), past medical history (e.g., whether there is liver or kidney dysfunction), long-term medication records, and recent laboratory test results. Simultaneously, the integrated medical platform connects to a vast medical knowledge graph, using the extracted drug entity as an index to retrieve detailed pharmacological knowledge of the drug, focusing on its contraindications (e.g., specific allergens, contraindications for specific underlying diseases) and interactions with other drugs.
[0057] Subsequently, a medical-specific multimodal large-scale model utilizes the aggregated information to perform logical verification and deep reasoning. On one hand, the model compares the extracted diagnostic and treatment entities with pre-defined standard clinical pathways to determine whether the current operational steps conform to the standard diagnostic and treatment procedures for the disease (e.g., checking whether a necessary skin test step has been omitted). On the other hand, the model performs complex medical logical reasoning, analyzing whether drug entities and dosage entities conflict with the patient's historical medical data (such as allergy history), retrieved contraindication knowledge, or drug interaction knowledge. The model uses causal reasoning algorithms to calculate whether medical logical conflicts exist. For example, the model might reason to determine whether the patient's allergy history includes the current drug, or whether the current drug might cause a severe adverse reaction with other drugs the patient is currently taking. It might even check the patient's liver and kidney function test results to determine whether the current dosage needs adjustment.
[0058] Finally, the cloud platform executes the report generation logic based on the inference results of the large model. If the inference does not detect any medical logic conflicts or violations, the cloud platform will call a preset electronic medical record template (such as a standard medical record template), and use the natural language generation capabilities of the large model to transform the structured information in the integrated medical event into fluent and professional medical terminology text (e.g., generating "100 mg of aspirin was administered orally at 10:30 AM"), automatically filling it into the corresponding section of the medical record template (such as the "Treatment Measures" or "Temporary Medical Orders" section), directly generating a standard medical record report, and pushing it to the doctor's workstation for final signature confirmation. It should be noted that each generated text is accompanied by an evidence pointer, which can be clicked to jump to the corresponding original audio clip, video keyframe, or vital sign trend graph.
[0059] Conversely, if a medical logic conflict is detected during the reasoning process—for example, the model discovers that the patient does indeed have a history of penicillin allergy—the cloud platform will trigger a risk warning mechanism while generating the report. It will prominently highlight the conflicting treatment actions in the report document (e.g., with a red background or in bold), and provide detailed reasoning for the medical logic conflict in the sidebar or annotation column (e.g., "The system detected that the patient has a history of penicillin allergy; there is a risk of cross-allergy when using cephalosporin antibiotics") and corrective suggestions generated based on the knowledge base (e.g., "It is recommended to switch to another class of antibiotics, such as azithromycin, or use with caution after a skin test") for the doctor's review. Simultaneously, the generated feedback command containing risk warnings and corrective suggestions will be sent to the terminal device via a logic transmission channel.
[0060] Furthermore, in other embodiments of the present invention, the method further includes: If the cloud platform or edge computing node identifies a critical value in the vital signs of the current patient or a contraindication risk in the medical procedure during the generation of integrated medical events or deep reasoning, it will generate a real-time warning instruction. Real-time warning commands are sent to the smart glasses within the wearable device group via a high-priority downlink through the logical transmission channel; The smart glasses receive real-time warning commands and use eye-tracking technology to obtain the coordinates of the medical staff's current visual gaze point. The display layout is dynamically planned based on the coordinates of the visual gaze point. Real-time warning instructions are presented in the unobstructed area of the augmented reality display area in the form of highlighted text or graphics until confirmation feedback is received from medical staff.
[0061] Specifically, edge computing nodes focus on handling tasks with the highest real-time requirements. For example, in the process of generating integrated medical events, they continuously analyze vital sign data streams from vital sign monitoring devices. Once a critical indicator such as heart rate or blood pressure suddenly exceeds a preset critical value range (e.g., heart rate drops instantly to below forty beats per minute), the edge computing node immediately determines that this is an emergency physiological abnormality. Without waiting for in-depth analysis from the cloud platform, it immediately generates a real-time warning command with the message "Critical: Bradycardia, please pay immediate attention!"
[0062] Meanwhile, during deep reasoning, the cloud platform utilizes its vast knowledge base to make more complex logical judgments. For example, when analyzing a medication event, it might discover a clear contraindication between the medication the healthcare professional is preparing to administer and the patient's allergy history. This risk is something the edge case cannot detect solely through real-time data. Once this contraindication risk is identified, the cloud platform immediately generates a real-time alert: "High Risk: Patient is allergic to this medication; use is prohibited!"
[0063] Once an alert command is generated, the edge computing node or cloud platform immediately invokes the communication module to send it out. To ensure the timely delivery of this life-saving information, the edge computing node or cloud platform does not use conventional data channels, but instead utilizes the high-priority downlink of the previously established logical transmission channel to transmit the real-time alert command back to the smart glasses worn by medical personnel at millisecond speeds.
[0064] Upon receiving the warning command, the smart glasses' communication module immediately activates its built-in eye-tracking module. This module uses an infrared camera to capture minute movements of the medical staff's eyes and, through a gaze estimation algorithm, calculates in real time the coordinates of the medical staff's current visual focus on the augmented reality display. This allows for precise determination of which area in the real-world scene the medical staff is currently focusing on (e.g., a surgical incision or a monitor screen). This step aims to determine where the doctor is currently looking, preventing the pop-up warning from obstructing the doctor's view of the patient's wound or critical surgical sites.
[0065] Furthermore, the smart glasses dynamically plan the display layout based on the calculated visual focus coordinates. Specifically, the smart glasses automatically avoid the area surrounding the medical staff's current visual focus, rendering warning information at the edge or side of the display screen, in blank areas, or above the doctor's field of vision (i.e., unobstructed areas). Simultaneously, the smart glasses project real-time warning instructions onto the medical staff's retina using high-contrast red highlighted text (such as "Critical: Braxton Hicks, please pay immediate attention!") or eye-catching graphic icons (such as flashing warning symbols). In this way, the smart glasses ensure that the warning information is captured by the medical staff's peripheral vision immediately, while strictly avoiding virtual information obstructing the medical staff's normal line of sight to observe key medical areas, thus ensuring the safe conduct of diagnostic and treatment procedures and achieving real-time closed-loop intervention for medical safety.
[0066] In one embodiment of the present invention, the method further includes: Edge computing nodes receive first-view video from multimodal data streams, extract the optimal keyframes using video quality assessment algorithms, and input them into a deep semantic segmentation neural network for semantic segmentation to extract wound boundary masks. By combining the binocular depth data from smart glasses, the absolute physical area and depth of the wound boundary mask are calculated, and a histogram is calculated in the HSV color space to quantify the composition ratio of swollen tissue, granulation tissue and necrotic tissue. The longitudinal trend analysis algorithm is used to calculate the Euclidean distance and slope of change between the current quantitative indicators and the historical wound feature vectors of the current patients, and to generate a dynamic assessment conclusion that includes the percentage of healing progress.
[0067] Specifically, the optimal keyframe of the wound image is first extracted using video quality assessment algorithms (such as Laplacian sharpness evaluation). This frame is then fed into an improved U-Net++ neural network. This network fuses deep semantic features with shallow texture features through dense skip connections, outputting a pixel-level wound boundary mask that accurately segments the wound area. Combined with the binocular depth data from the smart glasses, the pixel area of the wound boundary mask is mapped to the absolute area (square centimeters) and depth (millimeters) of the physical world. The segmented region is converted to the HSV color space, and its histogram distribution is calculated. Based on preset color thresholds (red represents granulation tissue, yellow represents necrotic tissue, and black represents necrosis), the percentage composition of each type of tissue is quantified. Simultaneously, the wound feature vector from the patient's last visit is automatically retrieved, and a longitudinal trend analysis algorithm is used to calculate the Euclidean distance and slope of change between the current indicator and historical snapshots, automatically generating a dynamic assessment conclusion including the percentage of healing progress (e.g., wound shrinkage of 15%, increase in granulation tissue).
[0068] In one embodiment of the present invention, the method further includes: At edge computing nodes, a local AI model is trained using locally stored and privacy-de-identified multimodal data. After the preset training period ends, only the encrypted model parameters of the local AI model are uploaded to the cloud platform; The cloud platform aggregates model parameters from multiple edge computing nodes to generate a global model, and then distributes the updated global model to each edge computing node.
[0069] Specifically, in each independent medical institution (i.e., each edge computing node) that has deployed the system of this embodiment, the system uses its multimodal data, which has undergone strict privacy anonymization processing and is accumulated during daily operation, to periodically train a locally deployed copy of the AI model. This training data is securely stored on the hospital's internal servers after ensuring that all personally identifiable information (such as faces, names, and ID numbers) has been completely removed or replaced. For example, the system uses locally stored voice dialogues, vital sign waveforms, and operation video clips related to specific disease diagnoses to further train the local disease recognition model, making it more adaptable to the hospital's specific data distribution and treatment habits.
[0070] Next, after a preset training cycle, such as once a week or when a certain amount of new data has accumulated locally, the local training process is paused. At this point, the local AI model has updated its internal weights and parameters by learning from the new data. Crucially, the system does not upload any raw or anonymized multimodal data to the external network. Instead, it precisely calculates the changes in the model parameters of the local model compared to the global model downloaded from the cloud in the previous round—that is, it calculates the parameter update amount or gradient. Then, the system only uploads this model parameter update, representing the local learning outcome, to the central federated learning platform deployed in the cloud, through a secure encrypted channel, such as using Transport Layer Security (TLS), and with end-to-end encryption of the content itself.
[0071] Furthermore, the cloud-based federated learning platform, acting as a coordinator, collects these encrypted model parameter updates from edge computing nodes across different medical institutions participating in the federated learning task. Once a sufficient number of updates have been collected (e.g., exceeding a preset number of participants), the cloud platform initiates a secure aggregation process. It employs an advanced aggregation algorithm, such as federated averaging, to perform a weighted average of all received parameter updates. During this process, the cloud platform can calculate a global model update that integrates the learning outcomes of all participants without decrypting the specific parameter content uploaded by any individual node. This design fundamentally eliminates the possibility of inferring the original training data from the model parameters, providing the highest level of privacy protection.
[0072] Finally, the cloud platform uses the calculated global model update amount to update its maintained central global model, thereby generating a new generation of global model that incorporates data experience from multiple institutions, boasts stronger performance, and better generalization capabilities. This iteratively optimized new global model is then securely sent back to every edge computing node participating in federated learning. Upon receiving the new global model, each edge computing node replaces its local old model with it.
[0073] In summary, the multimodal medical data processing method based on wearable devices in the above embodiments of the present invention utilizes a group of wearable devices deployed on medical personnel and a vital sign monitoring device deployed on the patient via wireless connection. Based on a unified time reference, it synchronously collects a multimodal data stream including first-person perspective video, dialogue audio, vital signs from the patient, and medical personnel's gestures. The terminal device performs local preprocessing and privacy anonymization of the multimodal data stream, achieving full-dimensional perception and compliant collection of medical data from the source. This solves the problem of missing information dimensions caused by traditional single recording methods (such as pure voice or pure text), as well as the challenges of comprehensive video recording and patient privacy protection. The contradiction between the two is greatly improved, enhancing the integrity and security of raw medical data. By constructing a multi-dimensional state vector containing data content features, visual scene features, and network environment features using terminal devices, and using a deep reinforcement learning agent model for reasoning to output network slice selection instructions and resource allocation strategies, an intelligent routing mechanism that drives network behavior based on data content is realized. This ensures that critical medical information such as emergency instructions and critical signs can be transmitted through independent logical channels with millisecond-level latency, solving the problem that existing network transmission modes cannot meet the deterministic guarantee of high reliability and low latency communication in critical medical scenarios such as acute and critical illnesses and remote surgery. Furthermore, by utilizing a dual-mode approach at edge computing nodes... The cross-modal temporal alignment algorithm based on the asynchronous attention mechanism matches the semantic features of dialogue audio with the spatiotemporal features of first-person video within a dynamic backtracking window, generating fused medical events containing multi-dimensional evidence chains. This enables precise timestamp correction and factual verification of asynchronous diagnostic and treatment behaviors such as "doing first, then speaking," solving the technical challenge of traditional recording methods lacking effective mutual verification mechanisms between voice and operation, and subjective description and objective facts, making it difficult to verify the accuracy and temporal sequence of recordings. This significantly enhances the credibility of medical data. Furthermore, by utilizing a built-in medical-specific multimodal large model on a cloud platform, combined with related knowledge retrieved from a medical knowledge graph, the fused medical events undergo logical verification and deep integration. This system, through inference, automatically generates structured medical record reports, not only completely freeing medical staff from tedious paperwork and solving the deep-seated problems of low efficiency and error-proneness of manual data entry, as well as the difficulty in reusing unstructured data generated by clinical decision support systems, but also significantly improving diagnostic efficiency and data value. Furthermore, by identifying critical vital signs or operational contraindications at the edge or in the cloud, it generates real-time warning instructions and visualizes them in unobstructed areas using smart glasses based on eye-tracking technology. This establishes a real-time closed-loop intervention mechanism that does not interfere with normal operations, addressing the technical pain points of traditional post-event quality control models being lagging behind and unable to effectively prevent medical errors at the moment they occur.By using an inertial measurement unit to collect high-frequency motion data on a medical wristband and inputting it into a deep convolutional long short-term memory network model to recognize specific medical hand gestures, precise and objective capture of key operations such as injection and pressure is achieved. This overcomes the technical limitations of relying solely on visual perception, which is easily obstructed, or verbal descriptions, making it difficult to capture subtle and rapid movements. It provides a new and reliable data dimension for the quantitative analysis and quality control of operations. This addresses the problems of inefficiency, data fragmentation, poor real-time performance, insufficient intelligence, and weak privacy protection in existing medical record processes.
[0074] Example 2 Please see Figure 2 This is a schematic diagram of a multimodal medical data processing system based on a wearable device according to a second embodiment of the present invention. For ease of explanation, only the parts related to the embodiment of the present invention are shown. The system includes: Wearable device group 11 is used to wirelessly connect to vital sign monitoring devices deployed on patients. Based on a unified time reference, it synchronously collects multimodal data streams including first-person perspective video, dialogue audio, vital signs from patients, and medical operation gestures from medical staff. The wearable device group uses terminal devices to perform local preprocessing on the multimodal data streams. The local preprocessing includes at least unified timestamp marking and sensitive area blurring based on privacy policies. It also constructs a multidimensional state vector containing data content features, visual scene features, and network environment features. It uses a deep reinforcement learning agent model to reason on the multidimensional state vectors, outputs network slice selection instructions and resource allocation strategies, and transmits the preprocessed multimodal data streams to edge computing nodes through the established corresponding logical transmission channels. Edge computing node 12 is used to receive the multimodal data stream and, through a cross-modal temporal alignment algorithm based on a dual-stream asynchronous attention mechanism, map the semantic features of the dialogue audio in the multimodal data stream and the spatiotemporal features of the first-view video to a joint embedding space for matching within a preset dynamic backtracking window, generate a fusion medical event containing a multidimensional evidence chain, and upload the fusion medical event to the cloud platform. The cloud platform 13 is used to receive the fused medical events and automatically generate structured medical record reports by performing logical verification and deep reasoning based on the fused medical events and related knowledge retrieved from the medical knowledge graph through a built-in medical-specific multimodal big model.
[0075] Furthermore, in one embodiment of the present invention, the wearable device group 11 includes: The sound source direction estimation unit is used to collect multi-channel raw dialogue audio using the built-in microphone array of the smart badge in the wearable device group, calculate the arrival time difference or phase difference between each channel, and perform sound source direction estimation. The beamforming unit is used to employ an adaptive beamforming algorithm to point the main pickup beam directly in front of the patient area, while simultaneously creating nulls in the side-rear area to suppress ambient noise and facilitate conversations with others. The dialogue audio output unit combines voice endpoint detection and segmented clustering algorithms, uses pre-stored doctor voiceprint features for target extraction, separates the audio that matches the doctor's voiceprint features and comes from the wearer's direction into the doctor channel, and separates the remaining audio from the beam directly in front into the patient channel, outputting dual-channel independent dialogue audio streams.
[0076] Furthermore, in one embodiment of the present invention, the wearable device group 11 includes: The data acquisition unit is used to collect triaxial acceleration and triaxial angular velocity data of wrist movement using the high-frequency inertial measurement unit built into the medical bracelet worn on the wrist of the medical staff in the wearable device group. The motion data window generation unit is used to separate the gravity component using a complementary filtering algorithm to extract linear acceleration, and to generate overlapping motion data windows using a time series slicing algorithm. The medical and nursing operation gesture recognition unit is used to input the action data window into a deep convolutional long short-term memory network model to recognize the corresponding medical and nursing operation gestures. The deep convolutional long short-term memory network model uses a multi-layer one-dimensional convolutional neural network to extract local waveform features generated by the operation action along the time axis, and uses the long short-term memory network to receive the local waveform features and capture the long-term dependency relationship of the action data window, and outputs the posterior probability of each medical operation gesture category.
[0077] Furthermore, in one embodiment of the present invention, the wearable device group 11 includes: The video frame recognition unit is used to scan the first-view video frame by frame using a target detection algorithm on the terminal device in the wearable device group to identify key medical entities and non-medical background areas in the video frame. The binarized soft mask matrix generation unit is used to combine the head posture data of medical staff collected by the smart glasses in the wearable device group with the coordinates of the key medical entity to calculate the semantic region of interest, and generate a binarized soft mask matrix based on the semantic region of interest. The image processing unit is used to perform Gaussian blurring or pixelation on the background area of the original video frame other than the area covered by the binarized soft mask matrix. The desensitized video generation unit is used to synthesize a clear semantic region of interest with a blurred background region through pixel-level weighted calculation to generate a desensitized first-person perspective video.
[0078] Furthermore, in one embodiment of the present invention, the wearable device group 11 includes: The voice keyword feature vector extraction unit is used to scan the dialogue audio stream within the current preset time window using the voice keyword detection algorithm built into the terminal device, and extract the voice keyword feature vectors that represent emergency or surgical instructions. The vital signs abnormality index feature generation unit is used to read vital signs data uploaded by the connected vital signs monitoring device, calculate the deviation of the current heart rate, blood oxygen or blood pressure values from the normal physiological range, and generate vital signs abnormality index features. The data content feature generation unit is used to combine the voice keyword feature vector with the vital sign abnormality index feature to form data content features. The visual scene feature generation unit is used to identify key frames of first-person perspective video using the target detection model built into the terminal device, detect whether there are preset key medical devices or specific anatomical parts, and output scene semantic labels containing the detected object categories, confidence levels and quantities as visual scene features. The channel quality data acquisition unit is used to measure the wireless channel quality at the current camping location in real time using the communication baseband module of the terminal device, and to acquire channel quality data including signal-to-noise ratio, reference signal received power and block error rate. The transmission statistics unit is used to calculate the current transmission queue buffer occupancy rate and the round-trip delay of the backhaul link. A network environment feature generation unit is used to combine the channel quality data and transmission statistics data to form network environment features; The multidimensional state vector generation unit is used to numerically encode and normalize the data content features, the visual scene features, and the network environment features, and concatenate them into a single one-dimensional tensor as a multidimensional state vector.
[0079] Furthermore, in one embodiment of the present invention, the wearable device group 11 includes: The business priority level generation unit is used to input the data content features and visual scene features in the multidimensional state vector into a pre-trained lightweight classification model to output the business priority level of the current scene. The business priority level includes at least the critical business level corresponding to remote surgery or emergency resuscitation, and the ordinary business level corresponding to routine care or log uploading. The optimal action instruction generation unit is used to input the service priority level and the network environment characteristics into a pre-trained deep reinforcement learning agent model using a deep Q network or a near-end policy optimization algorithm. The deep reinforcement learning agent model searches based on a reward function that aims to maximize the transmission success rate and minimize the latency of critical services, and outputs the optimal action instruction including the target slice identifier, reserved bandwidth value, routing path hop count and scheduling weight parameters. The logical transmission channel establishment unit is used to initiate a resource reconfiguration request to the software-defined network controller on the network side according to the optimal action instruction to establish a logical transmission channel, and transmit the preprocessed multimodal data stream to the edge computing node through the established corresponding logical transmission channel.
[0080] Furthermore, in one embodiment of the present invention, the edge computing node 12 includes: The multimodal ring buffer maintenance unit is used to maintain a multimodal ring buffer and continuously retain multimodal data streams for a preset duration. A dynamic backtracking window activation unit is used to activate a dynamic backtracking window in the annular buffer pool when a key medical instruction is detected based on the dialogue audio. The feature vector extraction unit is used to extract the textual semantic feature vector of key medical instructions and the spatiotemporal feature vector of video frame sequence in the first-view video within the dynamic backtracking window, respectively, using a cross-modal dual-tower encoder structure. The semantic response calculation unit is used to map the text semantic feature vector and the spatiotemporal feature vector to the same joint embedding space, and calculate the multi-head attention weight or cosine similarity matrix of the text semantic feature vector pointing to the video frame sequence. The integrated medical event generation unit is used to select the video frame timestamp corresponding to the peak in the multi-head attention weight or cosine similarity matrix as the truth time of the action, and package the video frame, voice command and corresponding medical operation gesture corresponding to the truth time into the integrated medical event.
[0081] Furthermore, in one embodiment of the present invention, the cloud platform 13 includes: An entity extraction unit is used to extract the current diagnosis and treatment behavior entity, drug entity, and dosage entity from the fused medical event; The knowledge data acquisition unit is used to retrieve stored historical medical record data based on the unique identifier of the currently visiting patient, and to retrieve contraindication knowledge and drug interaction knowledge related to the drug entity from the medical knowledge graph; The judgment and reasoning unit is used to use the medical-specific multimodal large model to determine whether the diagnosis and treatment behavior entity conforms to the clinical pathway specifications, and to reason whether the drug entity and dosage entity have medical logical conflicts with the historical medical record data, contraindication knowledge or drug interaction knowledge. The first medical record report generation unit is used to convert the fused medical event into natural language text and fill it into the corresponding chapter according to the preset electronic medical record template if no conflict is detected, thereby generating the medical record report. The second medical record report generation unit is used to, if a conflict is detected, highlight the corresponding diagnosis and treatment behavior in the generated medical record report with a risk highlight, and attach the reasoning basis and correction suggestions for the medical logic conflict, and send the generated feedback instruction containing risk warnings and correction suggestions to the terminal device through the logic transmission channel.
[0082] Furthermore, in one embodiment of the present invention, the cloud platform 13 or edge computing node 12 is also used to generate a real-time warning instruction when, during the process of generating fusion medical events or performing deep reasoning, it is identified that the vital signs of the current patient are critical or that there is a risk of contraindication to medical operations; and to send the real-time warning instruction to the smart glasses in the wearable device group through the high-priority downlink of the logical transmission channel. The smart glasses in the wearable device group 11 are also used to receive the real-time warning command, obtain the coordinates of the medical staff's current visual gaze point using eye-tracking technology, and dynamically plan the display layout according to the visual gaze point coordinates, presenting the real-time warning command in the unobstructed area of the augmented reality display area in the form of highlighted text or graphics, until the medical staff's confirmation feedback is received.
[0083] The multimodal medical data processing system based on wearable devices provided in this embodiment of the invention has the same implementation principle and technical effects as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment.
[0084] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0085] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.
Claims
1. A method for processing multimodal medical data based on wearable devices, characterized in that, The method includes: By utilizing wearable device sets deployed on medical staff and vital sign monitoring devices deployed on patients via wireless connection, a multimodal data stream containing first-person perspective video, dialogue audio, vital signs from patients, and medical operation gestures from medical staff is synchronously collected based on a unified time reference. The multimodal data stream is then preprocessed locally using terminal devices within the wearable device set. The local preprocessing includes at least unified timestamp marking and sensitive area blurring based on a privacy policy. A multi-dimensional state vector containing data content features, visual scene features, and network environment features is constructed using the terminal devices in the wearable device group. The multi-dimensional state vector is then inferred using a deep reinforcement learning agent model deployed on the terminal devices. This generates network slice selection instructions and resource allocation strategies. The preprocessed multimodal data stream is then transmitted to the edge computing node through the established corresponding logical transmission channel. The edge computing node receives the multimodal data stream and, through a cross-modal temporal alignment algorithm based on a dual-stream asynchronous attention mechanism, maps the semantic features of the dialogue audio in the multimodal data stream to the spatiotemporal features of the first-view video within a preset dynamic backtracking window for matching, generating a fusion medical event containing a multidimensional evidence chain, and uploading the fusion medical event to the cloud platform. The cloud platform receives the fused medical events and, through its built-in medical-specific multimodal big data model, performs logical verification and deep reasoning based on the fused medical events and related knowledge retrieved from the medical knowledge graph, automatically generates a structured medical record report.
2. The multimodal medical data processing method based on wearable devices according to claim 1, characterized in that, The step of synchronously acquiring a multimodal data stream based on a unified time reference, including first-person perspective video, dialogue audio, vital signs from the patient, and medical staff's gestures, includes the following step: The system utilizes the built-in microphone array on the smart name tag in the wearable device group deployed among medical staff to collect multi-channel raw dialogue audio, calculate the arrival time difference or phase difference between each channel, and estimate the direction of sound source. An adaptive beamforming algorithm is used to point the main pickup beam directly in front of the patient area, while creating nulls in the side and rear areas to suppress environmental noise and conversations with others. Combining voice endpoint detection and segmented clustering algorithms, the system extracts targets using pre-stored doctor voiceprint features, separates audio from the wearer's direction that matches the doctor's voiceprint features into the doctor channel, and separates the remaining audio from the front beam into the patient channel, outputting a dual-channel independent dialogue audio stream.
3. The multimodal medical data processing method based on wearable devices according to claim 1, characterized in that, The step of synchronously acquiring a multimodal data stream based on a unified time reference, including first-person perspective video, dialogue audio, vital signs from the patient, and medical staff's gestures, includes the following step: The high-frequency inertial measurement unit built into the medical bracelet worn on the wrist of the wearable device group is used to collect triaxial acceleration and triaxial angular velocity data of wrist movement; The gravitational component is separated using a complementary filtering algorithm to extract linear acceleration, and an overlapping motion data window is generated using a time series slicing algorithm. The action data window is input into a deep convolutional long short-term memory network model to identify the corresponding medical operation gestures. The deep convolutional long short-term memory network model uses a multi-layer one-dimensional convolutional neural network to extract local waveform features generated by the operation action along the time axis, and uses the long short-term memory network to receive local waveform features and capture the long-term dependency relationship of the action data window, and outputs the posterior probability of each medical operation gesture category.
4. The multimodal medical data processing method based on wearable devices according to claim 1, characterized in that, The step of blurring sensitive regions based on privacy policies in the local preprocessing of the multimodal data stream using terminal devices within the wearable device group includes: The terminal devices within the wearable device group use a target detection algorithm to scan the first-view video frame by frame to identify key medical entities and non-medical background areas in the video frames. The semantic region of interest is calculated by combining the head posture data of medical staff collected by the smart glasses in the wearable device group with the coordinates of the key medical entity, and a binary soft mask matrix is generated based on the semantic region of interest. Perform Gaussian blurring or pixelation on the background area of the original video frame, excluding the area covered by the binarized soft mask matrix. By using pixel-level weighted calculations, a clear semantic region of interest is synthesized with a blurred background region to generate a desensitized first-person perspective video.
5. The multimodal medical data processing method based on wearable devices according to claim 1, characterized in that, The step of constructing a multidimensional state vector containing data content features, visual scene features, and network environment features using terminal devices within the wearable device group includes: The terminal device uses a built-in voice keyword detection algorithm to scan the dialogue audio stream within the current preset time window and extract voice keyword feature vectors that represent emergency or surgical instructions. Read vital sign data uploaded by the connected vital sign monitoring device, calculate the deviation of the current heart rate, blood oxygen or blood pressure values from the normal physiological range, and generate vital sign abnormality index features. The voice keyword feature vector is combined with the vital sign abnormality index feature to form data content features; The target detection model built into the terminal device is used to identify key frames of the first-person perspective video, detect whether there are preset key medical devices or specific anatomical parts, and output scene semantic labels containing the detected object categories, confidence levels and quantities as visual scene features. The wireless channel quality at the current camp location is measured in real time using the communication baseband module of the terminal device, and channel quality data including signal-to-noise ratio, reference signal received power and block error rate are obtained. Statistics on the current transmission queue buffer occupancy rate and round-trip latency of the backhaul link; The channel quality data and transmission statistics data are combined to form network environment characteristics; The data content features, visual scene features, and network environment features are numerically encoded and normalized, and then concatenated into a single one-dimensional tensor as a multi-dimensional state vector.
6. The multimodal medical data processing method based on wearable devices according to claim 1, characterized in that, The steps of using a deep reinforcement learning agent model deployed on a terminal device to reason about the multidimensional state vector, outputting network slice selection instructions and resource allocation strategies, and transmitting the preprocessed multimodal data stream to the edge computing node through the established corresponding logical transmission channel include: The data content features and visual scene features in the multidimensional state vector are input into a pre-trained lightweight classification model to output the business priority level of the current scene. The business priority level includes at least the critical business level corresponding to remote surgery or emergency resuscitation, and the ordinary business level corresponding to routine care or log uploading. The service priority level and the network environment characteristics are input into a pre-trained deep reinforcement learning agent model using a deep Q network or a near-end policy optimization algorithm. The deep reinforcement learning agent model searches based on a reward function that aims to maximize the transmission success rate and minimize the latency of critical services, and outputs the optimal action instruction including the target slice identifier, reserved bandwidth value, routing path hop count and scheduling weight parameters. According to the optimal action instruction, a resource reconfiguration request is initiated to the software-defined network controller on the network side to establish a logical transmission channel, and the preprocessed multimodal data stream is transmitted to the edge computing node through the established corresponding logical transmission channel.
7. The multimodal medical data processing method based on wearable devices according to claim 1, characterized in that, The step of generating a fused medical event containing a multidimensional evidence chain by using a cross-modal temporal alignment algorithm based on a dual-stream asynchronous attention mechanism to map the semantic features of dialogue audio and the spatiotemporal features of first-person video in the multimodal data stream to a joint embedding space for matching within a preset dynamic backtracking window includes: Maintain a multimodal circular buffer pool to continuously retain multimodal data streams for a preset duration; When a key medical instruction is detected based on the audio of the conversation, a dynamic backtracking window is activated in the circular buffer pool; A cross-modal dual-tower encoder structure is used to extract the textual semantic feature vector of key medical instructions and the spatiotemporal feature vector of video frame sequence in the first-person perspective video within the dynamic backtracking window, respectively. Map the text semantic feature vector and the spatiotemporal feature vector to the same joint embedding space, and calculate the multi-head attention weight or cosine similarity matrix of the video frame sequence pointed to by the text semantic feature vector; The video frame timestamp corresponding to the peak value in the multi-head attention weight or cosine similarity matrix is selected as the ground truth moment of the action, and the video frame, voice command and corresponding medical operation gesture corresponding to the ground truth moment are packaged into the fused medical event.
8. The multimodal medical data processing method based on wearable devices according to claim 1, characterized in that, The step of automatically generating a structured medical record report by using a built-in medical-specific multimodal large model to perform logical verification and deep reasoning based on the fused medical events and related knowledge retrieved from the medical knowledge graph includes: Extract the current diagnosis and treatment entity, drug entity, and dosage entity from the fused medical event; Based on the unique identifier of the current patient, the stored historical medical record data is retrieved, and contraindication and drug interaction knowledge related to the drug entity are retrieved from the medical knowledge graph. The medical-specific multimodal large model is used to determine whether the diagnosis and treatment behavior entity conforms to the clinical pathway specification, and to infer whether the drug entity and dosage entity have medical logical conflicts with the historical medical record data, contraindication knowledge or drug interaction knowledge. If no conflict is detected, the fused medical event is converted into natural language text and filled into the corresponding chapter according to the preset electronic medical record template to generate the medical record report. If a conflict is detected, the corresponding diagnosis and treatment behavior will be highlighted as a risk in the generated medical record report, along with the reasoning basis and correction suggestions for the medical logic conflict. The generated feedback instruction containing risk warnings and correction suggestions will be sent to the terminal device through the logic transmission channel.
9. The multimodal medical data processing method based on wearable devices according to claim 1, characterized in that, The method further includes: If the cloud platform or edge computing node identifies a critical value in the vital signs of the current patient or a contraindication risk in medical procedures during the generation of fusion medical events or deep reasoning, it generates a real-time warning instruction. The real-time warning command is sent to the smart glasses in the wearable device group via the high-priority downlink of the logical transmission channel; The smart glasses receive the real-time warning command and use eye-tracking technology to obtain the coordinates of the medical staff's current visual gaze point. The display layout is dynamically planned based on the coordinates of the visual gaze point, and the real-time warning command is presented in the unobstructed area of the augmented reality display area in the form of highlighted text or graphics until confirmation feedback is received from medical staff.
10. A multimodal medical data processing system based on wearable devices, characterized in that, The system includes: The wearable device group is used to wirelessly connect to the vital sign monitoring equipment deployed on the patient. Based on a unified time reference, it synchronously collects multimodal data streams including first-person view video, dialogue audio, vital signs from the patient, and medical operation gestures from medical staff. The wearable device group uses terminal devices to perform local preprocessing on the multimodal data streams. The local preprocessing includes at least unified timestamp marking and sensitive area blurring based on privacy policies. It also constructs a multidimensional state vector containing data content features, visual scene features, and network environment features. The multidimensional state vector is inferred using a deep reinforcement learning agent model to output network slice selection instructions and resource allocation strategies. The preprocessed multimodal data stream is then transmitted to edge computing nodes through the established corresponding logical transmission channels. Edge computing nodes are used to receive the multimodal data stream and, through a cross-modal temporal alignment algorithm based on a dual-stream asynchronous attention mechanism, map the semantic features of the dialogue audio in the multimodal data stream and the spatiotemporal features of the first-view video to a joint embedding space for matching within a preset dynamic backtracking window, generating a fusion medical event containing a multidimensional evidence chain, and uploading the fusion medical event to the cloud platform. The cloud platform is used to receive the fused medical events and automatically generate structured medical record reports by performing logical verification and deep reasoning based on the fused medical events and related knowledge retrieved from the medical knowledge graph through a built-in medical-specific multimodal big model.
Citation Information
Cited By
Round live broadcast method based on chest card recorder system and related device
CN122160531A
Artificial intelligence-based neonatal image record detection method, system, device and medium
CN122224397A