Medical data intelligent acquisition device based on visual large model and acquisition method thereof
Through the intelligent medical data acquisition device based on visual big models, the problems of limited data acquisition, inaccurate manual recording and untimely monitoring of medical equipment are solved, and efficient and accurate automated data acquisition and real-time monitoring are achieved, which is suitable for a variety of medical equipment.
Patent Information
- Application Number
- CN202510406237.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-22
AI Technical Summary
The lack of data output interface for medical equipment leads to limited data collection, inaccurate manual recording, untimely monitoring and difficult data integration.
The intelligent medical data acquisition device based on visual big model is adopted, including hardware terminals and backend servers, and the visual big model analysis engine is used to analyze medical image information, combined with a dual gimbal system and a multi-function fixing device to realize automated data acquisition and real-time monitoring.
It improves the efficiency and accuracy of medical data collection, realizes high-frequency and high-precision data recording without manual intervention, promptly detects abnormal situations, and supports the flexible application of a variety of medical equipment.
Smart Images

Figure CN120356591A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and particularly to an intelligent medical data acquisition device based on a vision large model and an acquisition method thereof. Background Art
[0002] In a medical environment, many traditional medical devices (such as extracorporeal membrane oxygenation ECMO, cardiopulmonary bypass machine, etc.) lack data output interfaces, making it difficult to achieve automatic data acquisition and recording. Currently, the following main problems exist in medical data acquisition:
[0003] 1) Limited device data acquisition: Many medical devices do not have data output interfaces, and data needs to be manually transcribed. Some monitoring contents are not electronic devices and require manual observation, such as fluid infusion, urine volume, and blood collection buckets.
[0004] 2) Inaccurate manual recording: Data such as fluid infusion volume and urine volume need to be manually recorded regularly. Due to the long recording time interval, data is easily missed or misrecorded, and the dynamic changes during the period cannot be reflected.
[0005] 3) Delayed device monitoring: Device alarms or abnormalities may not be detected in a timely manner. For example, the situation of an empty fluid infusion bottle can only rely on manual inspections, lacking an automatic monitoring and reminder mechanism.
[0006] 4) Difficult data integration: Data from different devices is difficult to uniformly collect and manage. Summary of the Invention
[0007] The purpose of the present application is to provide an intelligent medical data acquisition device based on a vision large model and an acquisition method thereof, which can improve the efficiency and accuracy of medical data acquisition.
[0008] To achieve the above purpose, the present application provides the following solutions:
[0009] In a first aspect, the present application provides an intelligent medical data acquisition device based on a vision large model, including: a hardware terminal and a background server;
[0010] The hardware terminal includes a main body platform, a dual gimbal system, a camera, a display screen, and a multi-functional fixing device; the main body platform includes an integrated processor, a memory, and a network module, which are used to process and analyze medical image information collected by the camera to obtain medical data, and to transmit the medical data to the background server through the network module; the dual gimbal system includes a device fixing gimbal and a camera gimbal; the display screen is used to display medical image information and medical data in real time; the multi-functional fixing device is used to adjust the position of the intelligent medical data acquisition device according to different medical scenarios;
[0011] The background server includes: a visual big model parsing engine and an alarm monitoring system linked to the visual big model parsing engine; the visual big model parsing engine is used to parse the medical image information displayed on the display screen; the alarm monitoring system is used to automatically trigger an alarm when an abnormality occurs in medical data or when equipment fails.
[0012] Optionally, the device fixed gimbal in the dual gimbal system adopts a three-axis mechanical gimbal; the three-axis mechanical gimbal is integrated with a human body tracking algorithm to capture the trajectory of set objects or monitored objects in real time; the camera gimbal in the dual gimbal system adopts a macro zoom gimbal; the macro zoom gimbal is equipped with an image stabilization compensation device.
[0013] Optionally, the camera includes an infrared fill light module, a microscopic shooting module and a dynamic frame rate adjustment module; the dynamic frame rate adjustment module is used to achieve interval acquisition according to user settings.
[0014] Optionally, the multifunctional fixing device specifically comprises:
[0015] An electromagnetic adsorption module, a flexible clamp module and a sterile hook assembly; the flexible clamp module is provided with a pressure sensor and an adaptive clamping mechanism.
[0016] Optionally, the hardware terminal further includes the network communication module;
[0017] The network communication module specifically includes:
[0018] Data encryption transmission unit, using AES-256 encryption algorithm;
[0019] Adaptive bandwidth allocator to ensure the quality of video streaming;
[0020] The breakpoint resume module is used for data caching and recovery when the network is interrupted.
[0021] Optionally, the hardware terminal further includes a power management system;
[0022] The power management system specifically includes:
[0023] Modular fast-charging battery pack for hot-swappable replacement;
[0024] Intelligent power consumption adjustment unit, used to dynamically adjust power consumption according to the working mode;
[0025] The magnetic charging base has a built-in wireless charging coil array, which can be used to charge multiple devices in any space.
[0026] Optionally, the alarm monitoring system specifically includes:
[0027] The multi-level alarm module includes a local audio-visual alarm unit for the device, a reminder unit for the nurse station, and a push unit for the mobile terminal.
[0028] Optionally, the background server further includes an API interface service module.
[0029] Optionally, the API interface service module specifically includes:
[0030] A medical device data standardization interface unit for HL7 / FHIR protocol conversion;
[0031] A real-time data push interface unit for performing WebSocket data stream services;
[0032] A medical record system docking interface unit for data update and sharing with the hospital HIS / LIS / PACS systems.
[0033] In a second aspect, the present application provides a medical data intelligent acquisition method based on the above-mentioned medical data intelligent acquisition device based on a visual large model, including:
[0034] Deploy the medical data intelligent acquisition device at the target acquisition location through a multi-functional fixing device;
[0035] Use a dual pan-tilt system to lock the operation area of medical staff and the medical monitoring interface respectively;
[0036] Set the acquisition parameters through the display screen; the acquisition parameters include the target device type, acquisition frequency, and patient information binding;
[0037] Start the intelligent acquisition mode, control the camera to execute an adaptive shooting strategy according to the preset parameters, and obtain medical image information;
[0038] Input the acquired medical image information into the background server, and perform multi-modal data joint analysis based on the visual large model parsing engine in the background server.
[0039] According to the specific embodiments provided by the present application, the following technical effects are disclosed:
[0040] This application provides a medical data intelligent acquisition device based on a visual large model and its acquisition method. In the hardware terminal, the main platform integrates a processor, a memory, and a network module, which can efficiently process the medical data collected by the camera and quickly transmit it to the background server through the network module, reducing the latency of data processing and improving the acquisition efficiency. The dual gimbal system allows the camera to take pictures from multiple angles with high precision, ensuring that the collected medical image information is comprehensive and accurate. The camera itself has high resolution and excellent imaging capabilities, further improving the accuracy of the data. The display screen real-time displays the collected medical image information and the monitoring screen, facilitating the operator to view and adjust immediately to ensure the accuracy of the acquisition process. The multi-functional fixing device can flexibly adjust its position according to different medical scenarios, ensuring the stability and applicability of the acquisition device. In the background server, the visual large model parsing engine can efficiently parse the medical image information displayed on the display screen, accurately identify and analyze the key information in the data using advanced algorithms and models, improving the accuracy of the data. The alarm monitoring system linked with the visual large model parsing engine can real-time monitor the abnormal situations in the data and send out alarms in a timely manner to ensure that the abnormal situations are handled promptly, further improving the efficiency and accuracy of the acquisition process. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0042] Figure 1 FIG. 9 is a schematic structural diagram of a medical data intelligent acquisition device provided by an embodiment of the present application.
[0043] Figure 2 FIG. 13 is a top view of the medical data intelligent acquisition device provided by an embodiment of the present application;
[0044] Figure 3 FIG. 17 is a side view of the medical data intelligent acquisition device provided by an embodiment of the present application;
[0045] Figure 4 FIG. 21 is a front view of the medical data intelligent acquisition device provided by an embodiment of the present application;
[0046] Figure 5 FIG. 25 is a data acquisition flow chart provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0048] To make the above objects, features, and advantages of the present application more apparent and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0049] Embodiment 1
[0050] As Figure 1 shown, this embodiment provides a medical data intelligent acquisition device based on a visual large model, including: a hardware terminal and a background server.
[0051] The hardware terminal includes a main body platform, a dual gimbal system, a camera, a display screen, and a multi-functional fixing device; the main body platform includes an integrated processor, a memory, and a network module, which are used to process and analyze the medical image information collected by the camera to obtain medical data, and to transmit the medical data to the background server through the network module; the dual gimbal system includes a device fixing gimbal and a camera gimbal; the display screen is used to display medical image information and medical data in real time; the multi-functional fixing device is used to adjust the position of the medical data intelligent acquisition device according to different medical scenarios;
[0052] The background server includes: a visual large model parsing engine and an alarm monitoring system linked with the visual large model parsing engine; the visual large model parsing engine is used to parse the medical image information displayed on the display screen; the alarm monitoring system is used to automatically trigger an alarm when the medical data is abnormal or the device fails.
[0053] Among them, as Figures 2 - 4 shown, the camera is arranged on the camera gimbal, a charging base is arranged under the display screen, the multi-functional fixing device can be a clamp, and control buttons such as lights and sensors can be set.
[0054] Among them, the device fixing gimbal in the dual gimbal system adopts a three-axis mechanical gimbal; the three-axis mechanical gimbal integrates a human tracking algorithm for real-time capturing of the trajectory of a set item or a monitoring object; the camera gimbal in the dual gimbal system adopts a macro zoom gimbal; the macro zoom gimbal is configured with an image stabilization compensation device.
[0055] The camera includes: an infrared fill light module, a microscopic shooting module, and a dynamic frame rate adjustment module; the dynamic frame rate adjustment module is used to achieve interval acquisition according to user settings.
[0056] The multifunctional fixing device specifically includes:
[0057] an electromagnetic adsorption module, a flexible fixture module, and a sterile hook assembly; a pressure sensor and an adaptive clamping mechanism are provided in the flexible fixture module.
[0058] Among them, the hardware terminal further includes the network communication module.
[0059] Specifically, the network communication module includes: a data encryption transmission unit, which adopts the AES-256 encryption algorithm; a bandwidth adaptive allocator, which is used to ensure the video stream transmission quality; a breakpoint resumption module, which is used for data caching and recovery when the network is interrupted. The network communication module supports 5G / Wi-Fi wireless transmission and Gigabit Ethernet wired transmission to achieve a stable connection with the central server.
[0060] The hardware terminal further includes a power management system. Specifically, the power management system includes: a modular fast-charging battery pack for hot-swap replacement; an intelligent power consumption adjustment unit for dynamically adjusting the power consumption according to the working mode; a magnetic charging base with a built-in wireless charging coil array for charging multiple devices placed arbitrarily in space. The rechargeable lithium battery pack using fast-charging technology can work continuously for 5-10 hours after a single charge; it is equipped with a magnetic charging base to support the simultaneous charging management of multiple devices.
[0061] Among them, the alarm monitoring system specifically includes:
[0062] a multi-level alarm module, including a device local sound and light alarm unit, a nurse station prompt unit, and a mobile terminal push unit.
[0063] In addition, the background server further includes an API interface service module.
[0064] The API interface service module specifically includes: a medical device data standardization interface unit for HL7 / FHIR protocol conversion; a real-time data push interface unit for executing WebSocket data stream services; a medical record system docking interface unit for data update and sharing with the hospital HIS / LIS / PACS systems.
[0065] Among them, in some embodiments, the visual large model parsing solution is to apply the visual large model to the operating room data collection, which can identify and parse various information such as medical device display screens, liquid levels, and pipeline states, realizing the intelligent acquisition of data that could not be collected originally.
[0066] Specifically, the visual big model is an artificial intelligence model that can understand and analyze the content of an image, and can identify objects, text, and their semantic relationships in the image. Currently, it is mainly divided into two categories: open source and closed source: open source models can freely obtain and modify source code, and closed source models only provide API interface calls. Common visual big models include: 1) Open source: Qwen2-VL (supports multi-resolution image understanding, excellent performance in multiple benchmarks), InternVL2_5 (the first open source model with an accuracy of over 70% in the MMMU benchmark), Paligemma2 (a new generation of visual language model that supports multiple languages), Vitron (supports pixel-level visual understanding and generation); 2) Closed source: GPT-4V (developed by OpenAI, with powerful visual understanding and multimodal interaction capabilities), Claude series visual models (developed by Anthropic, good at document understanding and image-assisted question and answer), Baidu Wenxin Yiyan visual module (supports image-generated text and image-text question and answer). These models generally adopt the Transformer architecture and process visual information through the self-attention mechanism.
[0067] In the past, when collecting medical data, artificial intelligence was used to identify objects in an image. It was necessary to train a unique model for a specific scenario. For example, medical staff used software (OCR) to identify the name and ID number on an ID card. Medical staff had to draw the name area and the ID card area for the ID card scenario. However, this model can only be used to identify the ID card. The general visual model can recognize information in any image. Medical staff only need to input a natural input image and the medical staff's prompt words, and the visual model can give a description of the image. In this way, medical staff can recognize the numbers, characters and non-monitoring images on the monitoring screen, such as the amount of fluid replacement, the remaining amount, the urine volume, and other data.
[0068] The present application also provides an application scenario of the medical data intelligent collection device, which is as follows:
[0069] In scenario 1, medical staff are busy in the operating room or ward, hanging up fluid bags for patients. Over time, the fluid bag may gradually become empty. If no one notices, although the empty bag will not cause air to enter the blood vessels due to the pressure difference, the lack of pressure for a long time may cause blood to flow back to the fluid bag and coagulate at the end of the infusion tube, making the fluid channel ineffective. Therefore, medical staff need to check the progress of fluid refilling frequently. With the help of this acquisition system, medical staff can monitor the remaining amount of fluid refilling in real time, and the system will automatically issue a low liquid level alarm.
[0070] In Scenario 2, during the operation, medical staff need to closely monitor urine output, especially the hourly urine output. However, due to the natural downward flow of urine and the urine bag being often placed under the bed and covered by drapes, inexperienced medical staff may overlook urine output monitoring. In children or high-risk patients, accurate measurement of urine output is particularly important. Currently, medical staff observe urine output visually, which is neither convenient nor easy to remember, unable to achieve high-frequency and high-precision recording, let alone automated recording. With this system, through image acquisition technology, the images of the urine bag at intermittent intervals of every 30s are captured by the supplementary light under the bed and the infrared sensor, enabling artificial intelligence analysis, providing high-precision urine output data collection, and realizing automatic recording of urine output.
[0071] In Scenario 3, although some devices such as ECMO machines or cardiopulmonary bypass machines are equipped with monitoring panels, their systems do not support data output, and existing acquisition systems cannot obtain this data. Medical staff can only regularly manually read the numbers on the monitoring screen and record them in another system, which not only reduces data accuracy but also is prone to errors during the transcription process. Through the solution of this application, high-frequency and high-precision data recording can be achieved.
[0072] Embodiment 2
[0073] This embodiment provides a medical data intelligent acquisition method based on the above-mentioned medical data intelligent acquisition device based on a visual large model, including:
[0074] Deploy the medical data intelligent acquisition device at the target acquisition position through a multi-functional fixing device.
[0075] Adopt a dual gimbal system to lock the operation area of medical staff and the medical monitoring interface respectively.
[0076] Set the acquisition parameters through the display screen; the acquisition parameters include the target device type, acquisition frequency, and patient information binding.
[0077] Start the intelligent acquisition mode, control the camera to execute an adaptive shooting strategy according to the preset parameters, and obtain medical image information.
[0078] Input the collected medical image information into the background server, and perform multi-modal data joint analysis based on the visual large model parsing engine in the background server.
[0079] Specifically, as Figure 5As shown, the collection method based on the collection device has the following operation process: First, the doctor fixes the collection device in place. Subsequently, the first pan-tilt unit adjusts its direction to align with the medical staff, and the doctor observes the image through the display screen. Then, the second pan-tilt unit adjusts its direction to align with the content to be collected. Through the touch screen of the host computer, the doctor sets the collection parameters, including information such as the collection target, frequency, and patient binding. After the settings are completed, the collection program is started, and the machine takes pictures according to the preset parameters. The taken pictures are then transmitted to the background server, where the visual model interprets the target information and stores the interpreted information. Finally, the doctor can review and analyze these data through the software platform.
[0080] In summary, the present application has the following technical effects:
[0081] 1) High degree of automation: It can complete the data collection work without manual intervention.
[0082] 2) Wide applicability: It is applicable to a variety of medical devices and monitoring environments, and is not restricted by the device brand and the presence or absence of data output interfaces.
[0083] 3) Easy installation: A variety of fixing methods are designed according to different collection requirements. For example, a magnetic fixing device is used for liquid level height collection, a clamping fixing device is used for device display screen collection, a hook fixing device is used for infusion drip rate collection, etc., and can be flexibly arranged and adjusted according to the actual application scenario.
[0084] 4) High data accuracy: The real-time collected data is accurate, improving the safety of surgery and the quality of anesthesia, enhancing the integrity of clinical research data, and helping to ensure the safety of patients.
[0085] 5) Good scalability: It can support new recognition requirements by updating the model.
[0086] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0087] In this article, specific examples are used to elaborate on the principle and implementation method of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An intelligent medical data acquisition device based on a large vision model, characterized in that, Comprising: Hardware terminal and back-end server; The hardware terminal includes a main body platform, a dual gimbal system, a camera, a display screen, and a multi-functional fixing device; The main body platform includes an integrated processor, a memory, and a network module, which are used to process and analyze the medical image information collected by the camera to obtain medical data, and to transmit the medical data to the back-end server through the network module; the dual gimbal system includes a device fixing gimbal and a camera gimbal; the display screen is used to display medical image information and medical data in real time; the multi-functional fixing device is used to adjust the position of the medical data intelligent acquisition device according to different medical scenarios; The back-end server includes: a visual large model parsing engine and an alarm monitoring system linked to the visual large model parsing engine; the visual large model parsing engine is used to parse the medical image information displayed on the display screen; the alarm monitoring system is used to automatically trigger an alarm when the medical data is abnormal or when a device fails.
2. The intelligent medical data acquisition device based on a large vision model according to claim 1, characterized in that The device fixing gimbal in the dual gimbal system adopts a three-axis mechanical gimbal; the three-axis mechanical gimbal integrates a human body tracking algorithm for real-time capturing of the trajectory of a set item or a monitoring object; the camera gimbal in the dual gimbal system adopts a macro zoom gimbal; the macro zoom gimbal is configured with an image stabilization compensation device.
3. The intelligent medical data acquisition device based on a large vision model according to claim 1, wherein, The camera includes an infrared supplementary light module, a microscopic shooting module, and a dynamic frame rate adjustment module; the dynamic frame rate adjustment module is used to achieve interval acquisition according to user settings.
4. The intelligent medical data acquisition device based on a visual large model according to claim 1, characterized in that The multi-functional fixing device specifically includes: An electromagnetic adsorption module, a flexible fixture module, and a sterile hook assembly; a pressure sensor and an adaptive clamping mechanism are provided in the flexible fixture module.
5. The intelligent medical data acquisition device based on a large vision model according to claim 1, wherein, The hardware terminal further includes the network communication module; The network communication module specifically includes: A data encryption transmission unit, adopting the AES-256 encryption algorithm; A bandwidth adaptive allocator for ensuring the quality of video stream transmission; A breakpoint resumption module for data caching and recovery during network interruption.
6. The intelligent medical data acquisition device based on a large vision model according to claim 1, characterized in that, The hardware terminal further includes a power management system; The power management system specifically includes: A modular fast-charging battery pack for hot-swap replacement; An intelligent power consumption adjustment unit for dynamically adjusting power consumption according to the working mode; A magnetic charging base with a built-in wireless charging coil array for charging multiple devices placed arbitrarily in space.
7. An intelligent medical data acquisition device based on a large vision model according to claim 1, characterized in that, The alarm monitoring system specifically includes: A multi-level alarm module, including a device local sound and light alarm unit, a nurse station prompt unit, and a mobile terminal push unit.
8. The intelligent medical data acquisition device based on a large vision model according to claim 1, characterized in that, The back-end server further includes an API interface service module.
9. An intelligent medical data acquisition device based on a large vision model according to claim 8, characterized in that, The API interface service module specifically includes: A medical device data standardization interface unit for HL7 / FHIR protocol conversion; A real-time data push interface unit for executing WebSocket data stream services; A medical record system docking interface unit for data update and sharing with the hospital HIS / LIS / PACS system.
10. A medical data intelligent acquisition method for a medical data intelligent acquisition device based on a visual large model according to any one of claims 1-9, characterized in that, Including: Deploy the medical data intelligent acquisition device at the target acquisition position through the multi-functional fixing device; Adopt a dual gimbal system to lock the operation area of medical staff and the medical monitoring interface respectively; Set acquisition parameters through the display screen; the acquisition parameters include the target device type, acquisition frequency, and patient information binding; Start the intelligent acquisition mode, control the camera to execute an adaptive shooting strategy according to the preset parameters, and obtain medical image information; Input the collected medical image information into the background server, and perform joint analysis of multi-modal data based on the visual large model parsing engine in the background server.