Hysteroscope uterine distention medium intelligent management and control system and method based on bimodal data fusion
By using a dual-modal data fusion system, the system can identify the stage of hysteroscopic surgery in real time and dynamically adjust the flow rate, thus solving the inaccuracy and safety issues of distension medium management in existing systems and realizing intelligent fluid management for hysteroscopic surgery.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-07
AI Technical Summary
Existing hysteroscopic distension media management systems lack the ability to intelligently perceive the surgical situation, making it difficult to accurately predict the fluid consumption rate. This poses risks of air embolism and excessive fluid absorption. Furthermore, the independent placement of the equipment prevents doctors from simultaneously obtaining information on the remaining fluid volume.
An intelligent control system based on dual-modal data fusion is adopted. The system collects the physical state data of the distending medium in real time through non-contact sensors, and combines it with deep learning algorithms to analyze the surgical video stream data, identify the operation stage, and dynamically adjust the flow rate consumption weight to achieve intelligent control of fluid management.
It improves the accuracy and safety of distension media management, ensures the continuity of the surgical procedure, reduces equipment upgrade costs, and provides real-time fluid control information and proactive safety protection.
Smart Images

Figure CN121811297A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of medical device and computer vision application technology, and in particular to an intelligent control system and method for hysteroscopic distension media based on dual-modal data fusion. Background Technology
[0002] Hysteroscopic surgery, as a core method of minimally invasive gynecological diagnosis and treatment, requires the continuous infusion of a fluid distending medium into the uterine cavity during the procedure. This expands the uterine cavity and flushes out accumulated blood, thereby maintaining a clear surgical field and appropriate intrauterine pressure for the surgeon. In current clinical practice, the management of distending medium infusion mainly relies on manual monitoring by medical staff or passive delivery via a basal pressure pump. Because the operating room environment is usually dimly lit to facilitate endoscopic visualization, and circulating nurses need to handle multiple tasks during critical surgical periods, blind spots in monitoring are inevitable. If the infusion bag is not detected empty in time, air can easily enter the uterine cavity through the tubing. This can not only cause a collapse of the surgical field, forcing an interruption, but may also lead to serious complications such as air embolism.
[0003] Furthermore, the surgeon's visual focus is always on the endoscopic image displayed on the monitor in front, while traditional fluid management equipment is usually placed independently behind or to the side of the surgeon. This spatial separation of visual information prevents the surgeon from obtaining information about the remaining fluid volume while simultaneously observing the surgical scene. Existing equipment lacks the ability to perceive the surgical context, and can only estimate the remaining time by making simple linear extrapolations based on the current physical flow rate. However, the rate of fluid consumption varies greatly at different stages, such as endoscopic observation, lesion resection, and high-flow irrigation. Mechanical predictions lacking contextual awareness are often highly inaccurate and cannot accurately guide the timing of fluid changes. At the same time, existing technologies lack active blocking mechanisms that combine video features with physical inflow and outflow monitoring to address potential abnormalities such as excessive fluid absorption or the risk of microbubble ingress during surgery. There are still some gaps in the intelligent management of surgical safety. Summary of the Invention
[0004] The purpose of this invention is to provide an intelligent control system and method for hysteroscopic distension media based on dual-modal data fusion, so as to solve the problems pointed out in the background art.
[0005] In a first aspect, the present invention provides an intelligent control system for hysteroscopic distension media based on dual-modal data fusion, which includes a distension media supply unit, a fluid infusion pipeline, a hysteroscopic camera host and a surgical display terminal, characterized in that the system further includes an intelligent sensing execution device and a central control server; The intelligent sensing and execution device is configured to be detachably installed on the outer wall of the fluid delivery pipeline for non-contact acquisition of the physical state data of the uterine expansion medium and to perform on / off control of the fluid delivery pipeline. The central control server is communicatively connected to the intelligent sensing execution device and the hysteroscopic camera host, and the central control server is equipped with a dual-modal data fusion engine. The dual-modal data fusion engine is configured to simultaneously receive surgical video stream data output by the hysteroscopic camera host and physical state data uploaded by the intelligent sensing execution device; the dual-modal data fusion engine performs real-time analysis of the surgical video stream data based on deep learning algorithms to identify the current surgical operation stage, and calculates the dynamic remaining available time of the distension medium by combining the first derivative features of the physical state data. The central control server generates hierarchical control commands based on the matching result between the dynamic remaining available time and the current surgical operation stage, and sends them to the intelligent sensing and execution device to adjust the fluid transmission status of the fluid infusion pipeline.
[0006] Optionally, the intelligent sensing and execution device includes a non-contact liquid level monitoring module, the specific structure of which is a capacitive sensing clamp or a gravity sensing bracket. When the capacitive sensing clamp is used, the capacitive sensing clamp is configured to be snapped into the dripping part of the fluid delivery pipeline, and generates droplet counting and flow rate monitoring signals in real time by detecting the change in dielectric constant caused by the droplet falling. When the gravity-sensing bracket is used, the gravity-sensing bracket is configured to suspend the uterine expansion medium supply unit, and the weight decay signal is collected in real time by a high-precision weighing sensor. The non-contact liquid level monitoring module is configured to send the liquid level height signal or the weight decay signal as physical state data to the central control server.
[0007] Optionally, the intelligent sensing and execution device further includes a pipeline blocking execution mechanism, which is located downstream of the non-contact liquid level monitoring module; The pipeline blocking actuator includes a miniature solenoid valve or a mechanical clamp; when the central control server determines that the dynamic remaining available time has reached zero or detects the risk of air embolism, it sends an emergency flow blocking command to drive the miniature solenoid valve to close or to drive the mechanical clamp to physically clamp the fluid delivery pipeline to prevent air from entering the patient's body.
[0008] Optionally, the intelligent sensing and execution device also integrates an audible and visual alarm module, which includes a multi-color LED indicator strip and a buzzer; The audible and visual alarm module is configured to respond to the warning commands of the central control server: display a solid green light when there is sufficient remaining distended uterine distension medium; display a flashing yellow light and emit an intermittent warning sound when the dynamic remaining available time is lower than the preset safe fluid replacement threshold; and display a strobe red light and emit a high-frequency continuous alarm sound when a pipeline blocking operation is performed.
[0009] Optionally, the surgical operation stage recognition logic in the dual-modal data fusion engine is specifically configured as follows: Keyframes of the surgical video stream data are extracted, and a pre-trained convolutional neural network model is used to extract texture features and identify instrument behavior in the keyframes. When the system detects bubble tumbling features or tissue fragment floating features with bubble density exceeding a preset threshold in the image, it is determined to be a high-flow cleaning stage. When the electrosurgical resection ring is detected to be moving back and forth at high frequency and accompanied by smoke features in the image, it is determined to be the lesion resection stage; when the image is static or has only a small displacement below the preset threshold, it is determined to be the endoscope observation stage. The central control server configures differentiated flow rate consumption weighting coefficients for the different stages mentioned above, and uses the weighting coefficients to make real-time corrections to the prediction model of the dynamic remaining available time.
[0010] Optionally, the central control server is also equipped with an augmented reality (AR) visualization processing module; The augmented reality (AR) visualization processing module is used to generate a semi-transparent virtual dashboard layer from the dynamic remaining available time, the current fluid flow rate value, and the pipeline blockage status. Image fusion technology is used to synchronously overlay the virtual dashboard layer frames onto the edge non-critical areas of the surgical video stream data, and output it through the surgical display terminal, so that doctors can intuitively obtain fluid control information while watching the surgical screen.
[0011] Optionally, the central control server may also include a fluid absorption calculation and alarm unit; The fluid absorption calculation and alarm unit is configured to: calculate the cumulative input of the distension medium based on the physical state data, and calculate the cumulative discharge based on the weight data of the surgical waste collection bucket; The difference between the cumulative input and the cumulative output is calculated in real time to obtain the amount of fluid absorbed by the human body. When the amount of fluid absorbed by the human body exceeds the set safety tolerance threshold, the central control server forcibly sends a blocking command to the intelligent sensing execution device and displays a risk warning of excessive fluid absorption on the surgical display terminal.
[0012] Optionally, the dual-modal data fusion engine also includes computer vision-based microbubble detection logic; The central control server performs inverse dark channel prior feature analysis on the surgical video stream data. When it detects bright particles in the video field of view that have a brightness value exceeding a preset threshold and exhibit directional flow characteristics, it determines that microbubbles have escaped from the pipeline. At this point, regardless of whether the physical state data triggers the threshold, the system will prioritize triggering a secondary warning for microbubbles and control the intelligent sensing and execution device to perform an emergency pipeline blocking action until the risk is manually confirmed to be eliminated.
[0013] Optionally, the system may also include a postoperative multidimensional data traceability subsystem; The postoperative multidimensional data traceability subsystem is used to link and store the surgical video stream data, the physical state data change curve, and the execution log of the hierarchical control instructions throughout the entire surgical process, using the time axis as an index, to form a tamper-proof digital surgical black box record for postoperative medical quality review and accountability tracing.
[0014] Secondly, the present invention provides an intelligent control method for hysteroscopic distension media based on dual-modal data fusion, which is applied to the system described in any one of the first aspects. The method includes the following steps: The intelligent sensing and execution device collects real-time physical quantity data of the distending medium at the fluid infusion pipeline and obtains real-time video stream data of the surgery through the hysteroscopic camera host. The physical quantity data and the video stream data are synchronized and aligned in time using a central control server. Based on computer vision algorithms, surgical operation stage features in the video stream data are identified, and the current fluid consumption weight is determined accordingly. By combining the rate of change of the physical quantity data with the fluid consumption weight, the time point for the evacuation of the expansion medium is dynamically predicted; The drainage time points and safety warning information are overlaid on the surgical screen using augmented reality. When it is predicted that the uterine distension medium is about to run out, microbubble intrusion is detected, or the amount of human body fluid absorption exceeds the standard, the actuator on the fluid delivery pipeline will automatically trigger an audible and visual alarm or block the pipeline.
[0015] The present invention has achieved the following beneficial effects: This invention constructs a dual-modal data fusion engine, deeply integrating the visual semantic features of surgical videos with the physical sensor data of the distending medium. The system utilizes deep learning algorithms to analyze the surgeon's actions in real time, accurately identifying the current cleaning, resection, or observation stage, and dynamically adjusting the calculation model for flow rate consumption weights accordingly. This behavior-understanding-based prediction mechanism corrects the bias of traditional equipment's single linear calculation, enabling more accurate prediction of the distending medium's emptying time. This ensures that the fluid management strategy is highly synchronized with the actual surgical rhythm, effectively guaranteeing the continuity of the surgical process.
[0016] This invention employs an external intelligent sensing and execution device, achieving non-contact data acquisition and active control without disrupting the existing sterile tubing closed loop. This avoids the risk of iatrogenic infection and reduces equipment upgrade costs. Combined with augmented reality visualization technology, the system integrates key data such as remaining time, flow rate, and safety alarms into a virtual dashboard overlaid on the edge area of the surgical display terminal, allowing doctors to intuitively obtain overall information without shifting their gaze, thus improving the human-computer interaction experience. Simultaneously, the system's integrated microbubble detection and human fluid absorption assessment functions can automatically trigger audible and visual alarms and execute physical flow control when potential risks are detected, providing a reliable active safety protection system against air embolism and water intoxication.
[0017] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.
[0018] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0019] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 A schematic diagram of the overall hardware architecture and data flow of a hysteroscopic distension media intelligent control system based on dual-modal data fusion provided in this embodiment of the invention; Figure 2 The flowchart illustrates the logical steps of an intelligent control method for hysteroscopic distension media based on dual-modal data fusion, as provided in this embodiment of the invention. Detailed Implementation
[0020] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0021] This invention provides an intelligent control system for hysteroscopic distension media based on dual-modal data fusion. For example... Figure 1 As shown, the system includes a distension medium supply unit, fluid infusion tubing, a hysteroscopic camera host, a surgical display terminal, an intelligent sensing and execution device, and a central control server. The system's hardware architecture also includes a wireless waste fluid monitoring terminal. This terminal is configured as a base-type structure independent of the trolley, placed at the bottom of the waste fluid collection tank under the operating table. It integrates an industrial-grade resistance strain gauge weighing sensor with a range of 0-50 kg and a low-power Bluetooth (BLE 5.0) transmitter module, used to collect real-time waste fluid weight increment data and synchronously upload it to the central control server, providing a data foundation for subsequent calculations of human fluid absorption.
[0022] The distension medium supply unit refers to the container used to hold medical distension fluid and its auxiliary pressurization facilities. In actual clinical applications, the distension medium is mostly physiological saline (0.9% sodium chloride solution), mannitol injection (5% concentration), or glucose injection. This supply unit can be a soft, sterile package suspended on a medical infusion rail or floor-standing infusion stand, with a capacity typically of 3000 ml or 5000 ml to accommodate prolonged surgical procedures. In some surgical scenarios requiring high intrauterine pressure, the distension medium supply unit also includes a pneumatic pressure cuff wrapped around the soft bag to increase the initial perfusion pressure of the fluid through inflation. The fluid infusion tubing is the physical channel connecting the distension medium supply unit to the inlet of the hysteroscopic instrument at the patient end. This tubing is made of medical-grade polyvinyl chloride or thermoplastic elastomer material, possessing good transparency, flexibility, and biocompatibility. A drip chamber (or Mofe's drip tube) is usually installed in the middle section of the tubing to buffer fluid pressure and facilitate visual observation of the drip rate by medical staff. The end of the fluid infusion line is connected to the inlet valve of the hysteroscopic instrument via a standard Luer locking connector. The hysteroscopic camera unit illuminates the uterine cavity with a cold light source transmitted through fiber optics and acquires high-definition real-time video signals from within the uterine cavity via a miniature image sensor (such as a CCD or CMOS array) on its head. This video signal is processed by the image processing unit (ISP) of the main unit and output as a digital video stream (such as SDI, HDMI, or DVI signals). The preferred resolution of the video stream is 1920×1080 pixels or 3840×2160 pixels, with a frame rate of no less than 30fps. The surgical display terminal is the main monitor placed in front of the operating table, used to present the endoscopic images acquired by the hysteroscopic camera unit to the surgeon in real time.
[0023] Based on this, the intelligent sensing and execution device is configured to be detachably installed on the outer wall of the fluid infusion pipeline. This device is a non-contact external attachment that does not require cutting or modifying the existing sterile fluid infusion pipeline, nor does it directly contact the medication inside the pipeline, thereby effectively avoiding the risk of iatrogenic cross-infection and reducing the hospital's upgrade costs for existing equipment. The intelligent sensing and execution device is positioned to non-contactly collect the physical state data of the distending medium and execute the on / off control of the fluid infusion pipeline. It integrates a microprocessor control unit (MCU), a power management module, a wireless communication module, a sensor array, and a motor drive unit. The intelligent sensing and execution device includes a non-contact liquid level monitoring module and a pipeline blocking execution mechanism located downstream of the non-contact liquid level monitoring module. Depending on the application scenario, the non-contact liquid level monitoring module can be configured as a snap-fit capacitor clamp or a suspended gravity bracket, while the pipeline blocking execution mechanism serves as the safe execution end of the device and is located downstream of the monitoring module. It is important to note that because the drip chamber acts as a buffer during infusion, its internal macroscopic liquid level is usually in dynamic equilibrium. Measuring only the static liquid level height cannot directly calculate the remaining liquid volume in the upstream infusion bag. Therefore, in this embodiment, the MCU operates using dynamic dielectric perturbation analysis: when a drop of distending medium falls from the dropper into the drip chamber, it excites tiny high-frequency oscillations on the liquid surface, causing characteristic pulse waveforms in the inter-electrode capacitance. The MCU performs time-domain analysis on the capacitance signal to identify the frequency of the oscillation waveform, thereby achieving accurate droplet counting and instantaneous flow rate calculation. The system allows medical staff to input the initial specifications of the distending medium supply unit (e.g., 3000 mL) via the surgical display terminal before the surgery begins. The central control server calculates the consumption based on the logic of subtracting the cumulative drop count from the initial specifications to obtain the accurate remaining physical quantity. .
[0024] The central control server is an industrial control computer (IPC) deployed behind the operating room information panel, or it can be a high-performance workstation integrated on a hysteroscopy trolley, or an edge computing node based on a private cloud architecture. The central control server is communicatively connected to the intelligent sensing and execution device and the hysteroscopy camera host, respectively. The communication link preferably uses a low-latency, high-reliability wireless protocol (such as Wi-Fi 6 or ZigBee 3.0) or a wired industrial bus (such as RS485 or Gigabit Ethernet) to ensure real-time communication of massive video data and sensor control signals.
[0025] The central control server is internally configured with a dual-modal data fusion engine. This engine is configured to simultaneously receive surgical video stream data output from the hysteroscopic camera host and physical state data uploaded by the intelligent sensing and execution device. Because the video data (visual modality) and sensor data (physical modality) originate from different hardware sources and have different sampling frequencies and transmission paths, they are often out of sync in time. Therefore, the dual-modal data fusion engine first performs time synchronization, using the Network Time Protocol (NTP) or hardware trigger signals to precisely align each frame of video image with the physical sensor values at the same time, constructing a unified multi-dimensional time-series dataset.
[0026] Subsequently, the dual-modal data fusion engine performs real-time analysis of the surgical video stream data based on deep learning algorithms. Utilizing computer vision technologies such as convolutional neural networks (CNNs), it performs frame-by-frame semantic understanding of the video footage to identify the current stage of the surgical procedure. The system can intelligently distinguish whether the doctor is currently performing endoscopic observation, high-flow irrigation, electrocautery, or hemostasis checks. Simultaneously, the engine calculates the dynamic remaining available time of the distending medium by combining the first-order derivative features of the physical state data. This first-order derivative feature, i.e., the fluid consumption rate (flow velocity), is obtained through real-time differentiation and filtering of remaining physical quantities (such as fluid level or weight). The system uses the surgical procedure stage as a key weighting factor to dynamically correct the prediction of the remaining time. For example, when the system visually recognizes that the doctor has begun electrocautery, even if the current physical flow rate has not yet reached its peak, the system will predict, based on the algorithm model, that the flow rate will significantly increase in the near future, thereby shortening the predicted remaining available time in advance and achieving proactive early warning.
[0027] Finally, the central control server generates tiered control instructions based on the matching result of the dynamic remaining available time and the current stage of the surgical operation. These instructions are sent to the intelligent sensing and execution device to adjust the fluid transmission status of the fluid infusion pipeline. The tiered control instructions include, but are not limited to: green status indication (normal), yellow warning indication (approaching depletion), and red alarm and blockage (depletion or bubble risk). Through this closed-loop control, the system achieves fully automated management of the entire process from sensing and decision-making to execution.
[0028] Furthermore, for the non-contact liquid level monitoring module in the intelligent sensing and execution device, this embodiment provides two specific preferred structural forms: a capacitive sensing clamp and a gravity sensing bracket. These two structures can be used individually or in combination to achieve data redundancy verification.
[0029] To achieve continuous linear liquid level monitoring, rather than simply detecting a threshold state of liquid presence or absence, the capacitive sensing fixture in this embodiment employs a multi-segment array electrode structure. Specifically, eight independent pairs of miniature copper foil electrodes (each 3mm high, 1mm vertically spaced) are evenly spaced vertically along the inner side of the fixture, forming a vertical interdigitated capacitor array. The MCU uses time-division multiplexing to sequentially scan the capacitance values of each electrode pair. When the liquid level is between two sets of electrodes, the capacitance value at that location undergoes a step change due to the much higher dielectric constant of the liquid compared to air. The system uses a linear interpolation algorithm to locate the capacitance step point, thereby calculating the liquid level height value accurate to the millimeter level.
[0030] When the capacitive sensing clamp is used, it is configured to snap onto the dripper portion of the fluid delivery pipeline. Its working principle no longer relies on static liquid level detection, but is based on dynamic dielectric perturbation analysis. Specifically, when a droplet of expansion medium drips from the dropper opening under gravity, passes through the capacitor plate area, and falls onto the liquid surface of the dripper, it causes a momentary disturbance in the electric field distribution between the plates, resulting in a characteristic pulse waveform in the capacitance value. The MCU inside the intelligent sensing actuator identifies the frequency of this pulse waveform to achieve accurate droplet counting and instantaneous flow rate calculation. The central control server subtracts the consumed amount calculated from the cumulative drop count from the initial total volume of the expansion medium supply unit (e.g., 3000 mL) to obtain an accurate remaining physical quantity. This method eliminates the measurement lag caused by the dripper's buffering effect.
[0031] When the gravity-sensing hanger is used, it is configured to suspend the uterine distension medium supply unit. The hanger is designed as a series suspension structure, with the upper end attached to the infusion stand and the lower end to the infusion bag. The hanger integrates a high-precision resistance strain gauge load cell, employing a full-bridge circuit structure to eliminate the effects of temperature drift. The high-precision load cell collects the weight decay signal in real time; this signal is a monotonically decreasing curve over time. The central control server performs differentiation on this curve to determine the current absolute remaining drug quantity and calculates the real-time flow rate. Compared to liquid level monitoring, gravity monitoring provides continuous analog data, making it more suitable for refined trend prediction. The non-contact liquid level monitoring module is configured to send the liquid level signal or the weight decay signal as physical state data to the central control server.
[0032] To ensure patient safety and prevent air embolism, the intelligent sensing and execution device also includes a pipeline blocking actuator. This actuator is positioned downstream of the non-contact liquid level monitoring module, closer to the patient. This arrangement ensures that if the monitoring module detects an anomaly (such as air bubbles or liquid voids), the downstream actuator has sufficient reaction time to intercept the abnormal fluid before it enters the body. The pipeline blocking actuator includes a miniature solenoid valve or a mechanical clamp. Considering that sterile tubing is generally not suitable for cutting off to install a solenoid valve, this preferred embodiment uses a mechanical clamp driven by a high-torque miniature DC motor or stepper motor, in conjunction with a reduction gearbox and a cam push rod mechanism. Under normal conditions, the clamp's pressure block is in the retracted position, and the fluid delivery pipeline is unobstructed. When the central control server determines that the dynamic remaining available time has reached zero or detects a risk of air embolism, it sends an emergency flow control command to drive the miniature solenoid valve to close or to drive the mechanical clamp to physically clamp the fluid delivery pipeline to prevent air from entering the patient's body.
[0033] To achieve intuitive human-computer interaction, the intelligent sensing and execution device also integrates an audible and visual alarm module, which includes a multi-color LED indicator strip and a buzzer. The audible and visual alarm module is configured to respond to warning commands from the central control server: displaying a solid green light when there is sufficient remaining distending medium, providing a safety cues to medical staff; displaying a flashing yellow light and emitting an intermittent beeping sound when the dynamic remaining available time is lower than a preset safe fluid replacement threshold (e.g., less than 3 minutes remaining), reminding the circulating nurse to prepare for fluid replacement; and displaying a flashing red light and emitting a high-frequency continuous alarm sound during tubing closure operations, warning of emergencies such as air bubble intrusion or severe water intoxication, requiring immediate intervention.
[0034] This embodiment has deeply optimized the surgical operation stage recognition logic in the dual-modal data fusion engine. The specific recognition process is as follows: First, keyframes are extracted from the surgical video stream data. To balance computational efficiency and real-time performance, the system employs an adaptive keyframe extraction strategy, extracting 5-10 frames per second. Next, a pre-trained convolutional neural network (CNN) model is used to extract texture features and recognize instrument behavior on the keyframes. This model is based on a ResNet or EfficientNet architecture and has undergone transfer learning training on massive amounts of hysteroscopic surgical video data.
[0035] Specific identification criteria include: when a large number of churning bubbles (manifested as densely packed, bright circular outlines) or floating tissue fragments (manifested as rapidly moving, irregular, dark-colored lumps) are detected in the image, it is determined to be the high-flow-rate cleaning stage. During this stage, fluid consumption is extremely high to flush the field of view. When the electrosurgical resection loop is detected moving at high frequency and accompanied by smoke-like characteristics (manifested as localized white turbidity), it is determined to be the lesion resection stage. During this stage, fluid is mainly used for cooling and maintaining pressure, with moderate consumption. When the image is static or shows only slight displacement, and there is no active instrument operation, it is determined to be the endoscope observation stage. During this stage, fluid consumption is lowest. The central control server configures differentiated flow rate consumption weighting coefficients for the above different stages (e.g., 2.0 for the cleaning stage, 1.2 for the resection stage, and 0.6 for the observation stage). These weighting coefficients are used to correct the prediction model of the dynamic remaining available time in real time, thereby improving the accuracy and clinical applicability of the prediction.
[0036] Furthermore, the specific data fusion architecture and prediction model are as follows: First, the network architecture: The dual-modal data fusion engine adopts a spatiotemporal two-stream structure. Specifically, the visual stream input is resized to a 224×224 pixel RGB image, extracted by ResNet-50, and then passed through a global average pooling layer to obtain a 2048-dimensional visual feature vector. The physical flow input consists of a 50-dimensional time sliding window of data (containing normalized flow velocity and pressure values), which is mapped to a 128-dimensional physical feature vector through three fully connected layers (MLP). To address the feature dimension imbalance issue, the system introduces a feature projection step before the fusion layer, which... Ascending to Dimensions Using the same dimensions, a weighted concatenation strategy is then employed to fuse the two data points. During training, a joint loss function is used. ,in (FocalLoss) is used to address the sample imbalance problem during the surgical phase. (Mean Squared Error Loss) is used to monitor the accuracy of the remaining duration prediction. To balance the hyperparameters of the loss weights for the two tasks (set to 0.5 in this embodiment), they are concatenated before the fully connected layer, and finally, the confidence level of the current surgical stage (observation / cleaning / resection) is output through a Softmax classifier.
[0037] Second, the calculation formula: Based on the identified surgical stage, the system calls a preset flow rate consumption weighting coefficient. (Cleaning stage) resection stage Observation phase Dynamic remaining available time. The calculation formula is: in, This represents the current remaining liquid volume. The average flow rate throughout the entire procedure; The current instantaneous flow velocity; This is the flow velocity weighting smoothing factor (value 0.6). This formula incorporates a visual coefficient. It can proactively correct the lag error caused by simple physical extrapolation.
[0038] To present the backend calculation results intuitively to the doctors, the central control server is also equipped with an augmented reality (AR) visualization processing module. This module generates a semi-transparent virtual dashboard layer from the dynamic remaining available time, current fluid flow rate, and pipeline blockage status. Image fusion technology is used to synchronously overlay this virtual dashboard layer onto the edges and non-critical areas (such as the four corners of the screen) of the surgical video stream data. The system automatically identifies and avoids regions of interest in the surgical view using a saliency detection algorithm, allowing doctors to intuitively obtain fluid control information while viewing the surgical screen without frequently looking up at the equipment tower.
[0039] To address the unique fluid perfusion abnormalities inherent in hysteroscopic surgery, the central control server also includes a fluid absorption calculation and alarm unit. This unit is configured to: calculate the cumulative input volume of the distension medium based on the physical state data, and calculate the cumulative discharge volume by combining the weight data of the surgical waste collection container (obtained via a bottom weighing sensor). The difference between the cumulative input volume and the cumulative discharge volume is calculated in real time to obtain the amount of fluid absorbed by the patient. When the amount of fluid absorbed by the patient exceeds a set safety tolerance threshold (e.g., 1000ml, which can be customized based on the patient's weight), the central control server forcibly sends a blocking command to the intelligent sensing execution device and displays a fluid over-absorption risk warning on the surgical display terminal, alerting the medical team to the discrepancy between fluid inflow and outflow.
[0040] Furthermore, for the detection of microbubbles, the dual-modal data fusion engine also includes microbubble detection logic based on computer vision. The central control server analyzes the surgical video stream data. When it detects bright particles in the video field of view with brightness values exceeding a preset threshold (e.g., pixel grayscale value > 220) and exhibiting directional flow characteristics (i.e., consistent optical flow vector field direction), it determines that microbubbles have entered the pipeline. This detection logic is based on the inverse dark channel prior principle of optics: the bubble surface undergoes strong specular reflection under coaxial light, causing it to exhibit abnormally bright characteristics in the dark channel image, which is significantly different from the diffuse reflection characteristics of red mucosal tissue. At this time, regardless of whether the physical data triggers the threshold, the system prioritizes triggering a secondary warning. To prevent false alarms caused by microbubbles statically adhering to the tube wall or sensor wall, the system first controls the actuator to perform a short-term high-frequency micro-vibration (i.e., driving the mechanical tube clamp to generate a slight vibration, rather than opening the tube) to physically shake off the bubbles from the tube wall. If the bubble characteristics show a continuous directional flow in the video after vibration, it is determined to be a real air embolism risk. The system will immediately send an emergency closure command to drive the pipeline to close completely. It is strictly forbidden to automatically open the pipeline when flowing bubbles are detected until the risk is eliminated by manual confirmation.
[0041] Finally, the system also includes a postoperative multidimensional data traceability subsystem. This subsystem uses a timeline as an index to link and store the surgical video stream data, the physical state data change curves, and the execution logs of the hierarchical control instructions throughout the entire surgical process. It then uses a hash digest algorithm or blockchain distributed ledger technology to encrypt and lock these data, forming a tamper-proof digital surgical black box record for postoperative medical quality review and accountability.
[0042] This invention also provides an intelligent control method for hysteroscopic distension media based on dual-modal data fusion, which is applied to the system described in Embodiment 1. Figure 2 As shown, the method includes the following detailed steps: Step S101: Multi-source sensing data acquisition.
[0043] After the system starts, the intelligent sensing and execution device begins to work, collecting real-time physical quantity data of the distended medium at the fluid infusion line. If a capacitive clamp is used, the dielectric constant change sequence is collected; if a gravity hanger is used, the weight change sequence is collected. Simultaneously, real-time video stream data of the surgery is acquired through the hysteroscopic camera host. To ensure real-time sensing, the sampling rate of physical data is preferably no less than 50Hz, and the frame rate of video data acquisition is no less than 30fps.
[0044] Step S102: Spatiotemporal synchronization of heterogeneous data.
[0045] A central control server is used to synchronize and align the physical quantity data and the video stream data in time. Because the transmission paths of the physical signals and the processing paths of the video signals are different, there may be a deviation of tens of milliseconds in their arrival time at the server. The server uses a high-precision system clock to tag each data packet and sets up a data buffer queue.
[0046] Step S103: Visual semantic feature recognition.
[0047] The surgical procedure stage features in the video stream data are identified using computer vision algorithms. Aligned video frames are input into a deep convolutional neural network model to extract high-dimensional feature vectors. Based on the classification results of the feature vectors, it is determined whether the current stage is high-flow cleaning, lesion resection, or endoscopic observation, and the current fluid consumption weight is determined accordingly. For example, if it is identified as a cleaning stage, a higher flow rate prediction weight is assigned.
[0048] Step S104: Dynamic remaining time prediction.
[0049] By combining the rate of change of the physical quantity data (i.e., real-time physical velocity) with the fluid consumption weight, the emptying time of the expanded medium is dynamically predicted. The prediction model uses the Kalman filter algorithm, taking the physical velocity as the observed value and the visual weight as the correction factor for the state transition matrix, thereby smoothing noise and accurately predicting future liquid volume trends.
[0050] Step S105: Augmented reality information presentation.
[0051] The drainage time points and safety warning information are overlaid on the surgical screen using augmented reality. The rendering engine generates a semi-transparent HUD layer and blends it into non-critical areas of the video stream, which is then output to the doctor through the display terminal.
[0052] Step S106: Risk closed-loop control.
[0053] When the system predicts that the distending medium is about to run out, detects microbubble intrusion, or that the body's fluid absorption exceeds the limit, the actuators on the fluid infusion line automatically trigger an audible and visual alarm or shut off the line. The system performs a tiered response based on the risk level: for low-risk situations, it only prompts for fluid replacement; for high-risk situations, it immediately physically cuts off the line to ensure the safety of the procedure to the greatest extent possible.
[0054] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A hysteroscopic distension medium intelligent control system based on dual-modal data fusion, comprising a distension medium supply unit, a fluid infusion pipeline, a hysteroscopic camera host, and a surgical display terminal, characterized in that, The system also includes intelligent sensing and execution devices and a central control server; The intelligent sensing and execution device is configured to be detachably installed on the outer wall of the fluid delivery pipeline for non-contact acquisition of the physical state data of the expansion medium and to perform on / off control of the fluid delivery pipeline; the intelligent sensing and execution device includes a non-contact liquid level monitoring module and a pipeline blocking execution mechanism located downstream of the non-contact liquid level monitoring module. The central control server is communicatively connected to the intelligent sensing execution device and the hysteroscopic camera host, and the central control server is equipped with a dual-modal data fusion engine. The dual-modal data fusion engine is configured to simultaneously receive surgical video stream data output by the hysteroscopic camera host and physical state data uploaded by the intelligent sensing execution device; The dual-modal data fusion engine uses a deep learning algorithm to analyze the surgical video stream data in real time to identify the current surgical operation stage, and combines the first derivative features of the physical state data to calculate the dynamic remaining available time of the distending medium. The central control server generates hierarchical control commands based on the matching result of the dynamic remaining available time and the current surgical operation stage, and sends them to the intelligent sensing and execution device to drive the pipeline blocking actuator to adjust the fluid transmission state of the fluid infusion pipeline.
2. The intelligent control system for hysteroscopic distension media based on dual-modal data fusion according to claim 1, characterized in that, The intelligent sensing and execution device includes a non-contact liquid level monitoring module, the specific structure of which is a capacitive sensing clamp or a gravity sensing bracket. When the capacitive sensing clamp is used, the capacitive sensing clamp is configured to be snapped into the dripping part of the fluid delivery pipeline, and generates droplet counting and flow rate monitoring signals in real time by detecting the change in dielectric constant caused by the droplet falling. When the gravity-sensing bracket is used, the gravity-sensing bracket is configured to suspend the uterine expansion medium supply unit, and the weight decay signal is collected in real time by a high-precision weighing sensor. The non-contact liquid level monitoring module is configured to send the liquid level height signal or the weight decay signal as physical state data to the central control server.
3. The intelligent control system for hysteroscopic distension media based on dual-modal data fusion according to claim 1, characterized in that, The intelligent sensing and execution device also includes a pipeline blocking execution mechanism, which is located downstream of the non-contact liquid level monitoring module; The pipeline blocking actuator includes a miniature solenoid valve or a mechanical clamp; when the central control server determines that the dynamic remaining available time has reached zero or detects the risk of air embolism, it sends an emergency flow blocking command to drive the miniature solenoid valve to close or to drive the mechanical clamp to physically clamp the fluid delivery pipeline to prevent air from entering the patient's body.
4. The intelligent control system for hysteroscopic distension media based on dual-modal data fusion according to claim 1, characterized in that, The intelligent sensing and execution device also integrates an audible and visual alarm module, which includes a multi-color LED indicator strip and a buzzer. The audible and visual alarm module is configured to respond to the warning commands of the central control server: display a solid green light when there is sufficient remaining distended uterine distension medium; display a flashing yellow light and emit an intermittent warning sound when the dynamic remaining available time is lower than the preset safe fluid replacement threshold; and display a strobe red light and emit a high-frequency continuous alarm sound when a pipeline blocking operation is performed.
5. The intelligent control system for hysteroscopic distension media based on dual-modal data fusion according to claim 1, characterized in that, The specific configuration of the surgical operation stage identification logic in the dual-modal data fusion engine is as follows: Keyframes of the surgical video stream data are extracted, and a pre-trained convolutional neural network model is used to extract texture features and identify instrument behavior in the keyframes. When the system detects bubble tumbling features or tissue fragment floating features with bubble density exceeding a preset threshold in the image, it is determined to be a high-flow cleaning stage. When the electrosurgical resection ring is detected to be moving back and forth at high frequency and accompanied by smoke features in the image, it is determined to be the lesion resection stage; when the image is static or has only a small displacement below the preset threshold, it is determined to be the endoscope observation stage. The central control server configures differentiated flow rate consumption weighting coefficients for the different stages mentioned above, and uses the weighting coefficients to make real-time corrections to the prediction model of the dynamic remaining available time.
6. The intelligent control system for hysteroscopic distension media based on dual-modal data fusion according to claim 1, characterized in that, The central control server is also equipped with an augmented reality (AR) visualization processing module; The augmented reality (AR) visualization processing module is used to generate a semi-transparent virtual dashboard layer from the dynamic remaining available time, the current fluid flow rate value, and the pipeline blockage status. Image fusion technology is used to synchronously overlay the virtual dashboard layer frames onto the edge non-critical areas of the surgical video stream data, and output it through the surgical display terminal, so that doctors can intuitively obtain fluid control information while watching the surgical screen.
7. The intelligent control system for hysteroscopic distension media based on dual-modal data fusion according to claim 1, characterized in that, The central control server also includes a fluid absorption calculation and alarm unit; The fluid absorption calculation and alarm unit is configured to: calculate the cumulative input of the distension medium based on the physical state data, and calculate the cumulative discharge based on the weight data of the surgical waste collection bucket; The difference between the cumulative input and the cumulative output is calculated in real time to obtain the amount of fluid absorbed by the human body. When the amount of fluid absorbed by the human body exceeds the set safety tolerance threshold, the central control server forcibly sends a blocking command to the intelligent sensing execution device and displays a risk warning of excessive fluid absorption on the surgical display terminal.
8. The intelligent control system for hysteroscopic distension media based on dual-modal data fusion according to claim 1, characterized in that, The dual-modal data fusion engine also includes computer vision-based microbubble detection logic; The central control server performs inverse dark channel prior feature analysis on the surgical video stream data. When it detects bright particles in the video field of view that have a brightness value exceeding a preset threshold and exhibit directional flow characteristics, it determines that microbubbles have escaped from the pipeline. At this point, regardless of whether the physical state data triggers the threshold, the system will prioritize triggering a secondary warning for microbubbles and control the intelligent sensing and execution device to perform an emergency pipeline blocking action until the risk is manually confirmed to be eliminated.
9. The intelligent control system for hysteroscopic distension media based on dual-modal data fusion according to claim 1, characterized in that, The system also includes a postoperative multidimensional data traceability subsystem; The postoperative multidimensional data traceability subsystem is used to link and store the surgical video stream data, the physical state data change curve, and the execution log of the hierarchical control instructions throughout the entire surgical process, using the time axis as an index, to form a tamper-proof digital surgical black box record for postoperative medical quality review and accountability tracing.
10. A method for intelligent control of hysteroscopic distension media based on dual-modal data fusion, characterized in that, Applied to the system according to any one of claims 1-9, the method comprises the following steps: The intelligent sensing and execution device collects real-time physical quantity data of the distending medium at the fluid infusion pipeline and obtains real-time video stream data of the surgery through the hysteroscopic camera host. The physical quantity data and the video stream data are synchronized and aligned in time using a central control server. Based on computer vision algorithms, surgical operation stage features in the video stream data are identified, and the current fluid consumption weight is determined accordingly. By combining the rate of change of the physical quantity data with the fluid consumption weight, the time point for the evacuation of the expansion medium is dynamically predicted; The drainage time points and safety warning information are overlaid on the surgical screen using augmented reality. When it is predicted that the uterine distension medium is about to run out, microbubble invasion is detected, or the calculated amount of human body fluid absorption exceeds the threshold, the actuator on the fluid infusion pipeline will automatically trigger an audible and visual alarm or block the pipeline.