Real-time image enhancement and fusion system in thoracic surgery

The multimodal fusion system based on U-Net convolutional neural network-based image enhancement and dynamic registration algorithm solves the problems of real-time performance and accuracy of intraoperative imaging technology in thoracic surgery, significantly improving surgical safety and efficiency.

CN121921270APending Publication Date: 2026-04-24中国人民解放军总医院第八医学中心
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
中国人民解放军总医院第八医学中心
Filing Date
2025-12-27
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing intraoperative imaging technologies in thoracic surgery suffer from insufficient real-time image quality, lagging multimodal image fusion, and inadequate real-time performance, leading to high surgical error rates, prolonged operation time, and increased postoperative complications. Existing systems have failed to effectively address the need for real-time image enhancement and fusion in dynamic intraoperative scenarios.

Method used

Image enhancement is achieved using the U-Net convolutional neural network model based on the attention mechanism, and multimodal image fusion is realized by combining dynamic registration algorithm. Low-latency data transmission is achieved through 5G communication and PCIe 4.0 transmission. It is equipped with anti-interference shielding components and high-precision display devices, and supports 3D image display and data interaction.

Benefits of technology

It achieves a real-time image clarity improvement of over 30%, a noise suppression rate of ≥85%, a registration error of ≤1mm, significantly reduces the risk of accidental injury, shortens the operation time by 10% to 15%, and reduces the incidence of complications by over 50%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921270A_ABST
    Figure CN121921270A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time image enhancement and fusion system in thoracic surgery, and belongs to the technical field of medical image processing. The system comprises an image acquisition module, a real-time enhancement unit, a multi-modal fusion unit, a data interaction unit, a display control unit and a storage unit. The image acquisition module is used for acquiring intraoperative endoscope and ultrasonic images and preoperative CT / MRI data; the real-time enhancement unit adopts an attention mechanism improved U-Net model to realize image real-time denoising and enhancement; the multi-modal fusion unit realizes accurate fusion through a dynamic registration algorithm (in combination with respiratory movement data); the data interaction unit guarantees low-delay transmission; the display unit supports three-dimensional image interactive display; the storage unit stores data in real time. The problems of low image definition, inaccurate fusion and poor real-time performance in the existing operation are solved, the key anatomical structure recognition precision can be improved by 30% or above, the operation time is shortened by 10%-15%, the complication occurrence rate is reduced by 50% or above, and the method is suitable for thoracic surgery operations such as lung lobe resection and esophageal cancer radical treatment and has remarkable clinical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing technology, specifically to a real-time intraoperative image enhancement and fusion system for thoracic surgery, which is particularly suitable for thoracic surgeries such as lobectomy, radical esophagectomy, and puncture localization of lung nodules. Background Technology

[0002] Thoracic surgery is characterized by complex anatomical structures, limited operating space, and dense concentrations of critical blood vessels and nerves. The accuracy of intraoperative lesion localization, tissue boundary identification, and assessment of adjacent structures directly impacts surgical safety. Current intraoperative imaging techniques suffer from the following core deficiencies:

[0003] Insufficient real-time image quality: Intraoperative images from endoscopy, ultrasound, and other techniques are easily obstructed by surgical instruments, tissue bleeding, and respiratory movements, resulting in problems such as high noise levels, low contrast, and blurred edges. This makes it difficult to identify critical structures such as small blood vessels (e.g., branches of pulmonary arteries) and nerve fibers. For example, in lung nodule puncture and localization surgery, nodules smaller than 5 mm in diameter are often missed due to insufficient image contrast, increasing the risk of puncture failure. In radical esophagectomy, small lymphatic vessels around the esophagus are easily damaged due to blurred edges, leading to postoperative lymphatic leakage.

[0004] Multimodal image fusion lag: The fusion of preoperative CT / MRI with intraoperative real-time images relies on manual registration or static algorithms, which cannot adapt to dynamic changes in intraoperative anatomical structures (such as lung expansion / collapse, tissue traction). Registration errors typically exceed 3mm, easily leading to accidental injury. Clinical data shows that pulmonary vascular injury cases caused by registration lag account for 62% of all intraoperative vascular injuries, and the proportion requiring secondary surgery for repair is as high as 38%.

[0005] Insufficient real-time performance: Traditional image enhancement algorithms (such as histogram equalization and wavelet transform) have high computational complexity and processing delays exceeding 200ms, which cannot meet the requirements of real-time intraoperative operations. In thoracoscopic lobectomy, surgeons need to quickly adjust instrument positions based on real-time images; delays exceeding 200ms can lead to operational lag and increase the probability of lung tissue damage.

[0006] High clinical risks: The aforementioned drawbacks result in an intraoperative injury rate as high as 8%-12%, a higher incidence of postoperative complications (such as massive hemorrhage and pneumothorax), and a 15%-20% increase in surgical time, increasing patient trauma and medical costs. According to statistics from the *Chinese Journal of Thoracic and Cardiovascular Surgery* in 2024, thoracic surgeries using traditional imaging techniques resulted in an average postoperative hospital stay 2.3 days longer than expected, and an average increase of 18,000 yuan in medical expenses per patient.

[0007] Existing technologies, such as a medical image fusion method, are not optimized for dynamic intraoperative scenarios in thoracic surgery and do not consider the impact of respiratory motion on image registration. Their real-time performance and registration accuracy cannot meet clinical needs. Similarly, existing intraoperative image enhancement systems do not incorporate multimodal fusion, optimizing only a single image and failing to provide information on the correlation between preoperative anatomical structures and real-time intraoperative images. Surgeons still need to rely on experience to judge tissue boundaries. Furthermore, some existing systems lack electromagnetic interference resistance; when using instruments such as electrocoagulation hooks and ultrasonic scalpels, images are prone to snow-like noise, further affecting surgical procedures.

[0008] Therefore, there is an urgent need to develop a real-time image enhancement and fusion system adapted to intraoperative scenarios in thoracic surgery to address clinical pain points. Summary of the Invention

[0009] To address the shortcomings of existing technologies, the present invention aims to provide a real-time image enhancement and fusion system for thoracic surgery, achieving the following objectives:

[0010] Improve the clarity and contrast of real-time intraoperative images and effectively suppress noise and interference;

[0011] Achieve dynamic and precise fusion of preoperative CT / MRI with intraoperative endoscopic and ultrasound images, with a registration error of ≤1mm;

[0012] Ensure low latency (≤50ms) for image processing and transmission to meet the needs of real-time intraoperative operation;

[0013] It helps doctors accurately identify key anatomical structures, reducing surgical risks and shortening operation time.

[0014] To achieve the above objectives, the present invention provides the following technical solution:

[0015] A real-time intraoperative image enhancement and fusion system for thoracic surgery includes:

[0016] The image acquisition module is used to acquire real-time endoscopic images, ultrasound images and preoperative CT / MRI images during thoracic surgery. The image acquisition module includes a 1080P high-definition endoscopic camera, a high-frequency ultrasound probe and a DICOM 3.0 standard interface. The frame rate of the 1080P high-definition endoscopic camera is not less than 30fps and the frequency range of the high-frequency ultrasound probe is 5-12MHz.

[0017] The real-time enhancement unit, connected to the image acquisition module, employs an improved U-Net convolutional neural network model based on an attention mechanism to perform adaptive noise suppression, edge enhancement, and contrast enhancement processing on real-time endoscopic and ultrasound images. The encoder of the model is a ResNet50 backbone network, and the decoder includes spatial attention gates and channel attention gates. The processing frame rate of the real-time enhancement unit is no less than 30fps, and the peak signal-to-noise ratio of the output image is no less than 35dB, and the structural similarity index is no less than 0.9.

[0018] The multimodal fusion unit is connected to the real-time enhancement unit and the image acquisition module respectively. It adopts a dynamic registration algorithm of feature extraction-dynamic update-registration optimization, and corrects the feature point matching relationship by combining intraoperative respiratory motion data. It achieves three-dimensional registration of the enhanced real-time image and the preoperative CT / MRI image through thin plate spline interpolation algorithm. Then, it outputs the multimodal fusion image based on the weighted fusion strategy of maximizing mutual information. The registration error of the multimodal fusion unit does not exceed 1mm, and the spatial resolution of the output multimodal fusion image is not less than 1mm×1mm×1mm.

[0019] The data interaction unit is used to realize low-latency data transmission between modules, supports 5G millimeter wave communication and PCIe 4.0 wired transmission. The transmission rate of the 5G millimeter wave communication is not less than 10Gbps, and the transmission latency of the PCIe 4.0 wired transmission is not more than 20ms. The data interaction unit uses H.265 encoding and AES-256 encryption to synchronously process data.

[0020] The display control unit is connected to the multimodal fusion unit and adopts a 4K ultra-high-definition medical display or holographic projection device. It supports three-dimensional image rotation, scaling, local magnification and key structure annotation. The adjustment response delay of the display control unit does not exceed 10ms.

[0021] The storage unit uses an SSD array with a storage capacity of no less than 2TB, which is used to store intraoperative image data and operation logs in real time and supports data export in DICOM 3.0 format.

[0022] Furthermore, the processing flow of the real-time enhancement unit includes:

[0023] S1: Adaptive noise suppression of real-time images is performed using a Gaussian mixture model;

[0024] S2: Extract the edges of anatomical structures using an improved Canny operator, which has an adaptive threshold adjustment function;

[0025] S3: An improved CLAHE algorithm is used to enhance the contrast of key areas. The improved CLAHE algorithm introduces a regional weight allocation mechanism.

[0026] S4: Output the enhanced image, with the processing delay of the real-time enhancement unit not exceeding 30ms.

[0027] Furthermore, the dynamic registration algorithm of the multimodal fusion unit specifically includes:

[0028] S1: The SIFT algorithm is used to extract anatomical feature points from the enhanced real-time images and the preoperative CT / MRI images. The number of anatomical feature points extracted from each image is no less than 50.

[0029] S2: Intraoperative respiratory motion data are collected through a pressure sensor with a sampling frequency of 100Hz. The feature point matching relationship is corrected in real time based on the collected intraoperative respiratory motion data.

[0030] S3: Image deformation registration is achieved using a thin-plate spline interpolation algorithm;

[0031] S4: When the registration error exceeds 1mm, re-registration will be automatically triggered and an alarm will be issued.

[0032] Furthermore, in the weighted fusion strategy of the multimodal fusion unit, the weight range of endoscopic images is 0.4-0.6, the weight range of ultrasound images is 0.2-0.3, and the weight range of CT / MRI images is 0.2-0.3, and the weights can be dynamically adjusted.

[0033] Furthermore, the medical display used in the display control unit has a brightness of no less than 1000 cd / m². 2 The contrast ratio is not less than 1000:1, and the holographic projection device used in the display control unit has a projection accuracy of not less than 0.5mm.

[0034] Furthermore, the image acquisition module is equipped with an anti-interference shielding component, which adopts a double-layer copper mesh shielding structure and has a shielding effect of over 60dB. The data interaction unit supports bidirectional data interaction with the surgical robot and navigation system.

[0035] Furthermore, the real-time enhancement unit is equipped with an NVIDIA A100 GPU, which optimizes model inference speed through the TensorRT inference engine. The NVIDIA A100 GPU has a video memory capacity of 40GB.

[0036] Furthermore, the storage unit uses an SSD array containing four SSDs in RAID 5 mode. The SSD array is equipped with an array controller that supports cache protection and has a cache capacity of 1GB.

[0037] Compared with the prior art, the present invention has the following significant advantages:

[0038] 1. Outstanding real-time performance: Overall processing and transmission latency ≤50ms, perfectly matching the rhythm of thoracic surgery operations and solving the lag problem of traditional systems;

[0039] 2. Excellent enhancement effect: Through the attention mechanism CNN model, the clarity of key anatomical structures is improved by more than 30%, and the noise suppression rate is ≥85%, effectively assisting doctors in identifying small blood vessels and nerves;

[0040] 3. High fusion accuracy: The dynamic registration algorithm adapts to changes in anatomical structures during surgery, with a registration error of ≤1mm, which is far superior to existing technologies (3-5mm), significantly reducing the risk of accidental injury;

[0041] 4. Significant clinical value: It can shorten the operation time by 10%-15%, reduce the incidence of complications such as intraoperative massive bleeding and pneumothorax by more than 50%, and reduce patient trauma and medical costs;

[0042] 5. High compatibility: It supports integration with existing surgical equipment and imaging systems, eliminating the need for large-scale modifications to the operating room and facilitating clinical application. Attached Figure Description

[0043] Figure 1 This is a diagram showing the overall system architecture and data flow in this invention;

[0044] Figure 2 This is a flowchart of the real-time enhancement unit algorithm processing in this invention;

[0045] Figure 3 This is a flowchart of the dynamic registration and fusion process of the multimodal fusion unit in this invention. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0047] The real-time intraoperative image enhancement and fusion system for thoracic surgery of the present invention is characterized by:

[0048] The system includes an image acquisition module, a real-time enhancement unit, a multimodal fusion unit, a data interaction unit, a display control unit, and a storage unit. These modules work collaboratively via a high-speed bus or 5G communication. The specific structure and functions are as follows:

[0049] (1) Image acquisition module

[0050] It includes a 1080P high-definition endoscopic camera, a high-frequency ultrasound probe, and a DICOM 3.0 standard interface. The endoscopic camera has a frame rate of no less than 30fps, enabling smooth capture of dynamic images of the surgical field. The high-frequency ultrasound probe has a frequency range of 5-12MHz and supports switching frequencies according to the type of surgery. For example, a 12MHz high-frequency mode is used for lung nodule puncture, while a 5MHz low-frequency mode is used for radical esophagectomy to expand the detection range.

[0051] It can acquire intraoperative endoscopic and ultrasound images in real time and import preoperative CT / MRI image data. It supports multiple formats such as DICOM, JPEG, and PNG and can be adapted to different brands of imaging equipment.

[0052] It is equipped with an anti-interference shielding component, which adopts a double-layer copper mesh shielding structure with a shielding effectiveness of over 60dB. This component can effectively avoid image distortion caused by electromagnetic interference from surgical instruments such as electrocoagulation hooks and ultrasonic scalpels, and ensure stable image signals when the instruments are in operation.

[0053] (2) Real-time enhancement unit

[0054] Core Algorithm: An improved U-Net convolutional neural network model based on the attention mechanism. The encoder uses a ResNet50 backbone network to extract deep features from images. The decoder introduces spatial attention gates and channel attention gates. The spatial attention gate can focus on key surgical areas (blood vessels, nerves), while the channel attention gate can enhance the information transmission of effective feature channels and reduce interference from useless information.

[0055] Processing flow:

[0056] ① Adaptive noise suppression is performed on the acquired real-time images. A Gaussian mixture model is used to analyze the gray-scale distribution of image pixels, automatically identify the noise type and match the suppression parameters. Compared with traditional fixed parameter denoising, the noise suppression effect is improved by 25%.

[0057] ② The edge detection module is an optimized version of the Canny operator, which adds an adaptive threshold adjustment function. It can automatically adjust the high and low thresholds according to the image brightness, avoiding the problem of missed detection in dark areas and false detection in bright areas by the traditional Canny operator.

[0058] ③ An adaptive contrast enhancement algorithm, an improved version of CLAHE, is adopted, which introduces a regional weight allocation mechanism. Higher enhancement weights are assigned to key areas such as blood vessels and nerves, while the weights are reduced for background areas to avoid background noise amplification caused by over-enhancement.

[0059] ④ The output enhanced image should have a frame rate of no less than 30fps, a peak signal-to-noise ratio (PSNR) of no less than 35dB, and a structural similarity index (SSIM) of no less than 0.9 to ensure that the image is clear and has high consistency with the original structure;

[0060] Hardware acceleration: Equipped with an NVIDIA A100 GPU, which has 40GB of video memory and supports TensorRT inference engine optimization, the inference speed is increased by more than 3 times through model quantization, layer fusion and other technologies, ensuring that the processing latency does not exceed 30ms.

[0061] (3) Multimodal fusion unit

[0062] Dynamic registration algorithm:

[0063] ①Feature extraction: The SIFT algorithm is used to extract anatomical feature points from the enhanced real-time images and preoperative CT images, such as bronchial bifurcation, vascular nodes, and lung nodule edges. The number of feature points extracted from each image is no less than 50 to ensure sufficient registration reference points.

[0064] ② Dynamic update: Based on intraoperative respiratory motion monitoring data, the data is collected by a pressure sensor installed on a soft patch around the surgical incision. The sampling frequency is 100Hz. The sensor can acquire tissue displacement data caused by respiratory motion in real time and correct the feature point matching relationship in real time according to the displacement data so that the registration adapts to the changes in the respiratory cycle.

[0065] ③ Registration optimization: Thin plate spline interpolation (TPS) algorithm is used to achieve image deformation registration. By establishing nonlinear mapping relationship between feature points, the preoperative CT image is deformed and adjusted to align with the anatomical structure of the real-time intraoperative image, with a registration error of no more than 1mm.

[0066] ④ Error detection: A registration error monitoring module is set up to calculate the mutual information value of the registered image in real time. When the mutual information value is lower than 0.8 (corresponding to a registration error of more than 1mm), re-registration is automatically triggered and an audible and visual warning is issued to remind the doctor to pause the operation until registration is completed.

[0067] Fusion Algorithm: Based on a weighted fusion strategy that maximizes mutual information, the weights of each modality of images are dynamically adjusted. The weights for endoscopic images range from 0.4 to 0.6, ultrasound images from 0.2 to 0.3, and CT images from 0.2 to 0.3. The weight adjustments are based on the clarity of each modality and clinical needs. For example, the weight of endoscopy is increased when the intraoperative field of view is clear, and the weights of CT and ultrasound are increased when the field of view is blurry. The output is a three-dimensional fused image with a spatial resolution of 1mm×1mm×1mm, which can realize three-dimensional rotation observation at any angle.

[0068] Error tolerance mechanism: When the registration error exceeds the threshold, re-registration is automatically triggered and an alert is issued. At the same time, the current registration data is saved for subsequent analysis of the cause of the error. In addition, it also has a manual registration trigger function, which allows doctors to start re-registration at any time according to operational needs.

[0069] (4) Data Interaction Unit

[0070] Transmission methods: Supports both 5G millimeter wave communication and PCIe 4.0 wired transmission. The 5G millimeter wave communication transmission rate is no less than 10Gbps, enabling wireless transmission and reducing cable interference in the operating room; the PCIe 4.0 wired transmission has a bandwidth of 8GB / s and a transmission latency of no more than 20ms. The two transmission methods can be automatically switched. When the 5G signal strength is below -85dBm, it will automatically switch to wired transmission.

[0071] Data format: Compression and encryption are processed simultaneously. The compression algorithm is H.265 encoding, which improves the compression rate by 50% compared to H.264 encoding, thus reducing the amount of data transmitted. The encryption algorithm is AES-256 encryption, which adopts a dynamic key mechanism. The key is updated every 30 seconds and stored through a dedicated key management module to prevent key leakage.

[0072] Interface compatibility: Supports bidirectional data interaction with surgical robots, navigation systems and other devices. Equipped with RS485, Ethernet and other interfaces, it can transmit fused image data to the surgical robot to guide the robot's precise operation. At the same time, it can receive the robot's operation position data and mark the instrument position on the fused image to achieve data closure.

[0073] (5) Display control unit

[0074] Display devices: including 4K ultra-high-definition medical displays and holographic projection devices; the brightness of the 4K ultra-high-definition medical displays is no less than 1000 cd / m². 2 The contrast ratio is no less than 1000:1, supports HDR display, and can restore image details; the projection accuracy of the holographic projection device is no less than 0.5mm, and the projection screen size can be adjusted within the range of 50-150cm. Doctors do not need to look down at the monitor and can intuitively observe the three-dimensional fused image through the holographic screen.

[0075] Interactive features: Supports rotation, scaling, and local magnification of 3D images, with a maximum magnification of 10x. Interpolation algorithms are used to maintain image clarity during magnification. Key structure annotations can be customized, such as marking blood vessels in red and nerves in blue. The annotation color and line thickness are adjustable, and the annotation information can be saved as a template for reuse in similar surgeries.

[0076] Parameter adjustment: Doctors can adjust parameters such as enhancement intensity, fusion weight, and display brightness in real time via touch screen or wireless handheld device. The adjustment response delay is no more than 10ms, and the image is updated in real time after parameter adjustment without the need to restart the system.

[0077] (6) Storage unit

[0078] It uses an SSD array with a storage capacity of no less than 2TB, consisting of four 2TB SSDs in RAID 5 mode, which has data redundancy. When a single SSD fails, the data can be reconstructed from the other SSDs to avoid data loss. It supports real-time storage of intraoperative raw images, enhanced images, fused images and surgical operation logs. The image storage resolution is consistent with the original acquisition resolution, with no compression loss.

[0079] It supports data backtracking and export, with the export format being the DICOM 3.0 standard format, which can be directly imported into the hospital's PACS system; the data backtracking function supports positioning by time axis, allowing doctors to input specific time points to quickly retrieve image data at that moment, facilitating postoperative review and medical research. Specific implementation examples:

[0081] The present invention will be further described below with reference to specific embodiments:

[0082] 1. Hardware Configuration

[0083] Image acquisition module: Utilizes an Olympus high-definition endoscope camera (model GIF-H190), featuring 1080P resolution and adjustable frame rates of 30fps and 60fps. In complex surgeries, the 60fps mode can be switched for improved image smoothness. A Philips high-frequency ultrasound probe (model L12-5) with a convex array design allows for a detection depth of 2-15cm and supports B-mode, M-mode, and color Doppler modes, selectable according to needs, such as color Doppler mode for viewing vascular blood flow. The DICOM interface card is Mindray Medical's MC-DICOM-01 model, supporting simultaneous connection of up to three imaging devices from different brands, with an interface transmission rate of 1Gbps.

[0084] Real-time Enhancement Unit: The server uses a Dell PowerEdge R7525 processor, equipped with an Intel Xeon Gold 6348 processor, which has 28 cores and 56 threads and a base frequency of 2.6GHz, capable of meeting the needs of multi-tasking parallel processing; the GPU is an NVIDIA A10040GB, equipped with an active cooling system to ensure stable operation during long-term operations; the memory is 128GB DDR4, with a frequency of 3200MHz, and supports memory error correction function to avoid system crashes caused by memory errors during data processing;

[0085] Data Interaction Unit: The 5G industrial module adopts Huawei ME909s-821, supports 5G NR SA / NSA dual-mode, with a maximum downlink speed of 2.3Gbps and a maximum uplink speed of 500Mbps. The module is equipped with an external antenna to enhance signal reception. The PCIe 4.0 data cable adopts Kingmax PCIe 4.0×4 data cable with a transmission length of 2 meters, supports hot-swapping, and facilitates equipment position adjustment during surgery.

[0086] Display control unit: The 4K ultra-high-definition medical display uses EIZO RadiForce RX850, with a screen size of 27 inches, a pixel density of 163 PPI, and supports DICOM Part 14 standard calibration to ensure that the image color is consistent with the actual tissue; the holographic projection device uses Microsoft HoloLens2, with a resolution of 1920×1080, a projection distance of 0.5-3m, a field of view of 52 degrees, and supports gesture interaction, allowing doctors to rotate and zoom the image through gestures;

[0087] Storage Unit: The SSD array uses four Samsung 870EVO 2TB SSDs in RAID 5 mode. This mode achieves data redundancy through distributed parity checking. In the event of a single SSD failure, lost data can be recovered from the other three SSDs and the parity check data, with a data recovery speed of over 100MB / s. The array controller is an LSIMegaRAID SAS 9400-8i, which supports cache protection and is equipped with a 1GB cache to improve data read and write speeds.

[0088] 2. Software Implementation

[0089] Operating System: Ubuntu 20.04LTS 64-bit version is used. This system has high stability, supports long-term updates and maintenance, and has good compatibility with GPU drivers and deep learning frameworks. The real-time kernel (RT_PREEMPT) is installed to reduce the system scheduling latency to less than 1ms, ensuring the priority execution of real-time image processing tasks.

[0090] Development framework: PyTorch 1.12 is used for model training, supporting dynamic computation graphs for easy algorithm debugging; TensorRT 8.4 is used for model inference optimization, converting PyTorch models to the TensorRT engine, reducing model size by 75% and increasing inference speed by 3 times through INT8 quantization; OpenCV 4.5 is used for image preprocessing and post-processing, such as image format conversion and edge detection, supporting GPU acceleration and achieving a processing speed 10 times faster than the CPU version.

[0091] Model Training: A dataset of 5000 thoracic surgery images from three tertiary hospitals was used. This dataset includes endoscopic, ultrasound, and CT images of 2000 lobectomies, 1500 esophagectomies, and 1500 lung nodule biopsies. All data underwent ethical review, and patient information was anonymized. The Adam optimizer was used during training, with an initial learning rate of 0.001, decreasing to 1 / 10 of the original rate every 2000 iterations. The batch size was 32, and the training iterations totaled 10000. The training loss function combined with the cross-entropy loss function and the Dice loss function was used to avoid model bias caused by data imbalance. After training, the model achieved 92% image enhancement accuracy and 95% registration accuracy on the test set.

[0092] Algorithm optimization: The computational load of the model is reduced by quantization training (INT8), converting the model parameters from 32-bit floating-point numbers to 8-bit integers, thereby reducing the consumption of computing resources while ensuring that the model accuracy loss does not exceed 5%. Layer fusion technology is adopted to merge layers such as convolution, activation, and batch normalization into a single computational unit, reducing the time loss of data transmission between layers. A dynamic batch processing mechanism is introduced to automatically adjust the batch size according to the complexity of real-time images. Large batch processing is used for simple images to improve efficiency, while small batch processing is used for complex images to ensure real-time performance.

[0093] 3. Intraoperative application procedure

[0094] This system can be applied to a variety of thoracic surgeries. The following sections will detail the application process using lobectomy and radical esophagectomy as examples:

[0095] (1) Application process of lobectomy surgery

[0096] Preoperative preparation: Within 24 hours before surgery, import the patient's preoperative CT images through the DICOM interface, complete image segmentation in the system software, and annotate anatomical structures such as lobes, pulmonary segmental arteries, veins, and bronchi; initiate model initialization, the system automatically loads the pre-trained enhancement and fusion model, adjusts model parameters according to the resolution and grayscale characteristics of the patient's CT images, and completes parameter calibration; check the connection status of each module, test the signal acquisition of the endoscope camera and ultrasound probe, and ensure that the data interaction unit transmits normally and the display device screen is clear;

[0097] Intraoperative data acquisition: After the surgery begins, the endoscopic camera is inserted into the thoracic cavity through the thoracoscopic incision. The camera angle is adjusted to align with the target lung lobe, and the 30fps frame rate mode is activated to acquire dynamic images of the surgical field in real time. The high-frequency ultrasound probe is placed close to the chest wall or inserted into the thoracic cavity through the operating port, and the 12MHz high-frequency mode is used to scan the vascular distribution of the target lung lobe to acquire ultrasound images. The system automatically and synchronously acquires both types of image data, with each frame acquisition time not exceeding 30ms.

[0098] Real-time enhancement: After receiving real-time image data, the enhancement unit first performs adaptive noise suppression, matching and optimizing parameters for surgical instrument reflection noise that may appear in endoscopic images and speckle noise in ultrasound images, achieving a noise removal rate of over 85%; then, edge detection is performed to automatically identify the edge contours of pulmonary vessels and bronchi, with an edge positioning error of no more than 0.5mm; finally, contrast enhancement is performed to increase the grayscale difference between blood vessels and lung tissue by 30%, making the blood vessel contours clearer, and the enhanced image is transmitted to the fusion unit in real time;

[0099] Multimodal fusion: The fusion unit receives enhanced endoscopic, ultrasound, and preoperative CT images. First, it uses the SIFT algorithm to extract feature points, extracting bronchial bifurcation and pulmonary segmental vascular nodes from the CT images, and corresponding feature points from the endoscopic images, extracting a total of 60-80 feature points. A pressure sensor collects patient respiratory motion data, acquiring displacement data every 10ms. The system calculates tissue deformation based on the displacement data and corrects the feature point matching relationship in real time. The TPS algorithm is used to perform deformation registration on the CT images, aligning the anatomical structures in the CT images with the real-time intraoperative images, with a registration error controlled within 0.8mm. Finally, a weighted fusion strategy is used, adjusting the weights according to the intraoperative visual clarity. When the endoscopic view is clear, the weight of the endoscopic image is set to 0.6, the weight of the ultrasound image to 0.2, and the weight of the CT image to 0.2, outputting a three-dimensional fused image. The fusion process delay does not exceed 20ms.

[0100] Surgical guidance: Doctors observe the fused images through a 4K medical monitor or holographic projection device. In the fused images, pulmonary segmental arteries, veins, and bronchi are marked with different colors (arteries red, veins blue, and bronchi green). Doctors can zoom in on the target area through the touch screen to view the details of the blood vessel branches and guide the surgical instruments to separate the blood vessels and remove the lung lobes. When the instruments approach the marked blood vessels, the system automatically issues a prompt to remind the doctor to avoid accidental damage.

[0101] Data storage: The storage unit stores intraoperative raw images, enhanced images, fused images, and surgical operation logs in real time. The image data is stored in categories according to the surgical stage (such as preoperative preparation, lobectomy, vascular treatment, lobectomy, and postoperative hemostasis). Each image segment contains a timestamp, which is convenient for retrospective review after surgery. After the operation, the doctor can export the image data in DICOM format and the operation log in PDF format through the system and import them into the hospital's PACS system for archiving.

[0102] (2) Application process of radical esophagectomy

[0103] Preoperative preparation: Import the patient's preoperative CT and MRI images, mark the location of the esophageal tumor, and anatomy of the surrounding blood vessels, lymphatic vessels, trachea, etc.; adjust the ultrasound probe parameters, switch to 5MHz low frequency mode, and test the probe's detection effect on the surrounding tissues; complete the model parameter calibration, and optimize the edge detection threshold of the enhancement algorithm for the characteristics of deep surgical field and multiple tissue layers in esophageal cancer surgery.

[0104] Intraoperative data acquisition: An endoscopic camera is inserted into the esophageal lumen to acquire images of the tumor and mucosa within the esophagus; an ultrasound probe is placed on the external chest wall or inserted into the pleural cavity during the operation to detect the distribution of blood vessels and lymphatic vessels around the esophagus; the system acquires both types of images simultaneously, ensuring that the image time synchronization error does not exceed 10ms;

[0105] Real-time enhancement and fusion: The enhancement unit optimizes contrast enhancement parameters for esophageal mucosal images to highlight the boundary between the tumor and normal mucosa; the fusion unit combines preoperative MRI images (clearly showing soft tissue) with intraoperative images, focusing on the positional relationship between the esophagus and trachea during the registration process to avoid intraoperative tracheal damage; the fused images mark the tumor boundary, the location of blood vessels around the esophagus and the trachea to assist doctors in determining the resection range.

[0106] 4. Clinical test results

[0107] To verify the effectiveness of the system, a clinical trial was conducted in three tertiary hospitals, enrolling a total of 200 patients undergoing thoracic surgery, including 100 patients undergoing lobectomy, 50 patients undergoing radical esophagectomy, and 50 patients undergoing pulmonary nodule biopsy. Patients were randomly divided into an experimental group (using this system) and a control group (using traditional imaging techniques), with 100 patients in each group. The test results are as follows:

[0108] Surgical time: The average surgical time in the experimental group was 125 minutes, while that in the control group was 145 minutes, representing a 13.8% reduction in surgical time. Specifically, the average time for lobectomy in the experimental group was 110 minutes, compared to 130 minutes in the control group, a reduction of 15.4%. The average time for radical esophagectomy in the experimental group was 150 minutes, compared to 175 minutes in the control group, a reduction of 14.3%. The average time for lung nodule biopsy in the experimental group was 25 minutes, compared to 35 minutes in the control group, a reduction of 28.6%.

[0109] Intraoperative injury rate: The vascular injury rate in the experimental group was 1% (1 case), while that in the control group was 8% (8 cases); the nerve injury rate in the experimental group was 0%, while that in the control group was 3% (3 cases); the lymphatic vessel injury rate in the experimental group was 1% (1 case), while that in the control group was 6% (6 cases).

[0110] Postoperative complication rates: The incidence of postoperative massive hemorrhage in the experimental group was 1% (1 case), while it was 4% (4 cases) in the control group; the incidence of pneumothorax in the experimental group was 2% (2 cases), while it was 5% (5 cases) in the control group; the incidence of lymphorrhea in the experimental group was 1% (1 case), while it was 4% (4 cases) in the control group; the overall complication rate in the experimental group was 3%, while it was 7% in the control group, representing a 57.1% reduction in the complication rate in the experimental group.

[0111] Doctor satisfaction: A satisfaction survey was conducted on 30 doctors who participated in the surgery after the operation. A 5-point rating system was used. The average satisfaction score of the experimental group was 4.8 points, while that of the control group was 3.2 points. The satisfaction score of the experimental group was significantly higher than that of the control group. The doctors reported that the main advantages were clear images, accurate registration, and low operation delay.

[0112] Clinical trials have verified that this system meets the clinical needs of thoracic surgery and can effectively improve surgical safety and efficiency.

[0113] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.

Claims

1. A real-time image enhancement and fusion system for thoracic surgery, characterized in that, include: The image acquisition module is used to acquire real-time endoscopic images, ultrasound images, and preoperative CT / MRI images during thoracic surgery. The image acquisition module includes a 1080P high-definition endoscopic camera, a high-frequency ultrasound probe, and a DICOM 3.0 standard interface. The frame rate of the 1080P high-definition endoscopic camera is not less than 30fps, and the frequency range of the high-frequency ultrasound probe is 5-12MHz. The real-time enhancement unit, connected to the image acquisition module, employs an improved U-Net convolutional neural network model based on an attention mechanism to perform adaptive noise suppression, edge enhancement, and contrast enhancement processing on real-time endoscopic and ultrasound images. The encoder of the model is a ResNet50 backbone network, and the decoder includes spatial attention gates and channel attention gates. The processing frame rate of the real-time enhancement unit is no less than 30fps, and the peak signal-to-noise ratio of the output image is no less than 35dB, and the structural similarity index is no less than 0.

9. The multimodal fusion unit is connected to the real-time enhancement unit and the image acquisition module respectively. It adopts a dynamic registration algorithm of feature extraction-dynamic update-registration optimization, and corrects the feature point matching relationship by combining intraoperative respiratory motion data. It achieves three-dimensional registration of the enhanced real-time image and the preoperative CT / MRI image through thin plate spline interpolation algorithm. Then, it outputs the multimodal fusion image based on the weighted fusion strategy of maximizing mutual information. The registration error of the multimodal fusion unit does not exceed 1mm, and the spatial resolution of the output multimodal fusion image is not less than 1mm×1mm×1mm. The data interaction unit is used to realize low-latency data transmission between modules, supporting 5G millimeter wave communication and PCIe 4.0 wired transmission. The transmission rate of the 5G millimeter wave communication is not less than 10Gbps, and the transmission latency of the PCIe 4.0 wired transmission is not more than 20ms. The data interaction unit uses H.265 encoding and AES-256 encryption to synchronously process data. The display control unit is connected to the multimodal fusion unit and adopts a 4K ultra-high-definition medical display or holographic projection device. It supports three-dimensional image rotation, scaling, local magnification and key structure annotation. The adjustment response delay of the display control unit does not exceed 10ms. The storage unit uses an SSD array with a storage capacity of no less than 2TB, which is used to store intraoperative image data and operation logs in real time and supports data export in DICOM 3.0 format.

2. The real-time intraoperative image enhancement and fusion system for thoracic surgery according to claim 1, characterized in that: The processing flow of the real-time enhancement unit includes: S1: Adaptive noise suppression of real-time images is performed using a Gaussian mixture model; S2: Extract the edges of anatomical structures using an improved Canny operator, which has an adaptive threshold adjustment function; S3: An improved CLAHE algorithm is used to enhance the contrast of key areas. The improved CLAHE algorithm introduces a regional weight allocation mechanism. S4: Output the enhanced image, with the processing delay of the real-time enhancement unit not exceeding 30ms.

3. The real-time image enhancement and fusion system for thoracic surgery according to claim 1, characterized in that: The dynamic registration algorithm of the multimodal fusion unit specifically includes: S1: The SIFT algorithm is used to extract anatomical feature points from the enhanced real-time images and the preoperative CT / MRI images. The number of anatomical feature points extracted from each image is no less than 50. S2: Intraoperative respiratory motion data are collected through a pressure sensor with a sampling frequency of 100Hz. The feature point matching relationship is corrected in real time based on the collected intraoperative respiratory motion data. S3: Image deformation registration is achieved using a thin-plate spline interpolation algorithm; S4: When the registration error exceeds 1mm, re-registration will be automatically triggered and an alarm will be issued.

4. The real-time image enhancement and fusion system for thoracic surgery according to claim 1, characterized in that: In the weighted fusion strategy of the multimodal fusion unit, the weight range of endoscopic images is 0.4-0.6, the weight range of ultrasound images is 0.2-0.3, and the weight range of CT / MRI images is 0.2-0.3, and the weights can be dynamically adjusted.

5. The real-time intraoperative image enhancement and fusion system for thoracic surgery according to claim 1, characterized in that: The display control unit uses a medical display with a brightness of no less than 1000 cd / m². 2 The contrast ratio is not less than 1000:1, and the holographic projection device used in the display control unit has a projection accuracy of not less than 0.5mm.

6. The real-time intraoperative image enhancement and fusion system for thoracic surgery according to claim 1, characterized in that: The image acquisition module is equipped with an anti-interference shielding component, which adopts a double-layer copper mesh shielding structure and has a shielding effect of over 60dB. The data interaction unit supports bidirectional data interaction with the surgical robot and navigation system.

7. The real-time intraoperative image enhancement and fusion system for thoracic surgery according to claim 1, characterized in that: The real-time enhancement unit is equipped with an NVIDIA A100 GPU, which optimizes model inference speed through the TensorRT inference engine. The NVIDIA A100 GPU has a video memory capacity of 40GB.

8. The real-time intraoperative image enhancement and fusion system for thoracic surgery according to claim 1, characterized in that: The storage unit uses an SSD array containing four SSDs in RAID 5 mode. The SSD array is equipped with an array controller that supports cache protection and has a cache capacity of 1GB.