Medical endoscope device and system based on naked eye 3D display
By combining a dual-spectrum endoscopic probe with the YOLOv5 and U-Net models, the problems of surgical smoke interference and sensor performance bottlenecks were resolved, achieving high-precision naked-eye 3D visualization, improving surgical safety and efficiency, and reducing equipment maintenance costs.
Patent Information
- Application Number
- CN202510477182.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-09-23
AI Technical Summary
Existing 4K high-definition naked-eye 3D endoscopy technology, when using high-frequency electric knives, lasers or ultrasonic knives, has problems such as surgical smoke interfering with the field of view and sensor performance bottlenecks, resulting in reduced surgical accuracy and safety. Negative pressure suction devices have problems such as low efficiency, high cost and interference with surgical operations.
A dual-spectrum endoscope probe combined with the YOLOv5 network and U-Net model is used for smoke detection and defogging. Through RGB and NIR image fusion, combined with FPGA hardware-level black level correction and TOF distance adjustment, naked-eye 3D rendering is achieved, and sensor temperature control is dynamically adjusted to avoid the use of negative pressure suction devices.
The smoke detection accuracy has been improved to 99.5%, and the clarity of the generated defogging images has been increased by 60%, which has reduced equipment maintenance costs, improved surgical continuity and operational safety, adapted to complex surgical environments, suppressed sensor noise, and provided a clear 3D field of view.
Smart Images

Figure CN120678375A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of detection devices, and in particular to a material detection device for municipal engineering construction. Background Art
[0002] With the continuous advancement of medical technology, minimally invasive surgery has become an important tool in modern surgical treatment. Performed through tiny incisions, this type of surgery significantly reduces patient trauma and recovery time. To further enhance surgical precision and visualization, 4K high-definition, glasses-free 3D endoscopy technology has emerged. Through high-resolution image acquisition and real-time 3D rendering, this technology provides surgeons with an unprecedented surgical field of view, significantly improving surgical efficiency and safety.
[0003] However, in practical application, 4K high-definition glasses-free 3D endoscopy technology still faces a number of technical challenges. In particular, when using energy devices such as high-frequency electrosurgery, lasers, or ultrasonic scalpels for cutting or hemostasis, large amounts of surgical smoke are generated. This smoke, primarily composed of vaporized tissue particles and blood aerosols, can quickly fill the surgical field, compromising surgical precision and safety.
[0004] Currently, the solution to the problem of surgical smoke mainly relies on negative pressure suction devices. Through negative pressure suction, smoke can be effectively extracted from the surgical area. This method has many limitations. First, the layout and performance of the negative pressure suction device are often limited by surgical instruments and surgical space, resulting in limited suction efficiency. Second, frequent negative pressure suction may interfere with the surgical operation space and increase surgical risks. In addition, the negative pressure suction device requires regular filter replacement, which not only increases surgical costs but may also affect the continuity and efficiency of the operation. In addition, when the 4K high-definition endoscope sensor module operates at high resolution, the dark current generated during the photon conversion process will increase exponentially with increasing temperature. Especially in long-term surgical modes requiring high frame rates (FPS), the dynamic range of the sensor module will drop significantly and the number of hot pixels will increase sharply.
[0005] Therefore, this application is filed. Summary of the Invention
[0006] In order to solve the problems raised in the above background technology, namely, surgical smoke interference and sensor performance bottleneck, a medical endoscopy device and system based on naked-eye 3D display is proposed.
[0007] A medical endoscope device based on naked-eye 3D display includes a display module and a dual-spectrum endoscope probe, including a probe housing, and a first sensor head, a second sensor head, an optical lens group, and a micro copper tube disposed in the probe housing, wherein a baseline distance between the first sensor head and the second sensor head is set; wherein the first sensor head acquires an RGB image and the second sensor head acquires a NIR image; a controller electrically connected to the dual-spectrum endoscope probe and configured to: receive a dual-spectrum image of the RGB image and the NIR image, and align the timestamps of the dual-spectrum images and activate the controller by an FPGA trigger signal; Dynamic black level correction; smoke detection is performed on the bispectral image using the YOLOv5 network model. If smoke is present, a dehazing network model is constructed by automatically searching and optimizing the U-Net model of the network structure, and the dehazing image and derivative image are output; based on the dehazing image or the original bispectral image, the RGB image and the NIR image are multimodally fused to obtain a fused image; the fused image is disparity layered and then recombined and the disparity map is optimized to obtain the intermediate layer image; the disparity is dynamically adjusted based on the baseline distance and TOF distance, and a dynamic 3D image is output through naked-eye 3D rendering.
[0008] In a further solution of the present application, in the setting of the YOLOv5 network model, the first layer is configured as a standard convolutional layer, one Bottleneck is retained, and the residual connection is removed for structural optimization; in the loss function, a composite loss of at least three parameters is used to calculate the edge loss; wherein, smoke detection on a dual-spectral image through the YOLOv5 network model includes: inputting image frames of an RGB image and an NIR image, obtaining a sampling feature map through comparative reasoning; performing threshold segmentation on the sampling feature map and then performing morphological separation to output a binary smoke mask; obtaining a confidence score of the binary smoke mask and performing a threshold judgment to output whether smoke exists or not.
[0009] In a further scheme of the present application, in the dehazing network model, the gradient-based DARTS algorithm searches in the image dataset of the bispectral image to generate an optimal structure U-Net model containing an MBConv block; wherein, the dehazing network model is constructed by automatically searching the U-Net model with the optimized network structure, and the dehazing image and the derived image are output, including: inputting the image frame of the binary smoke mask and the RGB image; performing deep separable convolution feature extraction through the optimal structure U-Net model, and attention-guided jump connection to output the dehazing image and the derived image.
[0010] In a further solution of the present application, multimodal fusion is performed on the RGB image and the NIR image to obtain a fused image, including: performing feature point matching on the RGB image and the NIR image for affine transformation; inputting the dehazed image and the aligned NIR image, obtaining a fusion weight based on a preset smoke attenuation coefficient, and fusing the RGB image and the NIR image to obtain a fused image.
[0011] In a further scheme of the present application, the fused image is disparity layered and then reorganized and the disparity map is optimized to obtain an intermediate layer map, including: inputting a stereo image pair of the fused image, dividing the stereo image pair into equal parts along the channel dimension, wherein the first branch performs conventional convolution, the second branch fuses the result of the first branch, and the third branch fuses the result of the second branch; reconstructing the right disparity according to the left disparity of the stereo image pair, calculating the inconsistent area between the left disparity and the right disparity, and filling the inconsistent area with surrounding pixels to obtain an intermediate layer map.
[0012] In a further solution of the present application, the parallax is dynamically adjusted based on the baseline distance and the TOF distance, and a dynamic 3D image is output through naked-eye 3D rendering, including: performing parallax scaling and / or boundary cropping according to the intermediate layer image and the TOF distance and the baseline distance to obtain a comfortable parallax map; obtaining a light field image based on the fused image, the comfortable parallax map, the preset grating parameters, and based on multi-viewpoint generation and sub-pixel rendering; and generating a lens control signal through the light field image to output a dynamic 3D image.
[0013] In a further solution of the present application, the dual-spectrum endoscope probe also includes a heat exchange plate; the probe housing includes a probe part and an operating part, the first sensor head, the second sensor head, the optical lens group and the micro copper tube are all located on the inside of the probe part, the heat exchange plate is located on the inside of the operating part, one end of the micro copper tube extends into the probe part and contacts the first sensor head and the second sensor head; the other end contacts the heat exchange plate.
[0014] In a further solution of the present application, a micro temperature sensor is also provided in the probe housing; the controller is electrically connected to the micro temperature sensor and is also configured to: obtain real-time temperature data of the micro temperature sensor; selectively adjust the power of the heat exchanger or switch the operating mode of the dual-spectrum endoscope probe according to the range mapped by the real-time temperature data.
[0015] In a further scheme of the present application, the power of the heat exchanger is selectively adjusted or the operating mode of the dual-spectrum endoscope probe is switched according to the range mapped by the real-time temperature data, including: when the temperature data is between the preset first temperature threshold and the second temperature threshold, the power of the heat exchanger is adjusted by controlling the current temperature data pid; when the temperature data is greater than the preset second temperature threshold, the dual-spectrum endoscope probe is controlled to dynamically skip frames through exponential backoff and control frequency reduction.
[0016] The second aspect of the present application also provides a medical endoscopy system based on naked-eye 3D display, including a medical endoscopy device based on naked-eye 3D display as described above, the medical endoscopy device including a frame and an operating module corresponding to the dual-spectrum endoscope probe; a mobile terminal for communicating with the controller of the medical endoscopy device.
[0017] Compared with the prior art, the present invention has the following beneficial effects:
[0018] This application uses dual-spectral imaging and U-Net model defogging technology to directly eliminate the visual interference of surgical smoke without relying on a negative pressure suction device, avoiding the impact of traditional suction methods on surgical space and operation fluency;
[0019] By optimizing the YOLOv5 network model for smoke detection, smoke detection accuracy has been increased to 99.5%, while maintaining low model complexity and fast inference speed (response speed <30ms), reducing the risk of false positives and missed detections during surgery. The automatically searched and optimized U-Net dynamically adjusts the network depth and width to adapt to different smoke concentrations, resulting in dehazed images with over 60% higher clarity, effectively restoring tissue detail and vascular texture, providing doctors with a clearer field of view. Combining the advantages of RGB and NIR images makes the system more robust in various surgical scenarios (such as bleeding and reflections), adapting to complex surgical environments. Furthermore, combined with FPGA hardware-level black level correction and the ability to effectively suppress dark current noise and hot pixel issues of 4K sensors under high-temperature and high-frame-rate conditions, the system achieves stable naked-eye 3D visualization while maintaining ultra-high-definition image quality. This reduces equipment maintenance costs, improves surgical continuity and operational safety, and compensates for RGB information loss in harsh environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a schematic structural diagram of a medical endoscopy device based on naked-eye 3D display provided in an embodiment of the present application.
[0021] Figure 2 A schematic diagram of a portion of the structure of a dual-spectrum endoscope probe in a medical endoscope device based on naked-eye 3D display provided in an embodiment of the present application;
[0022] Figure 3 A partial cross-sectional view of a medical endoscopy device based on naked-eye 3D display provided in an embodiment of the present application;
[0023] Figure 4 This is a general flow chart of a controller for a medical endoscopy device based on naked-eye 3D display provided in an embodiment of the present application;
[0024] Figure 5A node diagram of the medical endoscopy device controller based on naked-eye 3D display provided in an embodiment of the present application executing step S20; and
[0025] Figure 6 A further node diagram of the medical endoscopy device controller based on naked-eye 3D display provided in an embodiment of the present application executing step S20. Figure 7 This is a module diagram of a medical endoscopy system based on naked-eye 3D display provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0027] The technical solution of the present invention is further described in detail below through specific embodiments and in combination with the accompanying drawings: A medical endoscopy device 100 based on naked-eye 3D display includes a display module 10, a dual-spectrum endoscope probe 20 and a controller 30; the dual-spectrum endoscope probe 20 includes a probe housing 21, wherein a ring-shaped light-enhancing lamp 22 is provided on the end face of the probe housing 21, as well as a first sensor head 23, a second sensor head 24, an optical lens group 25 and a micro copper tube 26 arranged in the probe housing 21.
[0028] In the present application, the baseline spacing distance between the first sensor head 23 and the second sensor head 24 is set to form stereoscopic vision and enhance depth perception.
[0029] The first sensor head 23 acquires RGB images, and the second sensor head 24 acquires NIR images. In a specific embodiment, the first sensor head 23 uses a 4K high-definition RGB-CMOS, which provides 400-700nm visible light images with a resolution of 3840×2160@60fps. The second sensor head 24 uses an NIR sensor to penetrate smoke / blood and identify tissue oxygenation status. The optical lens group 25 uses a dual-channel spectroscopic prism as shown in the figure and can be electrically zoomed to ensure that the RGB and NIR optical paths are coaxial.
[0030] The micro copper tube 26 forms a spiral cooling channel that fits the first sensor head 23 and the second sensor head 24 to maintain the probe temperature, thereby ensuring that its performance is not affected when it works continuously for a long time.
[0031] In the present application, the baseline distance between the first sensing head 23 and the second sensing head 24 is 8 mm.
[0032] Based on a variation of the present application, the baseline distances of the first sensor head 23 and the second sensor head 24 can be designed to be adjustable, and a piezoelectric ceramic fine-tuning mechanism is used to compensate for mechanical deformation.
[0033] FPGA hardware trigger signal synchronizes dual sensor exposure,
[0034] The controller 30 is electrically connected to the dual-spectrum endoscope probe 20 and the display module 10 and is configured to:
[0035] Step S10: receiving a dual-spectrum image of an RGB image and an NIR image, aligning the timestamps of the dual-spectrum images through an FPGA trigger signal, and performing dynamic black level correction;
[0036] Step S20: Perform smoke detection on the bispectral image using the YOLOv5 network model. If smoke is present, a defogging network model is constructed by automatically searching and optimizing the U-Net model of the network structure, and the defogging image and the derived image are output.
[0037] Step S30: performing multimodal fusion on the RGB image and the NIR image according to the defogging image or the original bispectral image to obtain a fused image;
[0038] Step S40: performing disparity layering on the fused image, recombining the image, and performing disparity map optimization to obtain an intermediate layer image;
[0039] Step S50 : dynamically adjust the parallax based on the baseline distance and the TOF distance, and output a dynamic 3D image through naked-eye 3D rendering.
[0040] First, input RGB image (visible light band) and NIR image (near infrared band);
[0041] The RGB camera and near-infrared (NIR) camera are synchronously triggered by the FPGA to ensure the timestamp alignment of the dual-spectral images; the hardware-level signal synchronization accuracy is controlled at the microsecond level to avoid time offset of multimodal data.
[0042] The dark current noise of the sensor is detected in real time, the black level offset is calculated per frame, and the dynamic compensation formula is applied to the RGB and NIR images respectively:
[0043] I corrected (x,y)=I raw (x,y)-B dynamic (t)
[0044] Where Bdynamic(t) is the average value of black level over time; thus, a time-synchronized and noise-corrected dual-spectral image (RGB+NIR) is obtained. FPGA hardware acceleration ensures real-time performance, is suitable for high-speed scenes, improves high-speed performance, and dynamic correction improves the image signal-to-noise ratio and reduces subsequent algorithm errors.
[0045] The preprocessed bispectral image (RGB+NIR) is input into the improved YOLOv5s network (with the NIR channel input branch added). If a smoke area is detected, the dehazing process is triggered. Otherwise, the process jumps to S30 and uses the bispectral image as input.
[0046] Simply put, the improved YOLOv5s network is used to detect smoke on the bispectral image. When there is no smoke, the original bispectral image is output, otherwise the dehazed image and the derivative image are output.
[0047] In this application, based on the characteristics of the smoke area, the improved YOLOv5s network is automatically optimized, and through the automatic search and optimization of the U-Net network structure, the number of jump connection layers and the size of the convolution kernel are continuously adaptively adjusted to adapt to different smoke scenes, thereby avoiding the overfitting problem of the traditional dehazing algorithm, and finally outputting the dehazed RGB image and the corresponding derivative image.
[0048] If the image has been dehazed, weighted wavelet fusion (RGB texture + NIR radiation feature) is used. Conversely, if the image has not been dehazed, pyramid fusion is performed directly on the bispectral image to output a fused image, thereby improving the image information density and retaining more features.
[0049] Based on the fused image, the initial disparity is calculated and the disparity is completed to provide a depth basis for 3D rendering, optimize the disparity to reduce artifacts and reduce the occlusion area; finally, the intermediate layer image and disparity map are input and combined with TOF measurement.
[0050] The absolute depth value of the distance-corrected disparity map is used to generate multi-view images using rasterization or holographic technology, and finally output dynamic 3D images for naked eyes, improving depth accuracy and avoiding the scale ambiguity problem of pure binocular matching.
[0051] In summary, this application uses dual-spectral imaging and U-Net model defogging technology to directly eliminate the visual interference of surgical smoke without relying on a negative pressure suction device, avoiding the impact of traditional suction methods on surgical space and operation fluency;
[0052] By optimizing the YOLOv5 network model for smoke detection, smoke detection accuracy has been increased to 99.5%, while maintaining low model complexity and fast inference speed (response speed <30ms), reducing the risk of false positives and missed detections during surgery. The automatically searched and optimized U-Net dynamically adjusts the network depth and width to adapt to different smoke concentrations, resulting in dehazed images with over 60% higher clarity, effectively restoring tissue detail and vascular texture, providing doctors with a clearer field of view. Combining the advantages of RGB and NIR images makes the system more robust in various surgical scenarios (such as bleeding and reflections), adapting to complex surgical environments. Furthermore, combined with FPGA hardware-level black level correction and the ability to effectively suppress dark current noise and hot pixel issues of 4K sensors under high-temperature and high-frame-rate conditions, the system achieves stable naked-eye 3D visualization while maintaining ultra-high-definition image quality. This reduces equipment maintenance costs, improves surgical continuity and operational safety, and compensates for RGB information loss in harsh environments.
[0053] like Figure 5 In the YOLOv5 network model settings, the first layer is configured as a standard convolutional layer, one bottleneck is retained, and the residual connection is removed for structural optimization. In the loss function, a composite loss with at least three parameters is used to calculate the edge loss. Among them, smoke detection on bispectral images using the YOLOv5 network model includes:
[0054] Step S21: input image frames of RGB image and NIR image, and obtain sampling feature map through comparative reasoning;
[0055] Step S22: performing threshold segmentation on the sampled feature map and then performing morphological separation to output a binary smoke mask;
[0056] Step S23: Obtain the confidence score of the binary smoke mask and perform a threshold judgment to output whether smoke exists or does not exist.
[0057] It can be understood that the first standard convolution layer configuration uses dual-spectral 4-channel data (3 channels of RGB and 1 channel of NIR), uniformly processes multi-spectral input, and extracts spatial features through 3×3 convolution; retains 1 Bottleneck and optimizes the single Bottleneck: and removes residual connections to increase training speed; the number of channels is dynamically calculated to reduce the number of parameters, and an attention mechanism module is added to improve the response of smoke edge features.
[0058] Directly processing four-channel input (RGB + NIR) avoids the parameter expansion of traditional two-branch structures, reducing computational complexity by approximately 40% compared to parallel convolution. Standard convolution is 1.8 times faster than depthwise separable convolution on NPU accelerators, making it more hardware-friendly. Edge loss calculations are performed using a composite loss with at least three parameters, exemplified by the BSIoU, SCL, and EGL loss components, resulting in significantly higher efficiency than traditional loss calculations.
[0059] During execution, the RGB frame (H×W×3) and the NIR frame (H×W×1) are first input, the NIR channel is copied to a 3-channel pseudo RGB to achieve channel alignment of the image frame, and then feature comparison is performed to output a sampled feature map; the feature map is used for adaptive threshold segmentation, morphological optimization, and connected domain filtering to output a binary smoke mask (H×W×1); then the binary mask and the original feature map are input, and two calculation dimensions are used: morphological confidence and feature consistency confidence. The two confidences are weighted and fused to obtain a confidence score; a threshold judgment is performed on the confidence score to output whether smoke exists or not.
[0060] like Figure 6 In the dehazing network model, the gradient-based DARTS algorithm searches in the image dataset of bispectral images to generate the optimal structure U-Net model containing MBConv blocks; and constructs the dehazing network model by automatically searching the U-Net model with the optimized network structure, outputting the dehazing image and the derived image, including:
[0061] Step S24: input the binary smoke mask and the image frame of the RGB image;
[0062] Step S25: Perform depth-wise separable convolution feature extraction through the optimal structure U-Net model, and perform attention-guided jump connection to output the dehazed image and the derived image.
[0063] It can be understood that the RGB image frame and the binary smoke mask are taken as input; through the U-Net forward propagation, the encoder path (downsampling) and the decoder path (upsampling) are configured, and the attention jump connection is adopted to finally output the dehazed image and the derivative image; among them, the derivative image is the transmittance map, namely the T-map.
[0064] The DARTS algorithm automatically searches for the optimal MBConv configuration on bispectral data, improving dehazing accuracy compared to manually designed networks. Furthermore, the MBConv block reduces the number of parameters, which improves inference speed. For 4K endoscopes, real-time processing exceeds 50fps, achieving high-precision, real-time dehazing.
[0065] Multimodal fusion of RGB images and NIR images is performed to obtain a fused image, including:
[0066] Step S31, performing feature point matching on the RGB image and the NIR image with affine transformation;
[0067] Step S32: input the defogging image and the registered NIR image, obtain a fusion weight based on a preset smoke attenuation coefficient, and fuse the RGB image and the NIR image to obtain a fused image.
[0068] Input the original RGB image and the original NIR image, eliminate the parallax error between the RGB and NIR images through feature point matching (such as SIFT+ORB), ensure pixel-level alignment, and obtain the NIR image after registration with the RGB image. Then, the dehazed RGB image, the registered NIR image, the transmittance map T-map, and the dynamic smoke attenuation coefficient β are combined to perform the following operations:
[0069] Calculate the fusion weight: α = exp(-β * T-map); expand α from an H × W × 1 matrix to an H × W × 3 matrix. Then perform pixel-level fusion: the fused image matrix = α * dehazed RGB image matrix + (1 - α) * registered NIR image matrix. The resulting fused image shows that the haze areas are dominated by the highly penetrating NIR image, while the non-haze areas are dominated by the RGB image.
[0070] Similarly, through adaptive fusion driven by physical models, both penetration and authenticity can be achieved in smoke scenarios with high quantification, comprehensively improving surgical safety and efficiency.
[0071] In step S40, the fused image is disparity layered and then reassembled and the disparity map is optimized to obtain an intermediate layer map, including:
[0072] Step S41: Input a stereo image pair of the fused image, divide the stereo image pair into equal parts along the channel dimension, wherein the first branch performs conventional convolution, the second branch fuses the result of the first branch, and the third branch fuses the result of the second branch;
[0073] Step S42: reconstruct the right disparity according to the left disparity of the stereo image pair, calculate the inconsistent area between the left disparity and the right disparity, and fill the inconsistent area with surrounding pixels to obtain an intermediate layer image.
[0074] The stereo pairs are divided into three groups along the channel dimension, where the first branch uses conventional convolution to extract basic disparity features:
[0075] branch1=Conv2D(64,kernel=3)(input_group1);
[0076] The second branch uses skip connections to fuse the results of the first branch and enhance context awareness:
[0077] branch2=Conv2D(64,kernel=3)(branch1+input_group2);
[0078] The third branch fuses the second branch to capture large-scale disparity relationships:
[0079] branch3=DilatedConv2D(64,kernel=3,dilation=2)(branch2+input_group3);
[0080] Channel equalization reduces video memory usage by 50%, enabling real-time processing of 4K endoscopes.
[0081] The algorithm takes the three-branch feature map and the initial left / right disparity as input, reconstructs the right disparity map through visual translation from the left disparity map, and calculates the left and right disparity difference mask, which identifies occluded or mismatched areas. Inconsistent areas are weighted-filled with surrounding valid disparity values (see the OpenCV inpainting algorithm). The filled intermediate disparity map ensures 3D coherence of blood vessels and nerves, preventing display disruption in naked-eye 3D images. This significantly reduces the risk of vessel misidentification during thoracoscopic lobectomy, for example. It also exhibits high robustness to dynamic scenes, capable of handling temporary occlusions caused by endoscope movement, such as parallax jumps caused by an instrument occluding tissue.
[0082] In summary, through the collaborative optimization of three branches, multi-scale features are covered and the parallax accuracy of complex anatomical scenes is improved; and compared with the previous method of directly leaving the occluded area black, this application uses the above steps to adaptively fill in the surrounding pixels, and combines the equal division of channels and lightweight branch design to improve the continuity of 3D navigation and reduce the instrument positioning error.
[0083] In step S50, the parallax is dynamically adjusted based on the baseline distance and the TOF distance, and a dynamic 3D image is output through naked-eye 3D rendering, including:
[0084] Step S51: performing disparity scaling and / or boundary clipping according to the intermediate layer image, the TOF distance, and the baseline distance to obtain a comfortable disparity map;
[0085] Step S52: obtaining a light field image based on the fused image, the comfortable disparity map, preset grating parameters, multi-viewpoint generation, and sub-pixel rendering;
[0086] In step S53 , a lens control signal is generated through the light field image to output a dynamic 3D image.
[0087] The intermediate layer disparity map (from S42), the detected TOF distance (based on the actual viewing distance measured by TOF, such as the doctor's current viewing distance is 50 cm, this value can also be a non-detected fixed value), and the known baseline distance are used as input. The previous naked-eye 3D display parallax comfort zone is used to dynamically scale the parallax or cut off extreme values that exceed the naked-eye 3D display parallax comfort zone; the parallax is automatically reduced to avoid 3D dizziness.
[0088] The fused image is then taken as input based on DICOM data, a comfortable disparity map, and known grating parameters. First, sub-images of different perspectives are generated according to the comfortable disparity map. The grating arrangement of the naked-eye 3D display is matched to output a multi-viewpoint synthesized light field image. The adjustable focus lens is driven according to the light field image for dynamic rendering. Specifically, the closed loop from TOF distance and disparity adjustment to light field update is completed, thereby obtaining a high-definition dynamic 3D image that can be displayed based on naked-eye 3D.
[0089] It is understandable that traditional fixed parallax causes visual fatigue and is not suitable for long-term surgery. This improved solution dynamically adjusts the comfortable parallax through TOF, and achieves a light field naked-eye 3D effect that integrates TOF and binocular baseline. It is very suitable for high-precision surgical scenarios such as thoracoscopic lobectomy or neuroendoscopic surgery.
[0090] The dual-spectrum endoscope probe also includes a heat exchange plate (not shown in the figure). The probe housing 21 includes a probe part 211 and an operating part 212. The first sensor head 23, the second sensor head 24, the optical lens group 25 and the micro copper tube 26 are all located on the inside of the probe part. The heat exchange plate is located on the inside of the operating part 212. One end of the micro copper tube 26 extends into the probe part 211 and contacts the first sensor head 23 and the second sensor head 24; the other end contacts the heat exchange plate.
[0091] It can be understood that the probe part 211 is the front end, used to be inserted into the patient's body, with a diameter of ≤5mm, the operating part 212 is the part held by the doctor, the heat exchange plate is arranged inside, one end is the device interface and several operating buttons are provided on the shell; the micro copper tube 26 runs through the probe part to the operating part, the probe end is attached to the sensor head CMOS chip, and the operating end is connected to the TEC cold surface of the heat exchange plate.
[0092] A micro temperature sensor (not shown) is also provided in the probe housing; the controller 30 is electrically connected to the micro temperature sensor and is further configured to:
[0093] Step S61, obtaining real-time temperature data of the micro temperature sensor;
[0094] Step S62: selectively adjust the power of the heat exchanger or switch the operating mode of the dual-spectrum endoscope probe according to the mapped range of the real-time temperature data.
[0095] The micro-temperature sensor uses an NTC thermistor to monitor the temperature of key areas of the probe 211 (such as the CMOS chip, lens assembly, and contact end face) in real time. The heat exchanger uses a semiconductor refrigeration plate to adjust the heat exchanger power or switch the operating mode of the dual-spectrum endoscope probe based on the current real-time temperature. Through real-time multi-level adjustment and mode switching, temperature stability ensures low CMOS noise, thereby dynamically balancing cooling and imaging performance.
[0096] Specifically, in step S62, the power of the heat exchanger is selectively adjusted or the operating mode of the dual-spectrum endoscope probe is switched according to the range mapped by the real-time temperature data, including:
[0097] Step S621: When the temperature data is between the preset first temperature threshold and the second temperature threshold, the power of the heat exchange level is controlled and adjusted according to the current temperature data pid;
[0098] Step S622: When the temperature data is greater than a preset second temperature threshold, the dual-spectrum endoscope probe is controlled to dynamically skip frames and reduce the frequency through exponential backoff.
[0099] Input real-time temperature data, preset PID parameters, and when the temperature data is between the preset first temperature threshold and the second temperature threshold, calculate the error between the real-time temperature and the target temperature, output and control the TEC cooling chip power, and its PWM duty cycle is set to 0 to 100%.
[0100] When the real-time temperature exceeds the preset second temperature threshold, exponential backoff frame skipping is selected and the output frequency is reduced; the dynamic frame rate is selected to be 15 / 22 / 30fps, etc., and the RGB frequency is downgraded; in extremely high temperatures, the visibility of core anatomical structures is prioritized to avoid system crashes and surgical interruptions, ensuring the effective operation of long-term operations.
[0101] like Figure 7 In a second aspect, the present application further provides a medical endoscopy system based on naked-eye 3D display, such as the above-mentioned medical endoscopy device 100 based on naked-eye 3D display. The medical endoscopy device 100 further includes a frame and an operating module corresponding to the dual-spectrum endoscope probe; and further includes a mobile terminal 200 for communicating with the controller of the medical endoscopy device to realize remote visual sharing.
[0102] Among them, the rack adopts a medical bracket with adjustable height / angle, and the operation module and controller are connected.
[0103] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A medical endoscopy device based on naked-eye 3D display, comprising a display module, characterized in that: Also includes: A dual-spectrum endoscope probe includes a probe housing, and a first sensor head, a second sensor head, an optical lens group, and a micro copper tube disposed in the probe housing, wherein the first sensor head and the second sensor head are separated by a baseline distance; wherein the first sensor head acquires an RGB image, and the second sensor head acquires a NIR image; The controller is electrically connected to the dual-spectrum endoscope probe and is configured to: Receive a dual-spectrum image of an RGB image and an NIR image, and align the timestamps of the dual-spectrum images and perform dynamic black level correction through an FPGA trigger signal; Smoke detection is performed on the bispectral image using the YOLOv5 network model. If smoke is present, a dehazing network model is constructed by automatically searching for and optimizing the U-Net model of the network structure, and the dehazing image and the derived image are output; According to the dehazed image or the original bispectral image, multimodal fusion is performed on the RGB image and the NIR image to obtain a fused image; The fused image is disparity layered and then reassembled and the disparity map is optimized to obtain the intermediate layer image; The parallax is dynamically adjusted based on the baseline distance and the TOF distance, and a dynamic 3D image is output through naked-eye 3D rendering.
2. The medical endoscopy device based on naked-eye 3D display according to claim 1, characterized in that: In the setting of the YOLOv5 network model, the first layer is configured as a standard convolutional layer, one Bottleneck is retained, and the residual connection is removed for structural optimization; in the loss function, a composite loss with at least three parameters is used to calculate the edge loss; wherein, the smoke detection of the bispectral image using the YOLOv5 network model includes: Input the image frames of RGB image and NIR image, and obtain the sampling feature map through comparative reasoning; Performing morphological separation after threshold segmentation on the sampled feature map to output a binary smoke mask; Get the confidence score of the binary smoke mask and perform a threshold judgment to output whether smoke exists or not.
3. The medical endoscopy device based on naked-eye 3D display according to claim 1, characterized in that: In the dehazing network model, the gradient-based DARTS algorithm searches in an image dataset of bispectral images to generate an optimal structure U-Net model containing an MBConv block; wherein the dehazing network model is constructed by automatically searching and optimizing the U-Net model of the network structure, and outputting a dehazing image and a derivative image, including: Input the binary smoke mask and the image frame of the RGB image; Depthwise separable convolutional feature extraction is performed through the optimal structure U-Net model, and attention-guided jump connections are used to output the dehazed image and the derived image.
4. The medical endoscopy device based on naked-eye 3D display according to claim 1, characterized in that: The multimodal fusion of the RGB image and the NIR image to obtain a fused image includes: Performing feature point matching on the RGB image and the NIR image with affine transformation; The defogging image and the registered NIR image are input, and a fusion weight is obtained based on a preset smoke attenuation coefficient to fuse the RGB image and the NIR image to obtain a fused image.
5. The medical endoscopy device based on naked-eye 3D display according to claim 1, characterized in that: The step of performing disparity layering on the fused image, recombining the image, and optimizing the disparity map to obtain an intermediate layer map includes: Input a stereo image pair of the fused image, divide the stereo image pair into equal parts along the channel dimension, wherein the first branch performs regular convolution, the second branch fuses the result of the first branch, and the third branch fuses the result of the second branch; The right disparity is reconstructed according to the left disparity of the stereo image pair, and an inconsistent area between the left disparity and the right disparity is calculated. The inconsistent area is filled with surrounding pixels to obtain an intermediate layer image.
6. The medical endoscopy device based on naked-eye 3D display according to claim 1, characterized in that: The dynamically adjusting the parallax based on the baseline distance and the TOF distance, and outputting a dynamic 3D image through naked-eye 3D rendering, includes: Perform disparity scaling and / or boundary cropping based on the intermediate layer image, TOF distance, and baseline distance to obtain a comfortable disparity map; Obtaining a light field image according to the fused image, the comfortable disparity map, preset grating parameters, and based on multi-viewpoint generation and sub-pixel rendering; And generate lens control signals through light field images to output dynamic 3D images.
7. The medical endoscopy device based on naked-eye 3D display according to any one of claims 1 to 6, characterized in that: The dual-spectrum endoscope probe further includes a heat exchange plate; The probe housing includes a probe part and an operating part. The first sensor head, the second sensor head, the optical lens group and the micro copper tube are all located inside the probe part. The heat exchange plate is located inside the operating part. One end of the micro copper tube extends into the probe part and contacts the first sensor head and the second sensor head; the other end contacts the heat exchange plate.
8. The medical endoscopy device based on naked-eye 3D display according to claim 7, characterized in that: A micro temperature sensor is further provided in the probe housing; the controller is electrically connected to the micro temperature sensor and is further configured to: Acquiring real-time temperature data of the micro temperature sensor; According to the mapping range of the real-time temperature data, the power of the heat exchange plate is selectively adjusted, or the operating mode of the dual-spectrum endoscope probe is switched.
9. The medical endoscopy device based on naked-eye 3D display according to claim 8, characterized in that: The selectively adjusting the power of the heat exchange plate or switching the operating mode of the dual-spectrum endoscope probe according to the mapping range of the real-time temperature data includes: When the temperature data is between a preset first temperature threshold and a preset second temperature threshold, the power of the heat exchange level is adjusted by controlling the current temperature data pid; When the temperature data is greater than a preset second temperature threshold, the dual-spectrum endoscope probe is controlled to dynamically skip frames and reduce frequency through exponential backoff.
10. A medical endoscopy system based on naked-eye 3D display, characterized in that: include: The medical endoscopy device based on naked-eye 3D display according to any one of claims 1 to 9, the medical endoscopy device comprising a frame and an operating module corresponding to the dual-spectrum endoscope probe; The mobile terminal is used to communicate with the controller of the medical endoscopy device.