Dynamic endoscope video three-dimensional reconstruction method and system based on blood vessel information enhancement

By performing image enhancement and multi-scale fast-guided filtering in the HSV color space, combined with 3DGS technology, the problem of missing vascular details in endoscopic three-dimensional reconstruction is solved, and the enhancement and visualization of vascular information is achieved, providing better visual support.

CN119963735AInactive Publication Date: 2025-05-09GUANGDONG MECHANICAL & ELECTRICAL COLLEGE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510043137.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing endoscopic three-dimensional reconstruction methods have problems in the rendering of vascular details, and the vascular enhancement algorithm lacks a unified and sound evaluation benchmark, resulting in poor three-dimensional reconstruction of endoscopic videos.

Method used

The image enhancement method based on HSV color space is adopted, combined with multi-scale fast-guided filtering and grayscale mapping function, and the endoscopic video is preprocessed, and three-dimensional reconstruction is used to enhance the visualization of vascular information.

Benefits of technology

It significantly improves the display effect of blood vessels in three-dimensional images, enhances the significance and visualization of blood vessel information, and provides more accurate visual support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963735A_ABST
    Figure CN119963735A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a dynamic endoscope video three-dimensional reconstruction method and system based on blood vessel information enhancement, and the method comprises the steps: collecting endoscope video images, and obtaining the depth information of each endoscope image in the endoscope video images; converting each endoscope image in the endoscope video image from an RGB domain to an HSV color space, and performing image enhancement on each endoscope image in the HSV color space to obtain an enhanced image; performing three-dimensional reconstruction rendering on the endoscope video images based on the depth information of each endoscope image in the endoscope video images and the corresponding enhanced image to obtain reconstructed endoscope video images; according to the invention, the highlighting effect of the blood vessel in the three-dimensional image can be realized, and the significance of blood vessel information in the image is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a method and system for dynamic endoscopic video three-dimensional reconstruction based on vascular information enhancement. Background Art

[0002] Reconstructing dynamic 3D scenes from endoscopic videos is crucial for robot-assisted minimally invasive surgery. By visualizing the 3D model of the tissue, doctors can simulate real surgical scenes in a virtual reality (VR) or augmented reality (AR) environment to achieve preoperative planning and surgical training. In addition, reconstruction technology that supports real-time rendering can also be directly applied to intraoperative scenes, allowing surgeons to fully understand the situation at the surgical site, thereby achieving precise navigation and control of surgical instruments and laying the foundation for the automation of robotic surgery.

[0003] Among the related technologies, the endoscopic 3DGS method has problems such as missing vascular details after three-dimensional reconstruction and rendering, and needs to be further optimized to meet the complex reconstruction requirements of surgical scenes.

[0004] There is a common problem of lack of fixed evaluation datasets in vascular enhancement algorithms. Most studies usually only select 5 to 6 public or private images for processing and evaluate relevant indicators based on them, which cannot achieve good generalization. This reflects the current lack of unified and sound evaluation benchmarks in the field. Summary of the invention

[0005] The object of the present invention is to provide a method and system for dynamic endoscopic video three-dimensional reconstruction based on vascular information enhancement, so as to achieve enhancement and visualization of vascular three-dimensional information.

[0006] In order to achieve the above object, the present invention provides the following technical solutions:

[0007] On the one hand, an embodiment of the present invention provides a method for dynamic endoscopic video 3D reconstruction based on vascular information enhancement, the method comprising the following steps:

[0008] S100, collecting endoscopic video images, and obtaining depth information of each endoscopic image in the endoscopic video images;

[0009] S200, converting each endoscopic image in the endoscopic video image from the RGB domain to the HSV color space, and performing image enhancement on each endoscopic image in the HSV color space to obtain an enhanced image;

[0010] S300, performing three-dimensional reconstruction and rendering on the endoscopic video image based on the depth information of each endoscopic image in the endoscopic video image and the corresponding enhanced image to obtain a reconstructed endoscopic video image.

[0011] Optionally, in S200, the step of performing image enhancement on each endoscopic image in the HSV color space to obtain an enhanced image includes:

[0012] S210, obtaining three channels of the endoscope image in the endoscope video image in the HSV color space; wherein the three channels include a hue channel, a saturation channel, and a lightness channel;

[0013] S220, adjusting the brightness of the endoscope image by dynamic range compression in the brightness channel of the three channels to obtain a filtered input image;

[0014] S230, using multi-scale fast guided filtering to perform brightness adaptive adjustment, dark area recovery and detail enhancement on the filtered input image to obtain a filtered image;

[0015] S240, performing contrast stretching on the region of interest in the filtered image in a saturation channel using a grayscale mapping function to obtain an enhanced image; wherein the region of interest includes blood vessel details in the filtered image;

[0016] Optionally, in S220, the brightness of the endoscopic image is adjusted by dynamic range compression in the brightness channel of the three channels to obtain a filtered input image, including:

[0017] The nonlinear transfer function is used to adjust the brightness of the endoscopic image. The calculation process of the nonlinear transfer function is as follows:

[0018]

[0019] Wherein, Vn represents the filtered input image, V represents the brightness channel image of the endoscopic image, and z represents the adaptive adjustment factor.

[0020] The parameter z can realize the adaptive adjustment of brightness, which is related to the histogram of the brightness channel as follows:

[0021]

[0022] Where L represents the pixel intensity value.

[0023] Optionally, in S230, the multi-scale fast guided filtering is used to perform brightness adaptive adjustment, dark area recovery and detail enhancement on the filtered input image to obtain the filtered image, including:

[0024] S231, establishing a local linear relationship between the guide image and the filtered output image under windows of multiple different sizes;

[0025] S232, for each size of the window, minimize the reconstruction error between the filter input image and the filter output image based on the local linear relationship to obtain filter parameters, and obtain a filter output image corresponding to the filter input image based on the filter parameters;

[0026] S233, adjusting the intensity of the pixels in the filtered output image based on the intensity of the corresponding pixels in the filtered input image, to obtain a filtered adjustment image under the window size;

[0027] S234, taking the average of the filtered and adjusted images under each window size to obtain a filtered image.

[0028] Optionally, in S232, the step of minimizing a reconstruction error between a filter input image and a filter output image based on the local linear relationship to obtain a filter parameter, and obtaining a filter output image corresponding to the filter input image based on the filter parameter includes:

[0029] Given a filtered input image V, minimizing the reconstruction error between the filtered input image V and the filtered output image Q yields:

[0030]

[0031] bk=Vk-akμk (5);

[0032] Among them, a k , b k is the filter parameter of the kth window, i is the index of the pixel, I i is the i-th pixel in the guidance image, V i is the i-th pixel in the filtered input image, V k is the image area of ​​the kth window in the filtered input image, ω represents the local square window with radius r, ω k is a local square window with the kth window as the center and r as the radius, is the filtered input image V in window ω k The mean value within k and σ k are the mean and standard deviation of the guide image I in window k, respectively, and ∈ is the regularization parameter that controls the smoothness;

[0033] The filtered output image is calculated by the following formula:

[0034]

[0035] in, and are the window ω centered on the i-th pixel respectively. i Upper filter parameter a iand b i The average value of .

[0036] Optionally, S234, adjusting the intensity of pixels in the filtered output image based on the intensity of corresponding pixels in the filtered output image and the filtered input image to obtain a filtered adjustment image under a window of the size includes:

[0037] Compare the intensity of the corresponding pixels in the filtered output image and the filtered input image. If the intensity of the central pixel is higher than the average intensity of the surrounding pixels, increase the intensity of the corresponding pixel in the filtered output image; otherwise, reduce the intensity of the corresponding pixel to obtain a filtered adjusted image.

[0038] The process is carried out according to the following formula:

[0039]

[0040] Where (x, y) represents the coordinates of the pixel, j represents the index of the number of times the intensity of the pixel in the filtered output image is adjusted, and v j (x, y) represents the filtered and adjusted image output for the jth time, E j (x,y),Q j (x, y), V(x, y) represent the j-th adjustment index, the j-th output filtered image, and the filtered input image, respectively. V n (x, y) represents the brightness channel image of the processed endoscopic image, and P is the adaptive adjustment parameter;

[0041]

[0042] Where σ is the global standard deviation of the filtered output image.

[0043] Optionally, in S240, the mapping function for contrast stretching in the saturation channel is:

[0044] S*=[1+ds×(S-1-1)2]-1 (11);

[0045] Among them, S and S* are the grayscale values ​​of the saturation channel before and after mapping, respectively; d s is the grayscale mapping parameter of the saturation channel. The calculation formula of the grayscale mapping parameter ds is:

[0046]

[0047] Among them, S ave is the grayscale mean of the saturation channel before mapping.

[0048] Optionally, in S300, the three-dimensional reconstruction and rendering of the endoscopic video image based on the depth information of each endoscopic image in the endoscopic video image and the corresponding enhanced image to obtain a reconstructed endoscopic video image includes:

[0049] S310, modeling the surgical scene in the form of a 3D Gaussian point cloud in a world coordinate system, using a SfM algorithm to estimate and initialize the 3D point cloud from a set of enhanced images to obtain a set of Gaussian distributions;

[0050] S320, for each pixel, calculating the colors and opacities of N Gaussian distributions overlapping with the pixel, and mixing the N Gaussian distributions overlapping with the pixel to obtain a final color of the pixel;

[0051] S330, reconstructing the endoscopic image based on the final color of each pixel in the endoscopic image, and obtaining a three-dimensionally reconstructed and rendered endoscopic video image based on the reconstructed endoscopic image.

[0052] On the other hand, an embodiment of the present invention provides a dynamic endoscopic video 3D reconstruction system based on vascular information enhancement, comprising:

[0053] at least one processor;

[0054] at least one memory for storing at least one program;

[0055] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.

[0056] The beneficial effects of the present invention are as follows: the present invention discloses a dynamic endoscopic video 3D reconstruction method and system based on vascular information enhancement, which pre-processes endoscopic video images by combining an image enhancement method, and uses 3DGS technology to complete the 3D reconstruction of tissues, thereby achieving enhancement and visualization of vascular 3D information. The present invention can highlight the display effect of blood vessels in 3D images and enhance the significance of vascular information in images. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0058] Figure 1 It is a flow chart of a method for dynamic endoscopic video three-dimensional reconstruction based on vascular information enhancement according to an embodiment of the present invention;

[0059] Figure 2 It is a framework diagram of a dynamic endoscopic video three-dimensional reconstruction method based on vascular information enhancement according to an embodiment of the present invention;

[0060] Figure 3 is a comparison chart of rendering results in two different data sets in combination with the EndoGaussian method in an embodiment of the present invention;

[0061] Figure 4 It is the visualization and interactive effect diagram of the three-dimensional model in the MR device;

[0062] Figure 5 It is a structural schematic diagram of a dynamic endoscopic video three-dimensional reconstruction system based on vascular information enhancement according to an embodiment of the present invention. DETAILED DESCRIPTION

[0063] The following will be combined with the embodiments and drawings to clearly and completely describe the concept, specific structure and technical effects disclosed in the present invention, so as to fully understand the purpose, scheme and effect disclosed in the present invention. It should be noted that the embodiments and features in the embodiments in this application can be combined with each other without conflict.

[0064] The following first explains the technical terms involved in the embodiments of the present invention:

[0065] 3D Reconstruction: It refers to the establishment of a mathematical model suitable for computer representation and processing of three-dimensional objects. It is the basis for processing, operating and analyzing its properties in a computer environment. It is also a key technology for establishing virtual reality in computers to express the objective world.

[0066] Virtual reality technology (VR) is a computer simulation system that can create and experience a virtual world. It uses computers to generate a simulated environment and immerse users in it.

[0067] Augmented Reality (AR) is a technology that superimposes and integrates virtual scenes or information with the real physical environment, and interactively presents them to users, thus creating a space where virtual and real objects share the same space.

[0068] 3D reconstruction rendering method: In the field of surgical scene reconstruction, early SLAM-based methods estimated depth through stereo video and fused depth maps in 3D space for reconstruction. However, due to ignoring the influence of surgical tools and assuming that the scene is static, its reconstruction effect in actual surgical scenes is limited. Subsequent improved methods introduced a tool-masked depth estimation framework to better deal with the interference of surgical tools; SurfelWarp proposed a single-view 3D deformation reconstruction method for capturing deformable objects. However, these methods mainly rely on sparse distortion fields for deformation tracking, and there are still obvious limitations in the processing of complex dynamic scenes. In recent years, the implicit reconstruction method of the NeRF model has performed well in 3D reconstruction. EndoNeRF and EndoSurf further applied it to medical scene reconstruction to achieve high-fidelity 3D reconstruction of endoscopic images. The emergence of new methods based on 3DGS has greatly shortened the rendering time. It surpasses NeRF in rendering speed and quality, especially Defor m3DGS can complete scene training in just 1 minute. However, the existing endoscopic 3DGS method has problems such as missing vascular details after 3D reconstruction rendering, and needs to be further optimized to meet the complex reconstruction requirements of surgical scenes.

[0069] Vascular enhancement algorithm: In the field of vascular enhancement of medical images, existing algorithm research mainly focuses on retinal images and endoscopic video images. Unlike the segmentation and enhancement of static vascular structures in retinal images, endoscopic videos have real-time and dynamic characteristics. In addition, due to the complexity of the internal environment of the human body, including the diversity of tissue characteristics, changes in lighting conditions, and the interference of motion artifacts, the vascular enhancement process faces greater challenges, which puts higher requirements on the research of more targeted and adaptive algorithms. Existing endoscopic vascular enhancement research mainly focuses on the two-dimensional processing of endoscopic video frames, using two types of methods: learning-based and traditional. A key challenge facing learning-based methods is the acquisition of high-quality labels. Although the emergence of generative methods has provided new possibilities for solving this problem, its application scope is still limited due to the lack of clear and public endoscopic image datasets. In addition, traditional methods dominate current research because they are more suitable for the needs of real-time endoscopy processing. For example, the histogram equalization method improves contrast by redistributing image grayscale values; the filtering-based method uses enhancement techniques of different frequencies and directions to highlight vascular details; the Retinex-based method significantly improves the overall quality of the image by improving uneven illumination and enhancing detail performance. However, existing vascular enhancement algorithms generally lack a fixed evaluation dataset, and most studies usually only select 5-6 public or private images for processing and evaluate relevant indicators based on them. This reflects the current lack of a unified and sound evaluation benchmark in the field.

[0070] refer to Figure 1 and Figure 2 ,like Figure 1 A method for dynamic endoscopic video 3D reconstruction based on vascular information enhancement provided by an embodiment of the present invention is shown, and the method comprises the following steps:

[0071] S100, collecting endoscopic video images, and obtaining depth information of each endoscopic image in the endoscopic video images;

[0072] S200, converting each endoscopic image in the endoscopic video image from the RGB domain to the HSV color space, and performing image enhancement on each endoscopic image in the HSV color space to obtain an enhanced image;

[0073] Specifically, the endoscopic image is first converted from the RGB domain to the HSV color space, and the hue, saturation and brightness are extracted, and the obtained three channels include a hue channel, a saturation channel and a brightness channel.

[0074] S300, performing three-dimensional reconstruction and rendering on the endoscopic video image based on the depth information of each endoscopic image in the endoscopic video image and the corresponding enhanced image to obtain a reconstructed endoscopic video image.

[0075] Specifically, the depth information of each enhanced image and the corresponding endoscopic image is input into the endoscopic 3DGS algorithm to achieve three-dimensional reconstruction and rendering of the endoscopic video image.

[0076] The present invention proposes an endoscopic video three-dimensional reconstruction method that effectively enhances vascular information. The endoscopic video image is preprocessed by combining an image enhancement method, and the 3DGS technology is used to complete the three-dimensional reconstruction of the tissue, thereby realizing the enhancement and visualization of the three-dimensional information of the blood vessels. The blood vessels are highlighted in the three-dimensional image, and the significance of the vascular information in the image is enhanced.

[0077] As an improvement of the above embodiment, in S200, the image enhancement is performed on each endoscopic image in the HSV color space to obtain an enhanced image, including:

[0078] S210, obtaining three channels of the endoscope image in the endoscope video image in the HSV color space; wherein the three channels include a hue channel, a saturation channel, and a lightness channel;

[0079] S220, adjusting the brightness of the endoscope image by dynamic range compression in the brightness channel of the three channels to obtain a filtered input image;

[0080] S230, using multi-scale fast guided filtering to perform brightness adaptive adjustment, dark area recovery and detail enhancement on the filtered input image to obtain a filtered image;

[0081] S240, performing contrast stretching on the region of interest in the filtered image in a saturation channel using a grayscale mapping function to obtain an enhanced image; wherein the region of interest includes blood vessel details in the filtered image;

[0082] It should be noted that the saturation channel represents the purity or intensity of the color and determines the vividness of the color. The saturation channel can independently adjust the vividness of the color without affecting the brightness or hue. Through the mapping algorithm, the contrast and clarity of the blood vessels can be improved, so that the location and morphology of the blood vessels can be more effectively displayed. In addition, the color saturation of the blood vessel area is enhanced to make it more vivid and bright in the image, which is convenient for doctors to identify and analyze, providing accurate visual support for doctors.

[0083] In some improved embodiments, in S220, the brightness of the endoscopic image is adjusted by dynamic range compression in the brightness channel of the three channels to obtain a filtered input image, including:

[0084] The nonlinear transfer function is used to adjust the brightness of the endoscopic image. The calculation process of the nonlinear transfer function is as follows:

[0085]

[0086] Wherein, Vn represents the filtered input image, V represents the brightness channel image of the endoscopic image, and z represents the adaptive adjustment factor.

[0087] The parameter z can realize the adaptive adjustment of brightness, which is related to the histogram of the brightness channel as follows:

[0088]

[0089] Where L represents the pixel intensity value.

[0090] Specifically, L is a pixel intensity value of a gray level determined according to the Cumulative Distribution Function (CDF) of the image. In this embodiment, it is a pixel intensity value that makes the cumulative distribution function reach 0.1.

[0091] It should be noted that if L≤50, it indicates that the endoscopic image is dark and needs to have its brightness enhanced; if 50<L<150, it indicates that the endoscopic image has moderate brightness and needs less brightness enhancement; if L≥150, it indicates that the endoscopic image has sufficient brightness and does not need to be enhanced.

[0092] In some improved embodiments, in S230, the multi-scale fast guided filtering is used to perform brightness adaptive adjustment, dark area recovery and detail enhancement on the filtered input image to obtain the filtered image, including:

[0093] S231, establishing a local linear relationship between the guide image and the filtered output image under windows of multiple different sizes;

[0094] S232, for each size of the window, minimize the reconstruction error between the filter input image and the filter output image based on the local linear relationship to obtain filter parameters, and obtain a filter output image corresponding to the filter input image based on the filter parameters;

[0095] S233, adjusting the intensity of the pixels in the filtered output image based on the intensity of the corresponding pixels in the filtered input image, to obtain a filtered adjustment image under the window size;

[0096] S234, taking the average of the filtered and adjusted images under each window size to obtain a filtered image.

[0097] It should be noted that although overall brightness enhancement helps to show the information of dark areas in endoscopic images, it may also lead to the emergence of new noise information. Therefore, we designed a multi-scale fast guided filter, which can not only remove noise, but also well preserve edges and shorten calculation time. In addition, the edge information in the brightness channel also contains vascular information, and the enhancement algorithm will not change the color information of the original image, which can better preserve the information.

[0098] Specifically, the local linear relationship between the guide image I and the filtered output image Q is set as follows:

[0099]

[0100] Where i is the index of the pixel, k is the index of the window ω with radius r, and the window ω is a local square window. k is a local square window with window k as the center and r as the radius, Q i is the i-th pixel in the filtered output image, I i is the i-th pixel in the guidance image.

[0101] Given a filtered input image V, minimizing the reconstruction error between the filtered input image V and the filtered output image Q yields:

[0102]

[0103] bk=Vk-akμk (5);

[0104] Among them, a k 、b k is the filter parameter of the kth window, V i is the i-th pixel in the filtered input image, V kis the image area of ​​the kth window in the filtered input image, is the filtered input image V in window ω k The mean value within k and σ k are the mean and standard deviation of the guidance image I in window k, respectively, and ∈ is a regularization parameter that controls the smoothness.

[0105] In the calculation image, all windows ω k The filter parameter a k and b k After that, the filtered output image output by the filter is:

[0106]

[0107] in, and are the window ω centered on the i-th pixel respectively. i Upper filter parameter a i and b i The average value of .

[0108] Fast guided filtering reduces the computational complexity by downsampling the filter input image V and the guidance image I by a ratio. k and b k After the calculation is completed, the original resolution is restored by upsampling, which effectively reduces the calculation time while maintaining the spatial consistency and accuracy of the processing results.

[0109] After fast guided filtering, the intensity of the corresponding pixels in the filtered output image is compared with that in the filtered input image. If the intensity of the central pixel is higher than the average intensity of the surrounding pixels, the intensity of the corresponding pixel in the filtered output image is increased; otherwise, the intensity of the corresponding pixel is reduced.

[0110] The process is carried out according to the following formula:

[0111]

[0112] Where (x, y) represents the coordinates of the pixel, j represents the index of the number of times the intensity of the pixel in the filtered output image is adjusted, and v j (x, y) represents the filtered and adjusted image output for the jth time, E j (x,y),Q j (x, y), V(x, y) represent the j-th adjustment index, the j-th output filtered image, and the filtered input image, respectively. V n (x, y) represents the brightness channel image of the processed endoscopic image, and P is the adaptive adjustment parameter;

[0113] The parameter P enables adaptive adjustment of the enhancement process and is related to the global standard deviation σ of the filtered output image as follows:

[0114]

[0115] Where σ is the global standard deviation of the filtered output image.

[0116] If E j If (x,y)<1, it means that the central pixel is brighter than the surrounding pixels; otherwise, it means that the central pixel is darker than the surrounding pixels. In this way, the problem of image detail information changing or degrading after the brightness is increased can be effectively avoided.

[0117] In fast guided filtering, the size of the window ω directly affects the filtering effect: a smaller window provides brightness information in the local neighborhood, while a larger window close to the image size reflects the changing characteristics of the global brightness. To achieve better image enhancement, we use three different sizes of windows (4, 16, 64) to filter the input image multiple times and enhance the contrast. Finally, the three filtering results are averaged to obtain the filtered image.

[0118]

[0119] Among them, V * (x,y) represents the filtered image.

[0120] The mapping function for contrast stretching in the saturation channel is:

[0121] S*=[1+ds×(S-1-1)2]-1 (11);

[0122] Among them, S and S* are the grayscale values ​​of the saturation channel before and after mapping, respectively; d s is the grayscale mapping parameter of the saturation channel. The calculation formula of the grayscale mapping parameter ds is:

[0123]

[0124] Among them, S ave is the grayscale mean of the saturation channel before mapping;

[0125] The mapping function can adaptively adjust the distribution characteristics of the saturation component S, thereby stretching the saturation contrast of the region of interest (blood vessel details) in the filtered image.

[0126] In some improved embodiments, in S300, the three-dimensional reconstruction and rendering of the endoscopic video image based on the depth information of each endoscopic image in the endoscopic video image and the corresponding enhanced image to obtain the reconstructed endoscopic video image includes:

[0127] S310, modeling the surgical scene in the form of a 3D Gaussian point cloud in a world coordinate system, using a SfM algorithm to estimate and initialize the 3D point cloud from a set of enhanced images to obtain a set of Gaussian distributions;

[0128] It should be noted that 3DGS (3D Gaussian Splatting) is a 3D scene representation and rendering technology. 3DGS is an explicit 3D scene representation in the form of point cloud. It is a static 3D scene representation that models the scene in the form of 3D Gaussian point cloud in the world coordinate system. In the reconstruction part, the SfM algorithm is used to estimate the 3D point cloud from a set of enhanced images and initialize it to obtain a set of 3D Gaussian functions. 3DGS uses 3D Gaussian functions to represent 3D scenes and achieves high-quality scene reconstruction and rendering by optimizing the parameters of the Gaussian function.

[0129] S320, for each pixel, calculating the colors and opacities of N Gaussian distributions overlapping with the pixel, and mixing the N Gaussian distributions overlapping with the pixel to obtain a final color of the pixel;

[0130] It should be noted that in the rendering part, each 3D Gaussian function is represented by a covariance matrix Σ and a center point X, which is called the mean value of the 3D Gaussian function. The 3D Gaussian function is expressed as:

[0131]

[0132] Among them, G(X) represents the value of the 3D Gaussian function, which represents the function response at point X. X is the coordinate of a point in 3D space, represented as a column vector.

[0133] For differentiable optimization, the covariance matrix Σ can be decomposed into a scaling matrix S and a rotation matrix R:

[0134] Σ = RSSTRT (14);

[0135] Gaussian Splatting is a technique for 3D Gaussian functions in the camera plane, which is mainly used for the representation and rendering of 3D scenes. By using the observation transformation matrix W and the Jacobian matrix J of the affine approximation of the projective transformation, the covariance matrix Σ′ in the camera coordinates can be calculated as:

[0136] Σ′=JWΣWTJT (15);

[0137] Each 3D Gaussian function contains learnable properties including: position (position X∈R 3 ), colors defined by spherical harmonic coefficients (spherical harmonic coefficients Y∈R m, m is the number of spherical harmonics), opacity α, rotation angle (rotation angle r∈R 4 ) and scale (scale s∈R 3 ).

[0138] Specifically, for each pixel, the color and opacity of all Gaussian distributions are calculated using Equation (13). Each Gaussian distribution has its corresponding color and opacity, and these Gaussian distributions are mixed according to the depth order (from front to back) to finally get the color of the pixel.

[0139] The final color of the pixel is obtained by mixing N Gaussian distributions that overlap with the pixel. The calculation formula is:

[0140]

[0141] Where C is the final color of the pixel, which is the result of cumulative mixing of all Gaussian distribution colors overlapping the pixel in order of opacity weight and depth, N represents the number of Gaussian distributions overlapping the current pixel, ci and αi represent the color value and opacity of the i-th Gaussian distribution, respectively. The Gaussian distribution is calculated by multiplying the 3D Gaussian G with covariance Σ by the opacity of each Gaussian distribution that can be optimized and the color defined by the spherical harmonic coefficients.

[0142] S330, reconstructing the endoscopic image based on the final color of each pixel in the endoscopic image, and obtaining a three-dimensionally reconstructed and rendered endoscopic video image based on the reconstructed endoscopic image.

[0143] In some embodiments, after S330, the method further includes:

[0144] The optimization strategy of 3DGS is combined to improve the effect of endoscopic video image reconstruction.

[0145] Specifically, the reconstruction effect is improved by combining the optimization strategies of Endo Gaussian or Deform3DGS. Both frameworks are designed to improve the efficiency and quality of surgical scene reconstruction. Endo Gaussian focuses on real-time rendering and processing of dynamic endoscopic scenes, while Deform3DGS focuses on fast reconstruction and deformation modeling.

[0146] It should be noted that Endo Gaussian is a real-time endoscopic scene reconstruction framework based on 3D Gaussian Splatting (3DGS). The framework significantly improves the rendering speed to real-time level by integrating efficient Gaussian representation and highly optimized rendering engine. To adapt to endoscopic scenes, Endo Gaussian proposes two strategies: Holistic Gaussian Initialization (HGI) and Spatio-temporal Gaussian Tracking (SGT), which deal with non-trivial Gaussian initialization and tissue deformation problems respectively. In HGI, the latest depth estimation model is used to predict the depth map of the input binocular / monocular image sequence, based on which the pixels are reprojected and combined for holistic initialization. In SGT, a deformation field is proposed to simulate surface dynamics, which consists of efficient encoded voxels and a lightweight deformation decoder, allowing Gaussian tracking to be performed with less training and rendering burden. Experiments show that Endo Gaussian outperforms previous techniques in multiple aspects, including better rendering speed (195FPS real-time, 100 times improvement), better rendering quality (37.848PSNR), and less training overhead (within 2 minutes per scene).

[0147] Deform3DGS is a framework for fast surgical scene reconstruction, which is also based on Gaussian Splatting. Deform3DGS introduces 3DGS into surgical scenes through point cloud initialization, and proposes a novel Flexible Deformation Modeling (FDM) scheme to learn the tissue deformation dynamics of a single Gaussian point. FDM can model surface deformation using efficient representations, thereby achieving real-time rendering performance. More importantly, FDM significantly accelerates surgical scene reconstruction and demonstrates great clinical value, especially in intraoperative settings where time efficiency is critical. Experiments show that the application of Deform3DGS on da Vinci robotic surgery videos is effective, demonstrating excellent reconstruction fidelity (PSNR: 37.90) and rendering speed (338.8FPS), while significantly reducing the training time to only 1 minute per intraoperative scene.

[0148] After the endoscopic video image rendered by three-dimensional reconstruction is used as a three-dimensional reconstruction model, the method further comprises:

[0149] Design a Unity application for PC and import SDKs (Software Development Kits) such as MRTK3, OpenXR, and PXR.

[0150] By calling Holographic Remote Connect.cs, the MR device can realize real-time streaming between the desktop program and the holographic remote player in HoloLens2, placing the rendering of the 3D reconstructed model on the rendering workstation, and the display and interaction on the MR device.

[0151] The 3D reconstructed model under the path is read through the script. The model is bound to a parent object of a Game Object. Scripts such as XR Interaction Manager.cs, Constraint Manager.cs, and Object Manipulator.cs are bound to the Game Object.

[0152] In the application, hand tracking can be realized, and the three-dimensional reconstructed model can be grabbed by hand rays to perform functions such as movement and rotation.

[0153] refer to Figure 3 and Figure 4 ,Compared with the EndoGaussian and Deform3DGS methods, the results show that the ,method of the present invention has a significant improvement in the core DV / BV index (a higher DV / BV ratio indicates that ,the blood vessels are more clearly distinguished from the background in the ,image). Figure 3 It can be observed that when processing is performed only in the brightness channel (0ds), the method of the present invention can reveal the vascular information that was originally not displayed in the dark area of ​​the image, and the noise in the dark area can also be smoothed to a certain extent. After adding saturation channel processing (0.5ds / 1ds / 2ds), the blood vessels in the edge area of ​​the image are better preserved and displayed. In addition, after data set enhancement, the vascular details become more obvious in the entire image and are presented more intuitively.

[0154] In the embodiment provided by the present invention, the endoscopic video image is preprocessed by combining an image enhancement method, and the three-dimensional reconstruction of the tissue is completed using 3DGS technology, thereby achieving enhancement and visualization of the three-dimensional information of the blood vessels.

[0155] The designed image enhancement method is lightweight and efficient, and adaptively adjusts the image based on the HSV color space. At the same time, the SCARED and ENDONERF 3D reconstruction datasets are introduced, and DV / BV is used as the core indicator to evaluate the reconstructed and rendered 2D images, which can better improve the evaluation benchmark.

[0156] refer to Figure 5 The embodiment of the present invention further provides a dynamic endoscopic video 3D reconstruction system based on vascular information enhancement, comprising:

[0157] at least one processor;

[0158] at least one memory for storing at least one program;

[0159] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.

[0160] The contents of the above method embodiments are all applicable to this embodiment. The functions specifically implemented by this embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments, which will not be repeated here.

[0161] Although the description of the present disclosure has been quite detailed and specifically describes several described embodiments, it is not intended to be limited to any of these details or embodiments or any particular embodiment, but should be regarded as providing a broad possible interpretation of these claims by reference to the appended claims, taking into account the prior art, so as to effectively cover the intended scope of the present disclosure. In addition, the above description of the present disclosure is based on the embodiments foreseeable by the inventor, and its purpose is to provide a useful description, and those non-substantial changes to the present disclosure that have not yet been foreseen may still represent equivalent changes to the present disclosure.

Claims

1. A dynamic endoscopic video 3D reconstruction method based on vascular information enhancement, characterized in that: The method comprises the following steps: S100, collecting endoscopic video images, and obtaining depth information of each endoscopic image in the endoscopic video images; S200, converting each endoscopic image in the endoscopic video image from the RGB domain to the HSV color space, and performing image enhancement on each endoscopic image in the HSV color space to obtain an enhanced image; S300, performing three-dimensional reconstruction and rendering on the endoscopic video image based on the depth information of each endoscopic image in the endoscopic video image and the corresponding enhanced image to obtain a reconstructed endoscopic video image.

2. The method according to claim 1, characterized in that In S200, the image enhancement is performed on each endoscopic image in the HSV color space to obtain an enhanced image, including: S210, obtaining three channels of the endoscope image in the endoscope video image in the HSV color space; wherein the three channels include a hue channel, a saturation channel, and a lightness channel; S220, adjusting the brightness of the endoscope image by dynamic range compression in the brightness channel of the three channels to obtain a filtered input image; S230, using multi-scale fast guided filtering to perform brightness adaptive adjustment, dark area recovery and detail enhancement on the filtered input image to obtain a filtered image; S240, using a grayscale mapping function to perform contrast stretching on the region of interest in the filtered image in a saturation channel to obtain an enhanced image; wherein the region of interest includes blood vessel details in the filtered image.

3. The method according to claim 2, characterized in that In S220, the brightness of the endoscopic image is adjusted by dynamic range compression in the brightness channel of the three channels to obtain a filtered input image, including: The nonlinear transfer function is used to adjust the brightness of the endoscopic image. The calculation process of the nonlinear transfer function is as follows: Wherein, Vn represents the filtered input image, V represents the brightness channel image of the endoscopic image, and z represents the adaptive adjustment factor. The parameter z can realize the adaptive adjustment of brightness, which is related to the histogram of the brightness channel as follows: Where L represents the pixel intensity value.

4. The method according to claim 3, characterized in that In S230, the multi-scale fast guided filtering is used to perform brightness adaptive adjustment, dark area recovery and detail enhancement on the filtered input image to obtain a filtered image including: S231, establishing a local linear relationship between the guide image and the filtered output image under windows of multiple different sizes; S232, for each size of the window, minimize the reconstruction error between the filter input image and the filter output image based on the local linear relationship to obtain filter parameters, and obtain a filter output image corresponding to the filter input image based on the filter parameters; S233, adjusting the intensity of the pixels in the filtered output image based on the intensity of the corresponding pixels in the filtered input image, to obtain a filtered adjustment image under the window size; S234, taking the average of the filtered and adjusted images under each window size to obtain a filtered image.

5. The method according to claim 4, characterized in that In S232, the step of minimizing the reconstruction error between the filter input image and the filter output image based on the local linear relationship to obtain filter parameters, and obtaining the filter output image corresponding to the filter input image based on the filter parameters includes: Given a filtered input image V, minimizing the reconstruction error between the filtered input image V and the filtered output image Q yields: bk=Vk-akμk (5); Among them, a k 、b k is the filter parameter of the kth window, i is the index of the pixel, I i is the i-th pixel in the guidance image, V i is the i-th pixel in the filtered input image, V k is the image area of ​​the kth window in the filtered input image, ω represents the local square window with radius r, ω k is a local square window with the kth window as the center and r as the radius, is the filtered input image V in window ω k The mean value within k and σ k are the mean and standard deviation of the guide image I in window k, respectively, and ∈ is the regularization parameter that controls the smoothness; The filtered output image is calculated by the following formula: in, and are the window ω centered on the i-th pixel respectively. i Upper filter parameter a i and b i The average value of .

6. The method according to claim 5, characterized in that S234, adjusting the intensity of the pixels in the filtered output image based on the intensity of the corresponding pixels in the filtered output image and the filtered input image to obtain a filtered adjustment image under the window of the size, including: Compare the intensity of the corresponding pixels in the filtered output image and the filtered input image. If the intensity of the central pixel is higher than the average intensity of the surrounding pixels, increase the intensity of the corresponding pixel in the filtered output image; otherwise, reduce the intensity of the corresponding pixel to obtain a filtered adjusted image. The process is carried out according to the following formula: Where (x, y) represents the coordinates of the pixel, j represents the index of the number of times the intensity of the pixel in the filtered output image is adjusted, and v j (x, y) represents the filtered and adjusted image output for the jth time, E j (x,y),Q j (x, y), V(x, y) represent the j-th adjustment index, the j-th output filtered image, and the filtered input image, respectively. V n (x, y) represents the brightness channel image of the processed endoscopic image, and P is the adaptive adjustment parameter; Where σ is the global standard deviation of the filtered output image.

7. The method according to claim 6, characterized in that In S240, the mapping function for contrast stretching in the saturation channel is: S*=[1+ds×(S-1-1)2]-1 (11); Among them, S and S* are the grayscale values ​​of the saturation channel before and after mapping, respectively; d s is the grayscale mapping parameter of the saturation channel. The calculation formula of the grayscale mapping parameter ds is: Among them, S ave is the grayscale mean of the saturation channel before mapping.

8. The method according to claim 1, characterized in that In S300, the three-dimensional reconstruction and rendering of the endoscopic video image based on the depth information of each endoscopic image in the endoscopic video image and the corresponding enhanced image to obtain the reconstructed endoscopic video image includes: S310, modeling the surgical scene in the form of a 3D Gaussian point cloud in a world coordinate system, using a SfM algorithm to estimate and initialize the 3D point cloud from a set of enhanced images to obtain a set of Gaussian distributions; S320, for each pixel, calculating the colors and opacities of N Gaussian distributions overlapping with the pixel, and mixing the N Gaussian distributions overlapping with the pixel to obtain a final color of the pixel; S330, reconstructing the endoscopic image based on the final color of each pixel in the endoscopic image, and obtaining a three-dimensionally reconstructed and rendered endoscopic video image based on the reconstructed endoscopic image.

9. A dynamic endoscopic video 3D reconstruction system based on vascular information enhancement, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Endoscope image enhancement method based on histogram equalization and improved unsharpened mask

    CN113989147A

  • Endoscope image enhancement method based on improved Retinex and weighted guided filtering

    CN115526799A

  • Blood vessel enhancement method for endoscope image

    CN117218036A

  • Capsule endoscope image three-dimensional reconstruction method, electronic device, and readable storage medium

    US20240054662A1

  • Three-dimensional reconstruction method and apparatus for monocular endoscope image, and terminal device

    WO2021115071A1