Endoscope image depth estimation method and apparatus
By collecting endoscopic image data, using nonlinear filtering and near-beam radiation model, combined with multi-scale decomposition, the problems of blurred boundary information and missing visual far point depth in endoscopic image depth estimation are solved, more accurate depth estimation is achieved, and a new feature extraction method is provided for lesion detection and three-dimensional reconstruction.
Patent Information
- Application Number
- CN202310036781.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-10
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2043-01-10
AI Technical Summary
The existing endoscopic image depth estimation methods have problems such as blurred boundary information, missing visual far point depth, and unclear depth change levels.
The near-light radiation model and light source distance attenuation are used to complete the image of the area and obtain the initial depth estimation image. The method of completing the image in the highlight area through nonlinear filtering and significant pixel distribution to obtain the initial depth estimation map solves the problem of image completion in the area. Through the image completion module, multi-scale residuals and color space intensity are integrated to obtain the depth estimation information of the endoscopic image.
It effectively improves the expression of coherent boundary information in the initial depth map, optimizes the depth estimation of endoscopic images, and provides new ideas for feature extraction of lesion detection and three-dimensional reconstruction.
Smart Images

Figure CN116109687B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and in particular to an endoscopic image depth estimation method and an endoscopic image depth estimation device. Background Art
[0002] In recent years, with the expansion of computer vision research, depth estimation in endoscopic images has found widespread application in lesion detection, surgical instrument tracking, and 3D reconstruction. It provides feedback on the internal environment of the endoscope, thereby ensuring successful surgery and assisting doctors. It has also become an effective way to enhance the value of computer vision. Therefore, depth estimation technology for endoscopic images has become a key technology for capturing the surface morphology of objects. However, due to the complexity of the operating environment during clinical surgery, the limitations of the surgical operating space, and the differences in endoscopic equipment, the quality of the acquired endoscopic images varies greatly.
[0003] Minimally invasive surgery relies entirely on the endoscope's built-in light source for illumination. Due to size limitations, clinical endoscopes can only use monocular endoscopes. The camera and light source are close to the surface of the object, and the incident direction of the illumination cannot be considered parallel for different points on the surface of the object. In clinical surgery, the constant movement of the cold light source and camera cannot meet the requirements for the constancy of the light source during the operation. Currently, the widely used image depth estimation methods have blurred boundary information, missing the depth of the visual far point, and the depth change level is not obvious. These limiting factors will seriously affect subsequent visual applications. Summary of the Invention
[0004] In order to overcome the defects of the existing technology, the technical problem to be solved by the present invention is to provide an endoscopic image depth estimation method, which can effectively estimate the initial depth map from the original endoscopic image, effectively improve the expression of coherent boundary information of the initial depth map, further optimize the depth estimation of the endoscopic depth estimation, and provide a new idea for feature extraction for applications such as endoscopic image depth estimation, lesion detection, and three-dimensional reconstruction, solving the problems of blurred boundary information, missing visual far point depth, and unclear depth change levels in the current endoscopic image depth estimation.
[0005] The technical solution of the present invention is: this endoscopic image depth estimation method comprises the following steps:
[0006] (1) Collect clinical surgical endoscopic image data and obtain highlight areas based on nonlinear filtering and significant pixel distribution. The highlight areas are obtained by traversing the original endoscopic image.
[0007] (2) Based on the detected endoscope highlight area, the near-light radiation model and light source distance attenuation are used to complete the image of the area and obtain the initial depth estimation map;
[0008] (3) The initial depth map is decomposed into multiple scales through nonlinear filtering, and the multi-scale residuals are fused with the color space intensity to obtain the depth estimation information of the endoscopic image.
[0009] The present invention is based on endoscopic images obtained during clinical surgery and defines an adaptive mirror reflection removal method according to the radiation characteristics of the endoscopic environment and light source. By utilizing color space information constraints, the initial depth map can be effectively estimated from the original endoscopic image. Then, multi-scale decomposition is used to restore the depth gradient details of the local area of the initial depth map, effectively improving the expression of the coherent boundary information of the initial depth map. Finally, the depth information of the image is converted into a fusion of multi-scale decomposition residuals. This process can further optimize the depth estimation of the endoscopic depth estimation. The present invention provides a new idea for feature extraction for applications such as endoscopic image depth estimation, lesion detection, and three-dimensional reconstruction. It solves the problems of blurred boundary information, missing visual far point depth, and unclear depth change levels in the current endoscopic image depth estimation.
[0010] Also provided is an endoscopic image depth estimation device, comprising:
[0011] A highlight region extraction module is configured to collect clinical surgical endoscopic image data and obtain highlight regions based on nonlinear filtering and significant pixel distribution by traversing the original endoscopic image;
[0012] The completion module is configured to complete the image of the detected endoscope highlight area using the near-light radiation model and light source distance attenuation to obtain the initial depth estimate.
[0013] picture;
[0014] The decomposition and fusion module is configured to perform multi-scale decomposition of the initial depth map through nonlinear filtering, fuse the multi-scale residuals with the color space intensity, and obtain the depth estimation information of the endoscopic image. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 4 is a flow chart of the method for estimating depth of an endoscopic image according to the present invention.
[0016] Figure 2 FIG. 4 is a flow chart of a specific embodiment of the method for estimating depth of an endoscopic image according to the present invention.
[0017] Figure 3 1 is a flow chart of step (1) of the method for estimating depth of an endoscopic image according to the present invention.
[0018] Figure 4 3 is a flow chart of step (2) of the method for estimating depth of an endoscopic image according to the present invention.
[0019] Figure 5 3 is a flow chart of step (3) of the method for estimating depth of an endoscopic image according to the present invention. DETAILED DESCRIPTION
[0020] The purpose of this invention is to address the existing depth estimation issues in endoscopic images by providing a depth estimation method for endoscopic images. This method uses nonlinear filtering and significant pixel distribution to obtain highlight regions in endoscopic images. Image completion methods are then used to complete missing information in these highlight regions, eliminating visual misperceptions. Subsequently, an initial depth map is estimated by constructing a surgical space illumination variation model using the endoscope's near-beam radiation field. Finally, nonlinear filtering is used to perform a multi-scale decomposition of the initial depth map. The multi-scale residuals are then fused with color space intensity to obtain depth estimation information for the endoscopic image.
[0021] like Figure 1 As shown, this endoscopic image depth estimation method includes the following steps:
[0022] (1) Collect clinical surgical endoscopic image data and obtain highlight areas based on nonlinear filtering and significant pixel distribution. The highlight areas are obtained by traversing the original endoscopic image.
[0023] (2) Based on the detected endoscope highlight area, the near-light radiation model and light source distance attenuation are used to complete the image of the area and obtain the initial depth estimation map;
[0024] (3) The initial depth map is decomposed into multiple scales through nonlinear filtering, and the multi-scale residuals are fused with the color space intensity to obtain the depth estimation information of the endoscopic image.
[0025] The present invention is based on endoscopic images obtained during clinical surgery and defines an adaptive mirror reflection removal method according to the radiation characteristics of the endoscopic environment and light source. By utilizing color space information constraints, the initial depth map can be effectively estimated from the original endoscopic image. Then, multi-scale decomposition is used to restore the depth gradient details of the local area of the initial depth map, effectively improving the expression of the coherent boundary information of the initial depth map. Finally, the depth information of the image is converted into a fusion of multi-scale decomposition residuals. This process can further optimize the depth estimation of the endoscopic depth estimation. The present invention provides a new idea for feature extraction for applications such as endoscopic image depth estimation, lesion detection, and three-dimensional reconstruction. It solves the problems of blurred boundary information, missing visual far point depth, and unclear depth change levels in the current endoscopic image depth estimation.
[0026] like Figure 3 As shown, preferably, in step (1), the endoscopic image is enhanced using nonlinear filtering:
[0027] I Enhance=Filter(I (r,g,b) ) (1)
[0028] Among them, I Enhance To enhance the image, Filter is a nonlinear filter function, I (r,g,b) are the three channel components of the original image. In the RGB color model, the gradient variation of the highlight area in the endoscopic image is relatively large. In order to detect the center area of the highlight area, the distribution of significant pixels in the color space is compared respectively. The channel thresholds are:
[0029] ω r =M(I(r))+S(I(r))·α (2)
[0030] ω g =M(I(g))+S(I(g))·β (3)
[0031] Among them, ω r and ω g Represent the thresholds of the red and green channels respectively, I(r) and I(g) represent the red and green channel components of the image respectively, M represents the mean value of the matrix, S represents the standard deviation, α and β represent different weight parameters respectively; the final extracted highlight area is:
[0032] S region ={I(r)≥ω r ∩I(g)≥ω g} (4)
[0033] Finally, by traversing the original endoscopic image to obtain the highlight area, S region The constraints for detecting highlight areas must satisfy both weights.
[0034] like Figure 4 As shown, preferably, in step (2), a low beam radiation model is used to describe the lighting condition of the endoscope:
[0035] I(x)=R scene (x)*M ill (x) (5)
[0036] Where x is the pixel in the image, I(x) represents the image taken by the endoscope, and R scene (x)
[0037] Represents the scene radiation intensity, M ill (x) represents the scene lighting. The near-beam radiation model represents the depth of the image by the attenuation of the distance between the radiation surface and the light source;
[0038] The captured image is decomposed into the product of the scene radiation intensity and the scene illumination, and equation (5) is solved:
[0039]
[0040]
[0041] Where c represents the three channels of the RGB image, and ε represents a very small constant to avoid the denominator being zero. In the initial depth map, the local consistency of illumination is considered by considering the neighboring pixels in a small area around the target pixel. The initial depth map is defined as:
[0042]
[0043] Among them, c m Represents the main color space, c s represents the secondary color space, d map represents the initial depth map, θ(x) represents the region centered at pixel x, and y represents the position index within the region.
[0044] like Figure 5 As shown, preferably, in step (3), based on the initial depth estimation map, multi-scale filtering is used to refine the depth estimation level and clarify the boundaries. The local linear optimization and multi-scale decomposition methods are:
[0045]
[0046]
[0047] Among them, ∈ i represents the smoothing scale-related parameters, represents the linear optimization method, a and b are linear filter constants, ψ i~n is the multi-scale residual of the initial depth estimation map. Assuming two general features: the consistency of light source brightness attenuation distance and surface diffuse reflection, the local linear optimization method is used to perform multi-scale decomposition of the initial depth map. The depth estimation of the endoscopic image is:
[0048]
[0049] Among them, D map (x) represents the depth estimation of the endoscopic image, μ and σ represent ψ n The mean and variance of are used to perform multi-scale fusion using the variance-weighted average difference method, which represents the probability that a pixel is an ambient light pixel.
[0050] Preferably, in step (3), the mathematical model constructed above is solved, and the depth of the image pixels is estimated based on the solution result to obtain an endoscopic image depth estimation method.
[0051] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes the steps of the above-mentioned embodiment method, and the storage medium can be a ROM / RAM, a magnetic disc, an optical disc, a memory card, etc. Therefore, corresponding to the method of the present application, the present application also simultaneously includes an endoscopic image depth estimation device, which is usually represented in the form of a functional module corresponding to each step of the method. The device includes:
[0052] A highlight region extraction module configured to collect clinical surgical endoscopic image data, obtain a highlight region based on nonlinear filtering and significant pixel distribution, and obtain the highlight region by traversing the original endoscopic image;
[0053] A completion module configured to perform image completion on the detected endoscopic highlight region using a low-light radiation model and a light source distance attenuation to obtain an initial depth estimation map;
[0054] A decomposition fusion module configured to perform multi-scale decomposition on the initial depth map through nonlinear filtering, fuse multi-scale residuals and color space intensity, and obtain endoscopic image depth estimation information.
[0055] Preferably, in the highlight region extraction module, the endoscopic image is enhanced by nonlinear filtering:
[0056] I Enhance = Filter(I (r,g,b) ) (1)
[0057] where I Enhance is the enhanced image, Filter is a nonlinear filtering function, and I (r,g,b) is the three channel components of the original image. In the RGB color model, the gradient change of the highlight region in the endoscopic image is large. In order to detect the center region of the highlight region, the distribution of significant pixels in the color space is compared respectively, and the channel threshold is:
[0058] ω r = M(I(r))+S(I(r))·α (2)
[0059] ω g = M(I(g))+S(I(g))·β (3)
[0060] where ω r and ω gThresholds of the red and green channels, I(r) and I(g) represent the red and green channel components of the image, M represents the mean value of the matrix, S represents the standard deviation, and a and b represent different weight parameters; the high-light region is finally extracted as:
[0061] S region = {I(r) ≥ ω r ∩I(g) ≥ ω g} (4)
[0062] Finally, the high-light region S region is obtained by traversing the original endoscopic image.
[0063] Preferably, in the completion module, a low-beam radiation model is used to describe the illumination condition of the endoscope:
[0064] I(x) = R scene (x) * M ill (x) (5)
[0065] where x is a pixel point in the image, I(x) represents the image captured by the endoscope, R scene (x) represents the scene radiance, M ill (x) represents the scene illumination, and the low-beam radiation model represents the depth of the image through the attenuation of the distance between the radiation surface and the light source.
[0066] The captured image is decomposed into the product of the scene radiance and the scene illumination, and formula (5) is solved:
[0067]
[0068]
[0069] where c represents the three channels of the RGB image, and e represents a very small constant to avoid a zero denominator; in the initial depth map, the local consistency of the illumination is considered by considering the neighboring pixels in a small region around the target pixel, and the initial depth map is defined as:
[0070]
[0071] where c m represents the primary color space, c s represents the secondary color space, d map represents the initial depth map, theta(x) represents a region centered at pixel x, and y represents the position index in the region.
[0072] Preferably, in the decomposition and fusion module, based on the initial depth estimation map, multi-scale filtering is used to refine the depth estimation level and clarify the boundaries. The local linear optimization and multi-scale decomposition methods are:
[0073]
[0074]
[0075] Among them, ∈ i represents the smoothing scale-related parameters, represents the linear optimization method, a and b are linear filter constants, ψ i~n is the multi-scale residual of the initial depth estimation map. Assuming two general features: the consistency of light source brightness attenuation distance and surface diffuse reflection, the local linear optimization method is used to perform multi-scale decomposition of the initial depth map. The depth estimation of the endoscopic image is:
[0076]
[0077] Among them, D map (x) represents the depth estimation of the endoscopic image, μ and σ represent ψ n The mean and variance of are used to perform multi-scale fusion using the variance-weighted average difference method, which represents the probability that a pixel is an ambient light pixel.
[0078] Preferably, in the decomposition and fusion module, the mathematical model constructed above is solved, and the depth of the image pixels is estimated according to the solution result to obtain an endoscopic image depth estimation method.
[0079] The present invention is based on endoscopic images obtained during clinical surgery and defines an adaptive mirror reflection removal method according to the radiation characteristics of the endoscopic environment and light source. By utilizing color space information constraints, the initial depth map can be effectively estimated from the original endoscopic image. Then, multi-scale decomposition is used to restore the depth gradient details of the local area of the initial depth map, effectively improving the expression of the coherent boundary information of the initial depth map. Finally, the depth information of the image is converted into a fusion of multi-scale decomposition residuals. This process can further optimize the depth estimation of the endoscopic depth estimation. The present invention provides a new idea for feature extraction for applications such as endoscopic image depth estimation, lesion detection, and three-dimensional reconstruction. It solves the problems of blurred boundary information, missing visual far point depth, and unclear depth change levels in the current endoscopic image depth estimation.
[0080] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention are still within the scope of protection of the technical solution of the present invention.
Claims
1. A method for estimating depth of an endoscopic image, characterized by: It includes the following steps: (1) Collect clinical surgical endoscopic image data and obtain highlight areas based on nonlinear filtering and significant pixel distribution. The highlight areas are obtained by traversing the original endoscopic image. (2) Based on the detected endoscope highlight area, the near-light radiation model and light source distance attenuation are used to complete the image of the area and obtain the initial depth estimation map; (3) Perform multi-scale decomposition of the initial depth map through nonlinear filtering, fuse the multi-scale residuals with the color space intensity, and obtain the depth estimation information of the endoscopic image; In step (2), the low beam radiation model is used to describe the illumination of the endoscope: I(x)=R scene (x)*M ill (x) (5) Where x is the pixel in the image, I(x) represents the image taken by the endoscope, and R scene (x) represents the scene radiation intensity, M ill (x) represents the scene lighting. The near-beam radiation model represents the depth of the image by the attenuation of the distance between the radiation surface and the light source; The captured image is decomposed into the product of the scene radiation intensity and the scene illumination, and equation (5) is solved: Where c represents the three channels of the RGB image, and ε represents a very small constant to avoid the denominator being zero. In the initial depth map, the local consistency of illumination is considered by considering the neighboring pixels in a small area around the target pixel. The initial depth map is defined as: Among them, c m Represents the main color space, c s represents the secondary color space, d map represents the initial depth map, θ(x) represents the area centered on pixel x, and y represents the position index within the area; in step (3), based on the initial depth estimation map, multi-scale filtering is used to refine the depth estimation hierarchy and clarify the boundaries. The local linear optimization and multi-scale decomposition methods are as follows: Among them, ∈ i represents the smoothing scale-related parameters, represents the linear optimization method, a and b are linear filter constants, ψ i~n is the multi-scale residual of the initial depth estimation map. Assuming two general features: the consistency of light source brightness attenuation distance and surface diffuse reflection, the local linear optimization method is used to perform multi-scale decomposition of the initial depth map. The depth estimation of the endoscopic image is: Among them, D map (x) represents the depth estimation of the endoscopic image, μ and σ represent ψ n The mean and variance of are used to perform multi-scale fusion using the variance-weighted average difference method, which represents the probability that a pixel is an ambient light pixel.
2. The method for estimating depth of an endoscopic image according to claim 1, wherein: In step (1), the endoscopic image is enhanced using nonlinear filtering: I Enhance =Filter(I (r,g,b) ) (1) Among them, I Enhance To enhance the image, Filter is a nonlinear filter function, I (r,g,b) are the three channel components of the original image. In the RGB color model, the gradient variation of the highlight area in the endoscopic image is relatively large. In order to detect the center area of the highlight area, the distribution of significant pixels in the color space is compared respectively. The channel thresholds are: ω r =M(I(r))+S(I(r))·α (2) oh g =M(I(g))+S(I(g))·β (3) Among them, ω r and ω g Represent the thresholds of the red and green channels respectively, I(r) and I(g) represent the red and green channel components of the image respectively, M represents the mean value of the matrix, S represents the standard deviation, α and β represent different weight parameters respectively; the final extracted highlight area is: S region ={I(r)≥ω r ∩I(g)≥ω g } (4) Finally, by traversing the original endoscopic image to obtain the highlight area, S region The constraints for detecting highlight areas must satisfy both weights.
3. The method for estimating depth of an endoscopic image according to claim 2, wherein: In the step (3), the mathematical model constructed above is solved, and the depth of the image pixels is estimated based on the solution result to obtain an endoscopic image depth estimation method.
4. An endoscope image depth estimation device, characterized in that: It includes: A highlight region extraction module is configured to collect clinical surgical endoscopic image data and obtain highlight regions based on nonlinear filtering and significant pixel distribution by traversing the original endoscopic image; The completion module is configured to complete the image of the detected endoscope highlight area using the near-light radiation model and light source distance attenuation to obtain an initial depth estimation map; A decomposition and fusion module is configured to perform multi-scale decomposition of the initial depth map through nonlinear filtering, fuse the multi-scale residuals with the color space intensity, and obtain the depth estimation information of the endoscopic image; In the completion module, the low beam radiation model is used to describe the lighting conditions of the endoscope: I(x)=R scene (x)*M ill (x) (5) Where x is the pixel in the image, I(x) represents the image taken by the endoscope, and R scene (x) represents the scene radiation intensity, M ill (x) represents the scene lighting. The near-beam radiation model represents the depth of the image by the attenuation of the distance between the radiation surface and the light source; The captured image is decomposed into the product of the scene radiation intensity and the scene illumination, and equation (5) is solved: Where c represents the three channels of the RGB image, and ε represents a very small constant to avoid the denominator being zero. In the initial depth map, the local consistency of illumination is considered by considering the neighboring pixels in a small area around the target pixel. The initial depth map is defined as: Among them, c m Represents the main color space, c s represents the secondary color space, d map represents the initial depth map, θ(x) represents the area centered on pixel x, and y represents the position index within the area; In the decomposition and fusion module, based on the initial depth estimation map, multi-scale filtering is used to refine the depth estimation level and clarify the boundaries. The local linear optimization and multi-scale decomposition methods are as follows: Among them, ∈ i represents the smoothing scale-related parameters, represents the linear optimization method, a and b are linear filter constants, ψ i~n is the multi-scale residual of the initial depth estimation map. Assuming two general features: the consistency of light source brightness attenuation distance and surface diffuse reflection, the local linear optimization method is used to perform multi-scale decomposition of the initial depth map. The depth estimation of the endoscopic image is: Among them, D map (x) represents the depth estimation of the endoscopic image, μ and σ represent ψ n The mean and variance of are used to perform multi-scale fusion using the variance-weighted average difference method, which represents the probability that a pixel is an ambient light pixel.
5. The endoscopic image depth estimation device according to claim 4, characterized in that: In the highlight region extraction module, nonlinear filtering is used to enhance the endoscopic image: I Enhance =Filter(I (r,g,b) ) (1) Among them, I Enhance To enhance the image, Filter is a nonlinear filter function, I (r,b,b) are the three channel components of the original image. In the RGB color model, the gradient variation of the highlight area in the endoscopic image is relatively large. In order to detect the center area of the highlight area, the distribution of significant pixels in the color space is compared respectively. The channel thresholds are: ω r =M(I(r))+S(I(r))·α (2) oh g =M(I(g))+S(I(g))·β (3) Among them, ω r and ω g Represent the thresholds of the red and green channels respectively, I(r) and I(g) represent the red and green channel components of the image respectively, M represents the mean value of the matrix, S represents the standard deviation, α and β represent different weight parameters respectively; the final extracted highlight area is: S region ={I(r)≥ω r ∩I(g)≥ω g } (4) Finally, by traversing the original endoscopic image to obtain the highlight area, S region The constraints for detecting highlight areas must satisfy both weights.
6. The endoscopic image depth estimation device according to claim 5, characterized in that: In the decomposition and fusion module, the mathematical model constructed above is solved, and the depth of the image pixels is estimated based on the solution result to obtain an endoscopic image depth estimation method.
Citation Information
Patent Citations
Endoscope weak texture image enhancement method and device
CN113538295A
Endoscopic stereo matching method and apparatus using direct attenuation model
US20190365213A1