An unsupervised multi-frame endoscopic scene depth estimation method and device
By using an unsupervised multi-frame method and a learnable block matching module in endoscopic scene depth estimation, the depth estimation model is optimized, and the robustness of endoscopic scene depth maps in brightness fluctuations and non-Lambertian reflection areas is solved, achieving higher precision depth map generation.
Patent Information
- Application Number
- CN202210582504.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-26
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-05-26
AI Technical Summary
The existing endoscopic scene depth maps are not characterized by low-texture and uniform texture areas due to non-Lambertian reflections and mutual reflections in adaptive propagation, and have poor robustness in brightness fluctuations.
The unsupervised multi-frame endoscopic scene depth estimation method is adopted, and the unsupervised multi-frame monocular training depth estimation model is optimized through the learnable block matching module and the autonomous teaching and cross-teaching paradigm, and the optimized endoscopic scene depth map is generated.
It improves robustness in areas with large brightness variations, can better adaptive propagation, induce more unique characterization in low-texture and uniform texture areas, and improves the accuracy of endoscopic scene depth maps.
Smart Images

Figure CN115661224B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of surgical navigation, and in particular, to an unsupervised multi-frame endoscopic scene depth estimation method and device. Background Art
[0002] Monocular depth estimation has proven to be a practical and general-purpose technology with various applications, such as in simultaneous localization and mapping, surgical navigation, and augmented reality. Although hardware sensors can capture depth ranges, professional hardware cannot collect dense depth maps. Currently, many methods position monocular depth of a single RGB camera as a new view synthesis method, thus eliminating the limitation of obtaining training depth data with expensive hardware sensors.
[0003] However, the accuracy of depth estimation is far from satisfactory. The predicted depth may be incorrect, and there may be a situation where the predicted depth has reached infinity while the photometric error still shows a low value. Secondly, when the camera and light source move, the inter-frame illuminance of the same anatomical structure varies greatly, so the assumption of constant brightness is not applicable in endoscopes. In addition, strong non-Lambertian reflections and mutual reflections occur on the surfaces of smooth tissues and organ fluids, resulting in unclear characterization of low-texture and uniform-texture regions. In these regions, the photometric error and matching cost are easily confused by drastic brightness fluctuations, thereby reducing the robustness of endoscopic scene images. Summary of the Invention
[0004] Embodiments of this application provide an unsupervised multi-frame endoscopic scene depth estimation method and device, which are used to solve the following technical problems: In the adaptive propagation of existing endoscopic scene depth maps, due to non-Lambertian reflections and mutual reflections, the characterization of low-texture and uniform-texture regions is not obvious, and the robustness effect in brightness fluctuation regions is poor.
[0005] Embodiments of this application adopt the following technical solutions:
[0006] On the one hand, an embodiment of the present application provides an unsupervised multi-frame endoscopic scene depth estimation method, which is characterized in that the method includes: determining a point cloud conversion view corresponding to an endoscopic scene image, and synthesizing a target frame and a source frame of the point cloud conversion view to obtain a synthesized frame; extracting key points in the synthesized frame, and determining the depth of sampled pixels according to the key points; combining the depth of the sampled pixels with the synthesized frame to obtain a photometric loss; obtaining a cross-teaching consistency loss and a self-teaching consistency loss according to the target frame; obtaining a total optimization loss through the photometric loss, the cross-teaching consistency loss, the self-teaching consistency loss, and a preset edge-aware smoothness loss; and optimizing and training an unsupervised multi-frame monocular training depth estimation model according to the total optimization loss, so that the unsupervised multi-frame monocular training depth estimation model outputs an optimized endoscopic scene depth map.
[0007] An embodiment of the present application optimizes a new unsupervised multi-frame endoscopic scene depth estimation model through a learnable block matching module and a self-teaching and cross-teaching paradigm, that is, optimizing the depth map estimation model of the endoscope, so as to obtain an optimized endoscopic scene depth map. It solves the problem that the endoscope is unavailable under constant brightness, and when the camera and the light source move together, the inter-frame illuminance of the same anatomical structure changes greatly. At the same time, it also solves the problem that strong non-Lambertian reflections and mutual reflections will occur on the surfaces of smooth tissues and organ fluids, which are likely to cause brightness fluctuation confusion. By using a learnable patch matching module, it can better adaptively propagate and induce more unique representations in low-texture and uniform-texture regions. The cross-teaching and self-teaching modes enhance the robustness to severe brightness fluctuations, which is beneficial to obtaining a better endoscopic scene depth map, and is of great significance for positioning and mapping, surgical navigation, and augmented reality, etc.
[0008] In a feasible implementation manner, determining a point cloud conversion view corresponding to an endoscopic scene image, and synthesizing a target frame and a source frame of the point cloud conversion view to obtain a synthesized frame specifically includes: predicting the depth of each pixel point of the acquired endoscopic scene image through a depth network and a pose network to obtain a relative pose image; performing point cloud conversion on the relative pose image to obtain a point cloud conversion view; obtaining a view synthesis formula according to the mapping framework from the source frame view to the target frame view, the parameters of the camera itself of the endoscopic scene, the relative pose from the target frame view to the source frame view, the pixel coordinates of the target frame view, and the depth map of the target frame; and obtaining the synthesized frame according to the source frame and the target frame.
[0009] In a feasible implementation, after determining the point cloud conversion view corresponding to the endoscopic scene image and synthesizing the target frame and the source frame of the point cloud conversion view to obtain a synthesized frame, the method further includes: obtaining a photometric loss of the synthesized frame according to a weight coefficient, a weighted loss, a structural similarity term, and a structural similarity term loss.
[0010] In a feasible implementation, before extracting key points in the synthesized frame and determining the depth of sampled pixels according to the key points, the method further includes: constructing a plane sweep stereo matching cost algorithm through a preset frame sequence; wherein, the plane sweep stereo matching cost algorithm is constructed based on the target camera frustum; performing depth feature mapping on the source frame at each different depth through the plane sweep stereo matching cost algorithm to obtain mapped features; wherein, the mapped features include each candidate depth, the camera intrinsic pose, and the relative pose projected into the camera space; calculating an average value of the distances between the feature mappings of the target frame according to the mapped features to obtain a final matching cost; performing depth regression on a depth network through the final matching cost and the mapped features to keep the updated depth range unchanged.
[0011] In a feasible implementation, extracting key points in the synthesized frame and determining the depth of sampled pixels according to the key points specifically includes: extracting key points in the synthesized frame through a visual odometry algorithm to obtain key points p in the synthesized frame k ; according to c(p k ; r) = F(p k ) ⊙ F(p′ k ) / l, obtaining key point related quantities; wherein, c is a related vector, r is a search range, F is a feature map, ⊙ is a point integral, l is the length of the feature descriptor, and p′ k is the feature similarity of pixels around the key point p k ; integrating the key point related quantities into a three-dimensional grid and constructing an autocorrelation volume; decoding an offset field in the autocorrelation volume through a preset two-layer convolutional network; according to obtaining a support domain wherein, is the pixel of each key point p k , is an additional offset field, is a 2D offset grid; according to obtaining the depth of the sampled pixels
[0012] The key points extracted in the embodiments of the present application are enhanced with local blocks, that is, the support domain. Sampling the neighboring pixels through adaptive propagation instead of using a static neighborhood set can make the adaptive propagation tend to aggregate pixels from the same surface, avoiding catastrophic errors caused by only using a fixed pattern.
[0013] In a feasible implementation manner, the photometric loss is obtained by combining the sampled pixel depth with the synthesized frame, specifically including: According to the depth synthesis formula: Obtain the support domain from the source frame view s to the target frame view t where K is the parameter of the camera itself for obtaining the endoscopic scene, and M t→s is the relative pose from the target frame view t to the source frame view s, is the pixel coordinate of the key point p in the target frame view t k ; is the depth map of the support domain of the target frame; According to Obtain the support domain synthesized frame from the source frame to the target frame where <·> is the warping degree calculation operation; According to Obtain the photometric loss L ph ; where I t (pk) is the key point of the target frame view t, and φ(I t (p), I s→t (p)) is the photometric loss of the synthesized frame.
[0014] In the embodiments of the present application, the photometric loss is accumulated on each support domain, which can make the area of the effective gradient larger and the convergence basin wider. And using two source frames effectively alleviates the negative impact of pixel occlusion on the photometric loss in terms of the photometric loss.
[0015] In a feasible implementation manner, according to the target frame, the cross-teaching consistency loss and the self-teaching consistency loss are obtained, specifically including: According to Obtain the cross-teaching consistency loss L ct ; where D t (p) is the depth map of the target frame, is the depth map of the discardable depth estimation network; According to Obtain the self-teaching consistency loss Lst; where R(p) is the unoccluded mask determined by the random mask, is the feature map after the consistency transformation of the frame exported by the appearance simulator and the original frame.
[0016] In the embodiments of the present application, through a deep network trained based on AF-SfMLeamer, the unsupervised multi-frame monocular training depth estimation model is guided to the correct depth, and even in regions with large brightness variations, better results can be obtained. Moreover, in the self-teaching paradigm, the target frame and the source frame exported by the appearance simulator are used to form the final matching cost, and the two generated depth maps are forced to be aligned consistently, which can better enhance the robustness to brightness fluctuations and occlusions.
[0017] In a feasible implementation manner, the appearance simulator includes: a random gamma correction module, a color jitter module, and a masking module; the random gamma correction module is a non-linear mapping module for adjusting the illuminance of the endoscopic scene image; the color jitter module is used to adjust the brightness jitter, saturation jitter, hue jitter, and contrast jitter of the endoscopic scene image; the masking module is used to simulate the occlusion between multiple input frames of the endoscopic scene image to improve the accuracy of the unsupervised multi-frame monocular training depth estimation model.
[0018] In the embodiments of the present application, masking is used to simulate the occlusion between multiple input frames. In the generated binary mask, some parts of the target frame are randomly cropped, and then the inverted binary mask is used as the real data simulating occlusion in the self-teaching consistency loss, which helps to make the model insensitive to occlusions.
[0019] In a feasible implementation manner, the total optimization loss is obtained through the photometric loss, the cross-teaching consistency loss, the self-teaching consistency loss, and the preset edge-aware smoothness loss, specifically including: according to to obtain the preset edge-aware smoothness loss; where is the gradient operator, I t (p) is the target frame, and D(p) is the edge-aware depth map; according to L = λ1L ph + λ2L ct + λ3L st + λ4L es , the total optimization loss is obtained; where λ is a weight parameter, L ph is the photometric loss, L ct is the cross-teaching consistency loss, and L st is the self-teaching consistency loss.
[0020] On the other hand, the embodiments of the present application also provide an implicit motion compensation video object segmentation device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions that can be executed by the at least one processor, so that the at least one processor can execute the method for unsupervised multi-frame endoscopic scene depth estimation according to any one of the above embodiments.
[0021] This application proposes an unsupervised multi-frame endoscopic scene depth estimation method. By optimizing the unsupervised multi-frame monocular training depth estimation model through the total optimization loss, an optimized endoscopic scene depth map is obtained. It solves the problems that under constant brightness, the endoscope is unavailable, and when the camera moves together with the light source, the inter-frame illuminance of the same anatomical structure varies greatly. At the same time, it also solves the problem that strong non-Lambertian reflections and mutual reflections occur on the surfaces of smooth tissues and organ fluids, which easily cause confusion due to brightness fluctuations. A learnable block matching module is adopted, which can better adaptively propagate and induce more unique representations in low-texture and uniform-texture regions. The cross-teaching and self-teaching modes enhance the robustness to severe brightness fluctuations, which is beneficial to obtaining a better endoscopic scene depth map. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:
[0023] Figure 1 It is a flowchart of an unsupervised multi-frame endoscopic scene depth estimation method provided by an embodiment of the present application;
[0024] Figure 2 It is a diagram of an unsupervised evaluation framework provided by an embodiment of the present application;
[0025] Figure 3 It is a diagram of an appearance simulator module provided by an embodiment of the present application;
[0026] Figure 4 It is a schematic structural diagram of an unsupervised multi-frame endoscopic scene depth estimation device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] In order to enable those skilled in the art to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0028] An embodiment of the present application provides an unsupervised multi-frame endoscopic scene depth estimation method, as Figure 1As shown, the multi-frame endoscopic scene depth estimation method specifically includes steps S101 - S106:
[0029] S101. Determine the point cloud conversion view corresponding to the endoscopic scene image, and synthesize the target frame and the source frame of the point cloud conversion view to obtain a synthesized frame.
[0030] Specifically, through a depth network and a pose network, the depth of each pixel point of the acquired endoscopic scene image is predicted to obtain a relative pose image. The relative pose image is subjected to point cloud conversion to obtain a point cloud conversion view. According to p s→t = KM t→s D t (p)K -1 p t , the view synthesis formula is obtained. Among them, p s→t is the mapping framework from the source frame view s to the target frame view t, K is the parameter of the camera itself for acquiring the endoscopic scene, M t→ s is the relative pose from the target frame view t to the source frame view s, p t is the pixel coordinate of the target frame view t, D t (p) is the depth map of the target frame. According to I s→t (p) = I s <p s→t >>, the synthesized frame is obtained. Among them, I t (p) is the target frame, I s (p) is the source frame, <·> is the warping degree calculation operation, and I s→t (p) is the synthesized frame from the source frame to the target frame.
[0031] Furthermore, according to the photometric loss φ(I t (p), I s→t (p)) of the synthesized frame is obtained. Among them, α is the weight coefficient, L1 is the weighted loss, SSIM is the structural similarity term, L SSIM is the structural similarity term loss, and L1 = ||I t (p) - I s→t (p)||.
[0032] As a feasible implementation, a view synthesis framework is constructed through a depth network and a pose network. When the depth of each pixel has been predicted, the pixel points can be back-projected into a 3D camera space. Then, according to the predicted relative pose, the generated point cloud is converted into a point cloud conversion view. Through the view synthesis formula, the target frame view and the source frame view are combined and transformed, and then, through the warping degree calculation operation, the appearance difference between the target frame and the synthesized frame is used as a monitoring signal for the entire framework training pipeline to be transmitted.
[0033] In one embodiment, Figure 2 FIG. is an unsupervised evaluation framework diagram provided by an embodiment of the present application. As Figure 2 shown, in the depth estimation branch, the source frame framework and the target frame framework are respectively the frameworks of the source frame view and the target frame view, and then view synthesis preprocessing is performed. After that, through a warping degree calculation operation, the synthesis of the source frame to the target frame is realized to obtain a synthesized frame, and a method using a quantization appearance difference criterion is used to obtain the photometric loss of the synthesized frame with a weighted loss and a structural similarity term loss.
[0034] S102. Extract key points in the synthesized frame, and determine the sampling pixel depth according to the key points.
[0035] Specifically, a plane sweep stereo matching cost algorithm is constructed through a preset frame sequence. Among them, the plane sweep stereo matching cost algorithm is constructed on the target camera frustum. Through the plane sweep stereo matching cost algorithm, depth feature mapping is performed on each source frame with a different depth to obtain a mapped feature. Among them, the mapped feature includes each candidate depth, the camera's intrinsic pose, and the relative pose projected into the camera space. According to the mapped feature, the average value of the distances between the feature mappings of the target frame is calculated to obtain the final matching cost. Through the final matching cost and the mapped feature, depth regression is performed on the depth network to keep the updated depth range unchanged.
[0036] In one embodiment, as Figure 2 shown, in planar flow, a multi-frame sequence is utilized, and a plane sweep stereo matching cost algorithm is constructed on the target camera frustum to evaluate the geometric coherence between the pixels from the target frame and the source frames with different depth values. In the linear interval between the maximum depth value and the minimum depth value of a set of front-parallel depth planes, depth feature mapping is performed on each source frame, and it is projected into the target camera space using each candidate depth, the camera's intrinsic pose, and the relative pose predicted from the pose network. And the final cost measure is the average value of the distances between the warped feature mappings and the feature mappings from the target frame among all source frames. Then, the final matching cost and the feature mapping are used as inputs and connected to the depth network to achieve depth regression.
[0037] In one embodiment, the maximum depth value and the minimum depth value are set to be learned in the time amount. In each training iteration, the average minimum value and maximum value of multi-batch depth predictions are utilized, and the maximum depth value and the minimum depth value are updated using an exponential moving average with a momentum of 0.99. The updated depth range is saved together with the weights of the unsupervised multi-frame monocular training depth estimation model and remains unchanged during the evaluation phase in the depth estimation branch.
[0038] Further, through visual odometry, key points are extracted from the synthesized frame to obtain key point p in the synthesized frame k . According to c(p k ; r) = F(p k ) ⊙ F(p′ k ) / l, key point related quantities are obtained. Wherein, c is the correlation vector, r is the search range, F is the feature map, ⊙ is point integration, l is the length of the feature descriptor, and p′ k is the feature similarity of the pixels around key point p k . The key point related quantities are integrated into a three-dimensional grid, and an autocorrelation volume is constructed. Through a preset two-layer convolutional network, the offset field in the autocorrelation volume is decoded. According to the support domain is obtained wherein, is the pixel of each key point p k , is the additional offset field, is the 2D offset grid. According to the sampled pixel depth is obtained
[0039] As a feasible implementation manner, as Figure 2 shown, in the feature extraction stage, key points are extracted. Representative key points are selected as the central pixels of the block matching module, and then each key point is combined with a local block to enhance their recognition ability, wherein the block matching mode is dynamic, and the neighborhood pixels are aggregated according to pixel-level correlation.
[0040] In one embodiment, for the detection of key points, an efficient visual odometry (DirectSparse Odometry, DSO) is used to implement the extraction of key points, and the support domain is used to enhance the algorithm. Then, the neighborhood pixels are sampled through adaptive propagation, and the pixels on the same surface are aggregated. Next, an offset field generator is preset to learn an additional offset field. First, the feature similarity between the target key point and its surrounding pixels is measured to facilitate the derivation of the additional offset field, and then the key point related quantities are obtained. Then, all the correlation vectors are integrated into a three-dimensional grid, and an autocorrelation volume is constructed. Finally, two-layer convolution is applied to decode the offset field in the autocorrelation volume, and then the sampled pixel depth at the key points is obtained through a warping operation.
[0041] S103. The depth estimation branch combines the sampled pixel depth with the synthesized frame to obtain a photometric loss.
[0042] Specifically, according to the depth synthesis formula: the support domain from the source frame view s to the target frame view t is obtained Among them, K is the parameter of the camera itself for obtaining the endoscope scene, M t→s is the relative pose from the target frame view t to the source frame view s, Ten-day target frame view t key point p k The pixel coordinates of is the depth map of the target frame support domain. Get the support domain synthetic frame from the source frame to the target frame Where <·> is the warpage calculation operation. The luminosity loss L is obtained ph Among them, I t (p k ) is the key point of the target frame view t, φ(I t (p), I s→t (p)) is the luminance loss of the synthetic frame.
[0043] In one embodiment, Figure 2 As shown in the figure, at the final stage of the depth estimation branch, a learnable block matching photometric loss is obtained. Based on the support domain synthetic frame and the synthetic frame loss in the previous article, a block matching based photometric loss can be obtained. The photometric loss is accumulated on each support domain, making the effective gradient area larger and the convergence basin wider. In addition, using two source frames to follow the one with the smallest photometric loss in terms of photometric loss can alleviate the negative impact of pixel occlusion on photometric loss.
[0044] S104. According to the target frame, a cross-teaching consistency loss and an autonomous teaching consistency loss are obtained.
[0045] Specifically, according to Get the cross-teaching consistency loss L ct Among them, D t (p) is the depth map of the target frame, is the depth map of the discardable depth estimation network. Get the autonomous teaching consistency loss L st . Where R(p) is the unoccluded mask determined by the random mask, Feature maps after consistent transformation between frames exported by the appearance simulator and the original frames.
[0046] In one embodiment, Figure 2As shown in the last block diagram, the self-teaching branch and the cross-teaching branch can also obtain the self-teaching consistency loss and the cross-teaching consistency loss based on the comparison depth map. In the cross-teaching paradigm, the depth network trained by AF-SfMLearner is used to enable the unsupervised multi-frame monocular training depth estimation model to perform correct depth estimation, and it can also memorize more effectively in areas with large brightness changes. In the cross-teaching consistency loss, instead of directly using their absolute difference, the absolute difference is normalized by their sum, so that points with different absolute depths can be treated equally during the optimization process. In addition, the output is scale-invariant, and its natural range is from 0 to 1, which is beneficial to improving the numerical stability during the entire training process.
[0047] The self-teaching paradigm can immunize the unsupervised multi-frame monocular training depth estimation model against harmful information in the final matching cost, such as harmful information caused by brightness fluctuations and occlusions, and focus on valuable elements. An appearance simulator composed of random gamma correction, color jitter, and masking is adopted to simulate edge cases. The target frame and the source frame derived from the appearance simulator are used to construct the matching cost respectively, and the two generated depth maps are forced to be consistent with each other, enhancing the robustness against brightness fluctuations and occlusions.
[0048] Among them, the appearance simulator includes: a random gamma correction module, a color jitter module, and a masking module. The random gamma correction module is a non-linear mapping module used to adjust the illuminance of the endoscopic scene image. The color jitter module is used to adjust the brightness jitter, saturation jitter, hue jitter, and contrast jitter of the endoscopic scene image. The masking module is used to simulate the occlusion between multiple input frames of the endoscopic scene image to improve the accuracy of the unsupervised multi-frame monocular training depth estimation model.
[0049] In one embodiment, Figure 3 is the module diagram of the appearance simulator provided by the embodiment of the present application. As Figure 3 shown, gamma correction is a non-linear mapping used to adjust the illuminance of the image and simulate different illuminations to achieve the situation of severe brightness fluctuations during the data acquisition process of the endoscopic scene due to the change of illumination. Color jitter simulates the complex non-Lambertian reflection and mutual reflection in the minimally invasive surgical environment by integrating random color jitter. The purpose of masking is to simulate the occlusion between multiple input frames. In a generated binary mask, some parts of the target frame are randomly cropped, and then the inverted binary mask is used as the real data simulating occlusion in the self-teaching consistency loss, which helps to make the model insensitive to occlusion. Even in the presence of occlusion, the model can still correctly predict the depth of the unoccluded area.
[0050] S105. Obtain the total optimization loss according to the photometric loss, the cross-teaching consistency loss, the self-teaching consistency loss, and the preset edge-aware smoothness loss.
[0051] Specifically, according to obtain the preset edge-aware smoothness loss. Among them, is the gradient operator, I t (p) is the target frame, and D(p) is the edge-aware depth map. According to L = λ1L pn + λ2L ct + λ3L st + λ4L es , obtain the total optimization loss. Among them, λ is the weight parameter, L ph is the photometric loss, L ct is the cross-teaching consistency loss, and L st is the self-teaching consistency loss.
[0052] In one embodiment, on the basis of the photometric loss, the cross-teaching consistency loss, and the self-teaching consistency loss, a preset edge-aware smoothness loss is further added to obtain the total optimization loss. The weight parameters can be set as λ1 = 1, λ2 = 0.02, λ3 = 0.002, and λ4 = 0.001.
[0053] S106. Optimize and train the unsupervised multi-frame monocular training depth estimation model according to the total optimization loss, so that the unsupervised multi-frame monocular training depth estimation model outputs an optimized endoscopic scene depth map.
[0054] Specifically, according to the obtained total optimization loss, optimize and train the unsupervised multi-frame monocular training depth estimation model to obtain a new optimized unsupervised multi-frame monocular training depth estimation model, and then perform optimized output according to this model to obtain an optimized endoscopic scene depth map.
[0055] In addition, the embodiment of the present application also provides an unsupervised multi-frame endoscopic scene depth estimation device, as Figure 4 shown, the unsupervised multi-frame endoscopic scene depth estimation device 400 specifically includes:
[0056] At least one processor 401, and a memory communicatively connected to the at least one processor 401. Among them, the memory 402 stores instructions that can be executed by the at least one processor 401, so that the at least one processor 401 can execute:
[0057] Determine the point cloud conversion view corresponding to the endoscopic scene image, and synthesize the target frame and the source frame of the point cloud conversion view to obtain a synthesized frame;
[0058] Extract key points in the synthesized frame, and determine the sampled pixel depth based on the key points;
[0059] Combine the sampled pixel depth with the synthesized frame to obtain the photometric loss;
[0060] Obtain the cross-teaching consistency loss and the self-teaching consistency loss based on the target frame;
[0061] Obtain the total optimization loss through the photometric loss, the cross-teaching consistency loss, the self-teaching consistency loss, and the preset edge-aware smoothness loss;
[0062] Optimize and train the unsupervised multi-frame monocular training depth estimation model according to the total optimization loss, so that the unsupervised multi-frame monocular training depth estimation model outputs an optimized endoscopic scene depth map.
[0063] This application proposes an unsupervised multi-frame endoscopic scene depth estimation method, which optimizes the unsupervised multi-frame monocular training depth estimation model through the total optimization loss to obtain an optimized endoscopic scene depth map. It solves the problem that the endoscope is unavailable under the condition of constant brightness, and the problem that the inter-frame illuminance of the same anatomical structure changes greatly when the camera and the light source move together. At the same time, it also solves the problem that strong non-Lambertian reflections and mutual reflections are generated on the surfaces of smooth tissues and organ fluids, which are likely to cause confusion in brightness fluctuations. By using a learnable block matching module, it can better adaptively propagate, induce more unique representations in low-texture and uniform-texture regions, and the cross-teaching and self-teaching modes enhance the robustness to severe brightness fluctuations, which is beneficial to obtaining a better endoscopic scene depth map.
[0064] Each embodiment in this application is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0065] The above describes specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0066] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present application shall be included within the scope of the claims of the present application.
Claims
1. An unsupervised multi-frame endoscopic scene depth estimation method, characterized in that The method includes: Determine the point cloud conversion view corresponding to the endoscopic scene image, and synthesize the target frame and the source frame of the point cloud conversion view to obtain a synthesized frame; Extract key points from the synthesized frame, and determine the sampled pixel depth according to the key points; Combine the sampled pixel depth with the synthesized frame to obtain a photometric loss, specifically including: According to the depth synthesis formula: Obtain the support region from the source frame view s to the target frame view t where K is the parameter of the camera itself for obtaining the endoscopic scene, and M t→s is the relative pose from the target frame view t to the source frame view s, is the pixel coordinate of the key point p in the target frame view t k , is the depth map of the target frame support region; According to obtain the support field synthesis frame from the source frame to the target frame where <·> is the warping degree calculation operation; According to obtain the photometric loss L ph ; where I t (p k ) is the key point of the target frame view t, and φ(I t (p), I s→t (p)) is the photometric loss of the synthesized frame; Obtain a cross-teaching consistency loss and an auto-teaching consistency loss according to the target frame. Obtain a cross-teaching consistency loss and an auto-teaching consistency loss according to the target frame, specifically including: According to obtain the cross-teaching consistency loss L ct ; where D t (p) is the depth map of the target frame, is the depth map of the discardable depth estimation network; According to obtain the autonomous teaching consistency loss L st ; where R(p) is the unoccluded mask determined by the random mask determined by the random mask, is the feature map after the consistency transformation of the frame exported by the appearance simulator and the original frame; Obtain a total optimization loss through the photometric loss, the cross-teaching consistency loss, the auto-teaching consistency loss, and a preset edge-aware smoothness loss. Obtain a total optimization loss through the photometric loss, the cross-teaching consistency loss, the auto-teaching consistency loss, and a preset edge-aware smoothness loss, specifically including: According to obtain the preset edge perception smoothness loss L es ; where is a gradient operator, I t (p) is the target frame, and D(p) is the edge perception depth map; According to \(L = \lambda_1L\) ph +\(\lambda_2L\) ct +\(\lambda_3L\) st +\(\lambda_4L\) es , the total optimization loss \(L\) is obtained; where \(\lambda\) is a weight parameter, \(L\) pj is the photometric loss, \(L\) ct is the cross-teaching consistency loss, \(L\) st is the self-teaching consistency loss; Optimize and train an unsupervised multi-frame monocular training depth estimation model according to the total optimization loss, so that the unsupervised multi-frame monocular training depth estimation model outputs an optimized endoscopic scene depth map.
2. The unsupervised multi-frame endoscopic scene depth estimation method according to claim 1, characterized in that Determine the point cloud conversion view corresponding to the endoscopic scene image, and synthesize the target frame and the source frame of the point cloud conversion view to obtain a synthesized frame, specifically including: Predict the depth of each pixel point of the acquired endoscopic scene image through a depth network and a pose network to obtain a relative pose image; Perform point cloud conversion on the relative pose image to obtain a point cloud conversion view; Obtain a view synthesis formula according to the mapping framework from the source frame view to the target frame view, the parameters of the camera itself in the endoscopic scene, the relative pose from the target frame view to the source frame view, the pixel coordinates of the target frame view, and the depth map of the target frame; Obtain the synthesized frame according to the source frame and the target frame.
3. A method for unsupervised multi-frame endoscopic scene depth estimation according to claim 2, characterized in that After determining the point cloud conversion view corresponding to the endoscopic scene image, synthesizing the target frame and the source frame of the point cloud conversion view to obtain a synthesized frame, the method further includes: Obtain a synthesized frame photometric loss according to a weight coefficient, a weighted loss, a structural similarity term, and a structural similarity term loss.
4. A method for unsupervised multi-frame endoscopic scene depth estimation according to claim 1, characterized in that Before extracting the key points from the synthesized frame and determining the sampled pixel depth according to the key points, the method further includes: Construct a plane sweep stereo matching cost algorithm through a preset frame sequence; wherein, the plane sweep stereo matching cost algorithm is constructed based on the target camera frustum; Perform depth feature mapping on the source frame at each different depth through the plane sweep stereo matching cost algorithm to obtain mapped features; wherein, the mapped features include each candidate depth, the camera intrinsic pose, and the relative pose projected into the camera space; Calculate the average value of the distances between the feature mappings of the target frame according to the mapped features to obtain a final matching cost; Perform depth regression on the depth network through the final matching cost and the mapped features.
5. A method for depth estimation of multi-frame endoscopic scenes based on unsupervised learning according to claim 1, characterized in that Extract the key points from the synthesized frame, and determine the sampled pixel depth according to the key points, specifically including: Extract key points from the synthesized frame through visual odometry to obtain the key point p in the synthesized frame k ; According to c(p k ; r) = F(p k ) ⊙ F(p′ k ), the quantities related to key points are obtained; where c is the related vector, r is the search range, F is the feature map, ⊙ is the point integral, l is the length of the feature descriptor, and p′ k is the feature similarity of the pixels around the key point p k . Integrate the key point-related quantities into a three-dimensional grid and construct an autocorrelation volume; Decode the offset field in the autocorrelation volume through a preset two-layer convolutional network; According to obtain the support domain wherein is the pixel of each of the key points p k and is the additional offset field and is the 2D offset grid; According to obtain the sampled pixel depth 6. The method for depth estimation of a multi-frame endoscopic scene based on unsupervised learning according to claim 1, characterized in that The appearance simulator includes: a random gamma correction module, a color jitter module, and a masking module; The random gamma correction module is a non-linear mapping module for adjusting the illuminance of the endoscopic scene image; The color jitter module is used to adjust the brightness jitter, saturation jitter, hue jitter, and contrast jitter of the endoscopic scene image; The masking module is used to simulate the occlusion between multiple input frames of the endoscopic scene image to improve the accuracy of the unsupervised multi-frame monocular training depth estimation model.
7. An unsupervised multi-frame endoscopic scene depth estimation device, characterized in that, The device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, enabling the at least one processor to execute a method for unsupervised multi-frame endoscopic scene depth estimation according to any one of claims 1-6.
Citation Information
Patent Citations
Self-supervised deep network training method, image depth acquisition method and device
CN113888613A
Monocular endoscope depth and pose estimation method and device based on unsupervised learning
CN114022527A