Spectral image processing method for three-dimensional endoscope
By using spectral image processing methods for 3D endoscopes, the problem of insufficient accuracy in spectral image fusion was solved, generating a high-quality interactive 3D navigation model and achieving high-precision 3D reconstruction and visual navigation.
Patent Information
- Application Number
- CN202511083540.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-28
AI Technical Summary
In existing technologies, the spectral image fusion accuracy of 3D endoscopes is insufficient, resulting in the loss of image details and inaccurate 3D reconstruction, which makes it difficult to meet the clinical demand for high-precision, high-realism visualization and navigation.
By receiving single-band spectral images synchronously acquired by multiple sensors at the endoscope, lens distortion correction is performed and the images are segmented into multiple groups of single-spectral image blocks. Cross-spectral registration and multi-level cross-fusion of multispectral images are then performed. Combined with spectral characteristic weights, multi-level cross-optimization fusion of spectral images is carried out to obtain a two-dimensional spectral fusion image. A three-dimensional point cloud coordinate set is constructed through parallax calculation, and back-projected onto the three-dimensional point cloud. Poisson fusion-driven inter-block stitching fusion is then performed to generate an interactive three-dimensional navigation model.
It improves the fusion effect of spectral images and the accuracy of 3D reconstruction, generating a high-quality interactive 3D navigation model that meets the clinical demand for high-precision, high-realism visual navigation.
Smart Images

Figure CN121033336A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of endoscopy, in particular to a spectral image processing method for a three-dimensional endoscope. BACKGROUND
[0002] In the technical field of endoscopy, although the traditional endoscope system can provide basic visualization support, it still faces multiple technical bottlenecks in complex cavity environments. On the one hand, multispectral imaging cannot accurately register across the spectrum due to differences in the physical positions of sensors at different wavelengths, and it is difficult to completely preserve the information of tissue characteristics. On the other hand, the fusion of spectral data and three-dimensional geometric structures often relies on inefficient global registration algorithms, which not only introduce stitching seam artifacts, but also cannot meet the real-time interaction requirements during surgery due to computational delays. In existing solutions, insufficient distortion correction can amplify the boundary error of image block segmentation, and the separate processing of visible light three-dimensional reconstruction and spectral feature mapping further leads to the misalignment of tissue function information and spatial structure, which cannot meet the clinical demand for high-precision and high-fidelity visualization navigation. SUMMARY
[0003] The present application provides a spectral image processing method for a three-dimensional endoscope, which aims to solve the technical problem of insufficient spectral image fusion accuracy in existing three-dimensional endoscope image processing methods, resulting in loss of image details and inaccurate three-dimensional reconstruction.
[0004] The present application provides a spectral image processing method for a three-dimensional endoscope, which includes: receiving a plurality of single-band spectral images synchronously collected and returned by a plurality of sensors at the end of an endoscope; after lens distortion correction of the plurality of single-band spectral images, dividing the plurality of single-band spectral images into a plurality of groups of single-spectral image blocks based on a preset inter-block overlap rate; after block-level cross-spectral registration of the plurality of groups of single-spectral image blocks, performing multi-level cross-optimization fusion of the plurality of groups of single-spectral image blocks based on a predefined spectral characteristic weight, to obtain a two-dimensional spectral fusion image; performing disparity calculation on a pair of binocular visible light images returned by left and right cameras to construct an endoscopic three-dimensional point cloud coordinate set; mapping the two-dimensional spectral fusion image in space to the endoscopic three-dimensional point cloud coordinate set based on a reverse projection sub-pixel mapping algorithm, to obtain a three-dimensional spectral point cloud; performing inter-block stitching fusion driven by Poisson fusion on the three-dimensional spectral point cloud, to output an interactive three-dimensional navigation model.
[0005] One or more technical solutions provided in the present application have at least the following technical effects or advantages:
[0006] The aforementioned spectral image processing method for 3D endoscopes first receives single-band spectral images synchronously acquired by multiple sensors at the endoscope. After lens distortion correction, these images are segmented into multiple small spectral image blocks according to a preset overlap rate. Subsequently, cross-spectral registration is performed to match and calibrate image blocks of different bands. Then, multi-level cross-fusion is performed based on spectral characteristic weights to obtain a two-dimensional spectral fusion image. On this basis, parallax calculation is performed on the visible light images transmitted from the left and right eye cameras to generate a 3D point cloud coordinate set for the endoscope. Then, back projection and sub-pixel mapping algorithms are used to map the 2D spectral fusion image onto the 3D point cloud coordinates to obtain a 3D spectral point cloud. Finally, Poisson fusion is performed on the 3D spectral point cloud, and inter-block stitching is performed to generate an interactive 3D navigation model. This process improves the fusion effect of the spectral images and enhances the accuracy of 3D reconstruction.
[0007] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a flowchart illustrating a spectral image processing method for a three-dimensional endoscope in one embodiment.
[0010] Figure 2 This is a schematic flowchart of lens distortion correction in a spectral image processing method for a three-dimensional endoscope, as shown in one embodiment. Detailed Implementation
[0011] This application provides a spectral image processing method for three-dimensional endoscopes, which solves the technical problem that insufficient spectral image fusion accuracy in existing three-dimensional endoscope image processing methods leads to loss of image details and inaccurate three-dimensional reconstruction.
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0013] It should be noted that the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such process, method, product, or device.
[0014] Examples, such as Figure 1 As shown, this application provides a spectral image processing method for a three-dimensional endoscope, the method comprising:
[0015] It receives multiple single-band spectral images synchronously acquired and transmitted from multiple sensors at the endoscope.
[0016] In this embodiment, the endoscopic device includes multiple spectral sensors capable of capturing images in different spectral bands (e.g., visible light, infrared light, etc.), with each sensor responsible for acquiring a spectral image in a specific band. During acquisition, each sensor performs acquisition operations synchronously to ensure the spatial consistency of spectral images in different bands, avoiding image distortion or misalignment due to time differences, thereby improving the accuracy of image registration and fusion. By summarizing the acquired spectral images, multiple single-band spectral images can be obtained. These single-band spectral images will undergo distortion correction, registration, and fusion in subsequent steps, ensuring accurate reconstruction of the final image in both spatial and spectral dimensions, providing high-quality data support for subsequent 3D reconstruction.
[0017] After lens distortion correction is performed on the multiple single-band spectral images, the multiple single-band spectral images are divided into multiple groups of single-spectral image blocks based on a preset inter-block overlap rate.
[0018] In one embodiment, after obtaining multiple single-band spectral images, radial and tangential distortion are first applied to each spectral image to eliminate distortion effects caused by lens optical characteristics, resulting in multiple corrected spectra. These corrected spectra ensure that the geometry of the image conforms to reality, making the image content more accurate. Subsequently, based on a preset inter-block overlap rate (e.g., 30%), each distortion-corrected spectrum is segmented into multiple smaller image blocks. The size of these image blocks can be set to 256×256 pixels to balance computational efficiency and image quality. During segmentation, the inter-block overlap rate ensures that the edges of each image block overlap with adjacent blocks, thus avoiding boundary artifacts and information loss caused by image segmentation. The existence of overlapping areas ensures a more natural and seamless transition between image blocks in subsequent image registration and fusion processes, thereby avoiding image quality degradation caused by segmentation. After segmentation, multiple sets of single-spectral image blocks are obtained, providing more accurate image data for subsequent block-level cross-spectral registration, spectral image fusion, and 3D reconstruction processing steps, ensuring that a high-quality 2D spectral fusion image and 3D spectral point cloud are ultimately obtained.
[0019] Furthermore, such as Figure 2 As shown, the method further includes:
[0020] The lens distortion parameters of a three-dimensional endoscope are obtained using a multi-scale calibration plate, wherein the multi-scale calibration plate includes a multi-scale sub-pixel level corner point array; the radial distortion coefficient and tangential distortion coefficient are output by nonlinear optimization of the lens distortion parameters; and the reverse mapping distortion correction of the multiple single-band spectra is performed based on the radial distortion coefficient and tangential distortion coefficient to obtain multiple corrected spectra.
[0021] Preferably, when performing lens distortion correction, a multi-scale calibration board is first constructed. This calibration board has sub-pixel-level corner point arrays (such as checkerboard or dot arrays) of multiple scales distributed on it. These corner points are arranged in different sizes and densities to cover multiple regions from the image center to the edge, adapting to the imaging characteristics of the endoscope at different working distances. This effectively improves the robustness and accuracy of distortion parameter estimation across the entire imaging field of view. Then, images of the calibration board are acquired from multiple viewpoints (different poses and distances) using the endoscope to ensure coverage of the entire field of view and depth of field. Subsequently, corner detection algorithms (such as Harris and Shi-Tomasi) are applied to each image, combined with sub-pixel optimization (such as OpenCV's cornerSubPix) to obtain the precise location of the corner points in the image, establishing a preliminary mapping relationship between image coordinates and actual physical coordinates, which serves as the lens distortion parameters. Afterward, a lens distortion model based on a pinhole imaging model is established, which includes radial distortion (such as k1, k2, k3) and tangential distortion (such as p1, p2). Using a nonlinear least squares optimization algorithm (such as the Levenberg-Marquardt algorithm), the objective function is to minimize the error between the actual corner positions of the image and the theoretically mapped corner positions. The distortion coefficients and camera intrinsic parameters are jointly solved, outputting a complete set of distortion parameters including radial and tangential distortion coefficients. Then, the calibrated radial and tangential distortion coefficients are used to perform inverse mapping on multiple single-band spectral images acquired by the endoscope. Specifically, for each pixel, its true position in the distortion-free image is calculated based on the distortion model, and the single-band spectrum is resampled using an interpolation method (such as bilinear interpolation) and filled into the corrected image, thereby generating multiple corrected spectra. These corrected spectra are multiple single-band spectra with geometric distortion eliminated, providing an accurate input image basis for subsequent image segmentation, registration, fusion, and 3D reconstruction, ensuring the imaging accuracy of the 3D endoscope in complex structures.
[0022] Furthermore, this application provides a method for segmenting the plurality of single-band spectral images into multiple groups of single-spectral image blocks based on a preset inter-block overlap rate, the method further comprising:
[0023] Predefined block size and inter-block overlap rate; after spatially aligning the multiple corrected spectra, the multiple corrected spectra are simultaneously segmented using a sliding window mechanism according to the block size and inter-block overlap rate to obtain the multiple sets of monospectral image blocks; wherein, each set of monospectral image blocks is assigned a block spatial index.
[0024] Optionally, the image block size and inter-block overlap rate are preset first. The block size refers to the pixel size of each image block, usually set to 256×256 pixels. The inter-block overlap rate refers to the proportion of overlap between adjacent blocks in the horizontal and vertical directions, usually set to 30%. Then, using the corrected spectrum of a certain band as a reference, the corrected spectra of other bands are aligned with this reference image through affine transformation, perspective transformation (or more complex nonlinear registration). Then, SIFT, SURF, or gray-level similarity-based methods are used to find points at the same positions in the corrected spectra of different bands, and transformation relationships (such as translation, rotation, and scaling) are calculated using these points. After obtaining the transformation relationships, these transformation relationships are applied to the corrected spectra, and all corrected spectra are transformed to the position coordinate system of the reference image through interpolation (such as bilinear interpolation) and resampling. Subsequently, based on the set block size and inter-block overlap rate, each calibration spectrogram is synchronously segmented using a sliding window mechanism. The sliding window starts from the upper left corner of the calibration spectrogram and slides gradually in both horizontal and vertical directions. The sliding step size is determined by both the block size and the overlap rate. For example, if the set block size is 256 pixels and the overlap rate is 30%, then the sliding step size is 256 × (1 - 0.3) = 179 pixels. After segmentation according to the sliding step size, multiple sets of monospectral image blocks with overlapping regions can be obtained. For each monospectral image block, a unique block spatial index is assigned. This block spatial index is used to identify the coordinate range of the monospectral image block in the calibration spectrogram, including its starting position of the upper left corner in the image coordinate system and the block size. For example, the spatial index of a monospectral image block can be represented as index ID = (x, y_start, 256, 256), which is used to accurately locate the position of the image block and its corresponding region in the subsequent registration and fusion process, improving the accuracy and efficiency of cross-spectral processing.
[0025] Furthermore, this application provides that the inter-block overlap rate is 20% to 40%.
[0026] Optionally, the preset inter-block overlap rate ranges from 20% to 40%, and is usually set to 30% to ensure that sufficient information redundancy is retained in the edge area of the image block, thereby effectively alleviating the artifact problem caused by segmentation at the image boundary and improving the stitching and fusion quality between image blocks.
[0027] After performing block-level cross-spectral registration on the multiple sets of monospectral image blocks, multi-level cross-optimization fusion of the multiple sets of monospectral image blocks is performed based on predefined spectral characteristic weights to obtain a two-dimensional spectral fusion image.
[0028] In one embodiment, to achieve high-quality spectral image fusion, the spatial index of each group of single-spectral image blocks is first used to register image blocks at the same location in different bands according to their corresponding positions in the image. Then, image blocks of different spectral channels are precisely aligned based on curvature fitting, thereby eliminating spatial misalignment caused by differences in band imaging. Subsequently, multi-level cross-optimization fusion is performed on the registered image blocks based on predefined spectral characteristic weights. These spectral characteristic weights refer to the fusion priority or importance assigned to different spectral bands during the fusion process. These weights can be set through user interaction or determined based on the spectral response characteristics of different tissues in specific bands. During the multi-level cross-optimization fusion process, pixel-level weighted averaging and edge enhancement are used to fuse single-spectral registered blocks of different bands, achieving smooth transitions between image blocks and avoiding edge breaks or brightness abrupt changes in the fused image. The final output two-dimensional spectral fusion image not only has high spatial consistency but also retains the significant features of each band in the spectral dimension, providing a high-quality image foundation for subsequent three-dimensional mapping and navigation modeling.
[0029] Furthermore, this application provides a method for block-level cross-spectral registration of the multiple sets of monospectral image blocks, the method comprising:
[0030] Based on the block spatial index, cross-spectral image block registration is performed on the multiple sets of monospectral image blocks to obtain multiple sets of cross-spectral image blocks; curvature registration is performed on the multiple sets of cross-spectral image blocks to obtain multiple sets of geometrically aligned cross-spectral blocks; the multiple sets of geometrically aligned cross-spectral blocks are restored to obtain multiple sets of monospectral registration blocks.
[0031] Optionally, to achieve high-precision alignment between different monospectral image patches, spatial indexing information is used to locate all monospectral image patches. That is, image patches from different spectral bands are acquired from multiple sets of monospectral image patches at the same spatial location, forming multiple sets of transspectral image patches with spatial correspondence. Each set of transspectral image patches contains image content at the same location but in different bands, providing a data foundation for subsequent registration. Then, for each set of transspectral image patches, SIFT (Scale Invariant Feature Transform) or similar feature extraction algorithms are used to extract structurally stable key points from each image patch. Bidirectional feature matching is then performed, and a closed-loop verification mechanism is used to eliminate erroneous matches, resulting in a set of reliable transspectral feature point pairs. Afterward, preliminary spatial alignment is performed using the matched feature point pairs. Based on this, the curvature distribution of each image patch in the local region is obtained by fitting a surface, the local radius of curvature is calculated, and then the transspectral image patches are geometrically aligned according to the calculation results, resulting in multiple sets of geometrically aligned transspectral patches. Then, for the multiple sets of geometrically aligned cross-spectral blocks that have completed curvature registration, the image content of each set of cross-spectral blocks is extracted band by band. That is, the spectral blocks of each band in each set of cross-spectral image blocks are extracted independently and labeled as monospectral image blocks under the corresponding band. The geometric transformation parameters applied to these monospectral image blocks in the cross-spectral registration process are retained to ensure that the accurate alignment relationship is maintained in the subsequent fusion process. Finally, the monospectral image blocks of each band after splitting are reorganized into independent sets according to spatial index, forming multiple sets of spatially aligned monospectral registration blocks. Each set of monospectral registration blocks contains image blocks of the corresponding band in different image frames at the same location and scale. These image blocks can be directly used for subsequent spectral fusion processing to ensure high-precision image integration.
[0032] Furthermore, this application provides a method for curvature registration of the multiple sets of transspectral image blocks to obtain multiple sets of geometrically aligned transspectral blocks, the method further comprising:
[0033] The first group of transspectral image patches is enumerated into multiple groups of mapped image patches. Based on improved SIFT feature extraction, bidirectional matching closed-loop verification of the multiple groups of mapped image patches is performed to filter out multiple groups of transspectral feature point pairs. The multiple groups of transspectral feature point pairs are spatially aligned to filter out the first group of transspectral feature points. Based on the first group of transspectral feature points, a surface is fitted to obtain the first local radius of curvature, and the first group of transspectral image patches is registered according to the first local radius of curvature to obtain the first group of geometrically aligned transspectral patches. The multiple groups of transspectral image patches are registered by analogy to obtain the multiple groups of geometrically aligned transspectral patches.
[0034] Optionally, in the first set of acquired transspectral image patches, one band is selected as the reference image patch, and the remaining band image patches are selected as target image patches. The reference image patch is then combined with each target image patch to form multiple sets of mapped image patches for subsequent registration calculations. Subsequently, an improved SIFT feature point extraction algorithm is executed on each set of mapped image patches. This involves first performing local contrast normalization on each image patch to eliminate brightness / contrast differences between transspectral image patches and enhance feature stability. Then, the single-channel image patch is expanded into a multi-channel patch (e.g., RGB+NIR four channels), and the gradient magnitude and direction are calculated for each channel. Weighted synthesis is then used to enhance feature discriminative power. Next, a Gaussian pyramid is used to construct a multi-scale image representation, extracting features for each image patch at multiple scales to enhance the ability to capture different structural details. Extreme points are detected in the scale space as candidate feature points, and a principal direction is assigned to each candidate feature point. Finally, a rotation-invariant local image gradient histogram is constructed to generate a descriptor. After obtaining the descriptors, a bidirectional matching closed-loop verification process is initiated. This involves calculating the Euclidean distance between each SIFT descriptor in the reference image block and all descriptors in the target image block, selecting the one with the smallest distance as its initial matching point, and then reversing this process, starting from each descriptor in the target image block and searching for the nearest neighbor matching point in the reference image block. After both forward and reverse matching are completed, consistency verification is performed on both sets of matching pairs. Only the nearest neighbor matching results are retained, i.e., the coordinate position error between the forward and reverse matching points does not exceed a threshold (e.g., 2 pixels), to eliminate a large number of mismatches and false features, resulting in multiple sets of cross-spectral feature point pairs. Then, for the obtained multiple sets of cross-spectral feature point pairs, preliminary spatial alignment is performed using coarse registration (e.g., based on affine transformation), and the reprojection error between each feature point and other points in its neighborhood is calculated. Feature point pairs with errors greater than the threshold are eliminated, thereby extracting the first set of high-quality cross-spectral feature points. Furthermore, utilizing the spatial distribution of the first set of transspectral feature points, a quadratic surface model is selected to fit these points, constructing a local structural surface in the region. The surface parameters are then solved using the least squares method, minimizing the sum of squared perpendicular distances from all feature points to the fitted surface, thus obtaining the best-fitting surface describing the local deformation trend. Based on the first and second-order partial derivatives of the fitted surface, differential geometry formulas are used to calculate the principal curvature of the region, and the first local radius of curvature is further solved. This first local radius of curvature reflects the degree of local geometric deformation of the image patch. After obtaining the first local radius of curvature, different image registration strategies are selected based on the radius's size. If the curvature is small (less than a threshold, flat region), affine transformation is used; if the curvature is large (greater than or equal to a threshold, deformed region), local non-rigid transformations (such as B-spline free deformation) are used for higher-order deformation compensation.Taking affine transformation as an example, the objective function is curvature radius-weighted least squares. The transformation parameters are iteratively solved using the least squares method or gradient descent method until convergence. Then, using the reference image block in the first group of transspectral image blocks as a baseline, the obtained transformation model is applied to the target image block in the first group of transspectral image blocks for bilinear interpolation or B-spline interpolation, generating aligned transspectral blocks as the first group of geometrically aligned transspectral blocks. Finally, for all remaining transspectral image block combinations, the above steps are repeated, from enumeration and combination, feature extraction, loop closure verification, surface fitting to curvature registration, completing the registration process group by group, ultimately obtaining the geometric alignment results of all image blocks, i.e., multiple groups of geometrically aligned transspectral blocks. In summary, the curvature-guided transspectral image block registration process described above not only improves the registration accuracy but also enhances the adaptability to structural changes between different spectral bands, providing a high-quality alignment foundation for subsequent spectral fusion and 3D reconstruction.
[0035] Furthermore, this application provides a method for multi-level cross-fusion of the multiple sets of single-spectral image blocks based on predefined spectral characteristic weights to obtain a two-dimensional spectral fusion image. The method further includes:
[0036] The spectral characteristic weights are obtained interactively; with the spectral characteristic weights as constraints, multi-level cross-fusion of the multiple sets of single-spectral registration blocks is performed according to the block spatial index mapping to output an initial edge enhancement fusion block; the weighted average value of the overlapping areas of the initial edge enhancement fusion block is taken for seam fusion to output the two-dimensional spectral fusion image.
[0037] Optionally, to achieve a two-dimensional spectral fusion image with edge sharpness and spectral fidelity, the relative importance weights of different spectral bands are first obtained as spectral characteristic weights through a user interface or preset task parameters. Then, these spectral characteristic weights are mapped as constraint information to corresponding groups of single-spectral registration blocks for multi-level cross-fusion. Specifically, based on the mapping relationship between the block spatial index and multiple groups of single-spectral registration blocks, single-spectral registration blocks belonging to the same location are grouped as a fusion candidate group. For pixels at the same location within the fusion candidate group, weighted fusion is performed according to the spectral characteristic weights. During the fusion process, to enhance image edge details and structural features, an edge enhancement factor (such as Sobel) is introduced to improve the edge response of the fusion result, giving it stronger edge preservation capabilities in boundary regions. The enhanced pixel values are then written into a temporary fusion image block, and spatial consistency control of neighboring pixels is achieved through local mean adjustment, further smoothing the fusion result and suppressing noise interference, thus obtaining the final initial edge-enhanced fusion block. Subsequently, due to the 20%–40% overlap between image patches, to avoid seams or brightness discontinuities in the fused image, the overlapping areas between all image patches are identified based on the block spatial index. For each overlapping pixel, a weighted average is calculated based on its distance from the center point to obtain the final pixel value of each overlapping pixel. This eliminates seam artifacts such as brightness jumps and texture breaks caused by image patch stitching, ensuring the overall coherence and visual consistency of the fused image. Finally, all the initial edge-enhanced fusion blocks are reconstructed into a complete image using spatial indexing, outputting a high-quality two-dimensional spectral fusion image. This two-dimensional spectral fusion image possesses both spectral representation capabilities of multi-band information and spatial structure consistency, providing a high-precision image foundation for subsequent tasks such as 3D reconstruction, object detection, and visual navigation.
[0038] By calculating the parallax of binocular visible light images transmitted from the left and right cameras, an endoscopic 3D point cloud coordinate set is constructed.
[0039] In one embodiment, to achieve accurate mapping of two-dimensional spectral fusion images to three-dimensional space, disparity calculation is first performed on the binocular visible light image pairs transmitted from the left and right eye cameras of the endoscope. The left and right eye cameras are two fixed-position, parameter-calibrated cameras equipped with the endoscope, which simultaneously acquire visible light images of the same scene from different left and right perspectives. During disparity calculation, a high-precision stereo matching algorithm, such as semi-global stereo matching (SGM), is employed. A matching search is performed on each pixel within the entire image range to obtain its corresponding disparity value and generate a grayscale disparity map. Then, combined with pre-calibrated camera intrinsic parameters (including focal length and principal point position) and the baseline distance between the two cameras, the three-dimensional point coordinates of each image pixel are calculated using the intrinsic parameter matrix. This constructs a complete endoscopic three-dimensional point cloud coordinate set. This endoscopic three-dimensional point cloud coordinate set can realistically reproduce the three-dimensional geometric structure of the scene, providing the necessary spatial foundation and structural support for subsequent spectral information mapping, three-dimensional spectral reconstruction, and interactive navigation model construction.
[0040] Furthermore, this application provides a method for constructing an endoscopic 3D point cloud coordinate set by calculating the disparity of binocular visible light images transmitted from the left and right eye cameras. The method further includes:
[0041] Semi-global stereo matching is performed on the binocular visible light images to output an optimized disparity map; global depth value transformation is performed on the optimized disparity map to output a registration depth map; based on the calibrated intrinsic parameters of the 3D endoscope camera, pixel-by-pixel 3D spatial coordinate transformation of the registration depth map is performed to output the endoscopic 3D point cloud coordinate set.
[0042] Preferably, after simultaneously acquiring binocular visible light images using a binocular camera, a semi-global stereo matching (SGM) algorithm is first employed. For each pixel in the left-eye image of the binocular visible light image, a corresponding matching pixel is searched in the right-eye image along the scan line direction. An initial matching cost value between pixels is calculated using a matching cost function based on gray-level difference, SAD (sum of absolute differences), or Census transform. To improve matching accuracy, SGM performs an energy minimization process in multiple directions (typically 8 or 16), achieving a trade-off between global consistency and local smoothness by aggregating matching costs from different paths. After cost aggregation, the disparity value with the minimum total cost is selected for each pixel to construct a preliminary disparity map. To improve the quality of the initial disparity map, post-processing operations are performed, including left-right disparity consistency detection (comparing the differences between the left and right disparity maps and removing occluded or mismatched regions), median filtering (eliminating isolated mismatched points), and hole filling (filling invalid regions using neighborhood interpolation). This results in an optimized disparity map, which possesses strong edge-preserving capabilities and high-density disparity information, improving the accuracy of image registration and spatial structure extraction in complex endoscopic scenes. Subsequently, using the camera's baseline length (the physical distance between the left and right cameras) and focal length, the disparity values are converted into depth values in the real scene. The calculation method involves multiplying the camera's focal length by the baseline length and then dividing by the disparity value. This calculation process is performed pixel-by-pixel, mapping the entire disparity map to a registration depth map to reflect the real spatial distance corresponding to each pixel, thus generating a registration depth map. Next, using the calibration data of the 3D endoscope (i.e., camera intrinsic parameters), each pixel in the registered depth map is back-projected into 3D space. During this process, the pixel's x-coordinate is subtracted from the camera's principal point x-coordinate, the difference is multiplied by the depth value, and then divided by the camera's focal length on the x-axis to obtain the pixel's x-coordinate in 3D space. Similarly, the pixel's y-coordinate is subtracted from the camera's principal point y-coordinate, the difference is multiplied by the depth value, and then divided by the camera's focal length on the y-axis to obtain the pixel's y-coordinate in 3D space. The pixel's depth value is used as its z-coordinate in 3D space. After performing the above processing on each pixel, a dense set of 3D coordinate points is obtained, forming a complete endoscopic 3D point cloud coordinate set. This endoscopic 3D point cloud coordinate set not only contains the spatial morphology of the target scene but also provides a spatial basis for subsequent 3D mapping of 2D spectral images, effectively supporting core functions such as subsequent spectral mapping, navigation modeling, and visualization analysis.
[0043] The two-dimensional spectral fusion image space is mapped to the coordinate set of the endoscopic three-dimensional point cloud based on the back projection subpixel mapping algorithm to obtain the three-dimensional spectral point cloud.
[0044] In one embodiment, after obtaining the coordinate set of the endoscopic 3D point cloud, in order to ensure that the 2D spectral fusion image and the 3D spatial structure can correspond accurately, an algorithm based on back projection and subpixel mapping is used to spatially map the fused 2D spectral image to the coordinate set of the endoscopic 3D point cloud, thereby constructing a 3D spectral point cloud with complete spatial structure and spectral information, providing a high-quality data foundation for subsequent inter-block stitching, image visualization and intelligent diagnostic analysis.
[0045] Furthermore, this application provides a method for mapping the two-dimensional spectral fusion image space to the endoscopic three-dimensional point cloud coordinate set based on a back-projection subpixel mapping algorithm to obtain a three-dimensional spectral point cloud. The method further includes:
[0046] The coordinate set of the endoscopic 3D point cloud is traversed, and the corresponding 2D pixel coordinate set is calculated by back-projection of the camera intrinsic parameters. Bilinear interpolation is used to extract the sub-pixel level spectral feature value set corresponding to the 2D pixel coordinate set in the 2D spectral fusion image mapping. According to the mapping relationship between the 2D pixel coordinate set and the endoscopic 3D point cloud coordinate set, the sub-pixel level spectral feature value set is bound to the endoscopic 3D point cloud coordinate set to obtain the 3D spectral point cloud.
[0047] Optionally, firstly, each point in the endoscopic 3D point cloud coordinate set is traversed. During this traversal, camera intrinsic parameters obtained from endoscopic calibration, such as focal length and principal point coordinates, are used to perform back-projection calculations on each 3D point using a pinhole camera model. This yields the corresponding pixel position of that point in the 2D image. The x-coordinate of the 2D pixel coordinate is obtained by multiplying the 3D point cloud's x-coordinate by the x-axis focal length, dividing by the 3D point cloud's z-coordinate, and adding the quotient to the principal point's x-coordinate. Similarly, the y-coordinate is obtained by multiplying the 3D point cloud's y-coordinate by the y-axis focal length, dividing by the 3D point cloud's z-coordinate, and adding the quotient to the principal point's y-coordinate. By summing the calculated 2D pixel coordinates, a 2D pixel coordinate set corresponding to the endoscopic 3D point cloud coordinate set can be obtained, representing the sub-pixel position of the 3D point in the spectral image. Subsequently, since the 2D spectral fusion image is composed of discrete pixels, and the calculated 2D pixel coordinates usually do not fall on integer pixels, a sub-pixel interpolation method is needed to extract its spectral values. Taking bilinear interpolation as an example, the four nearest integer pixels around the 2D pixel coordinate are determined, and the relative distance between the 2D pixel coordinate and these four points is calculated as weights. Then, the feature values of these four integer pixels are fused using an interpolation formula to obtain a sub-pixel spectral feature value set of the 2D pixel coordinate set. Afterwards, based on the mapping relationship between the 2D pixel coordinate set and the endoscopic 3D point cloud coordinates, the interpolated sub-pixel spectral feature values are bound to their corresponding 3D coordinate points to form complete 3D spectral points. After performing the above process on all 3D points, a 3D spectral point cloud containing rich spectral information is finally generated. This 3D spectral point cloud not only reflects the spatial structure of the scene but also expresses the spectral characteristics of different regions, providing crucial data support for subsequent medical tissue identification, anomaly detection, and navigation visualization.
[0048] The three-dimensional spectral point cloud is subjected to Poisson fusion-driven inter-block stitching and fusion to output an interactive three-dimensional navigation model.
[0049] In one embodiment, after obtaining the 3D spectral point cloud, a Poisson fusion-driven inter-block stitching fusion operation is performed on the 3D spectral point cloud. Specifically, overlapping regions of all adjacent point clouds are identified based on the block spatial index, and the normal vector field and spectral gradient field of points within each overlapping region are calculated. The normal vector field describes the geometric orientation of the local surface of the point cloud, while the spectral gradient field expresses the characteristic variation trend of each point in multiple spectral dimensions. During fusion, the goal is to minimize the normal vector deviation and spectral gradient difference between point clouds in the overlapping regions. The positions of points in the overlapping regions and their corresponding spectral attributes are adjusted to ensure a natural transition between blocks without obvious stitching marks, avoiding problems such as structural breaks, spectral abrupt changes, or texture discontinuities caused by image segmentation. After fusion, all spectral point cloud blocks are integrated into a complete three-dimensional data volume with a unified spatial coordinate system, and further constructed into an interactive three-dimensional navigation model. This model allows users to rotate, scale, and cross-section the three-dimensional tissue structure at any angle, and view the spectral feature information of any point in real time. This meets the high requirements for visualization accuracy and operational flexibility in application scenarios such as clinical diagnosis, intraoperative navigation, or precision detection, thereby significantly improving the practicality and intelligence of endoscopic images in actual operation.
[0050] Furthermore, this application provides a method for performing Poisson fusion-driven inter-block stitching fusion on the three-dimensional spectral point cloud to output an interactive three-dimensional navigation model. The method further includes:
[0051] Based on the block spatial index, the overlapping area of cloud blocks is located in the three-dimensional spectral point cloud index; the point cloud normal vector field and spectral gradient field of the overlapping area of cloud blocks are calculated; based on the point cloud normal vector field and spectral gradient field, the seams are eliminated by energy minimization Poisson fusion, and the interactive three-dimensional navigation model is output, wherein the interactive three-dimensional navigation model integrates the real-time cutting function of arbitrary plane.
[0052] Preferably, a block-based spatial index is used to traverse the 3D spectral point cloud block by block to identify overlapping regions. Within each overlapping region, for every point cloud sub-block, principal component analysis (PCA) or surface fitting based on local neighborhood is used to calculate the normal vector of that point. Taking PCA as an example, the k-nearest neighbors of each overlapping point are searched to construct a covariance matrix. Then, the eigenvector corresponding to the minimum eigenvalue is calculated as the normal vector to represent its surface orientation and local geometry. These normal vectors constitute the point cloud normal vector field of the entire overlapping region. In addition, the rate of change of each spectral band within its spatial neighborhood is used to obtain gradient information reflecting the spectral change trend, constructing a spectral gradient field to describe the transition characteristics between adjacent blocks in the spectral dimension. Subsequently, an energy function is constructed, which includes a geometric consistency term and a spectral smoothing term. The geometric consistency term is used to ensure a smooth transition of the point cloud normal vector field at the boundary within the overlapping region, while the spectral smoothing term is used to minimize the spectral gradient change within the overlapping region, improving the naturalness of spectral fusion. Next, the energy function is transformed into a system of linear Poisson equations. Using the point's position and spectral values as variables, a sparse matrix representation is constructed. Then, efficient numerical solvers such as multigrid methods or conjugate gradients are employed to obtain new adjusted point coordinate values and spectral fusion values. The optimization results are then updated in the 3D spectral point cloud, achieving seamless stitching and continuous spectral fusion between multiple point cloud blocks. After fusion, the entire 3D spectral point cloud is uniformly rendered and its data structure is encapsulated to generate a complete interactive 3D navigation model. This interactive 3D navigation model features real-time interactive operation, real-time sectionalization of arbitrary planes, and multi-channel spectral visualization, which can assist in tissue identification and lesion analysis, ensuring the overall consistency, continuity, and diagnostic usability of the 3D model in terms of geometric structure and spectral representation.
[0053] In summary, the embodiments of this application have at least the following technical effects:
[0054] This embodiment first receives multiple single-band spectral images synchronously acquired and transmitted from multiple sensors at the endoscope. Then, after lens distortion correction, the multiple single-band spectral images are segmented into multiple groups of single-spectral image blocks based on a preset inter-block overlap rate. Next, block-level cross-spectral registration is performed on the multiple groups of single-spectral image blocks, and multi-level cross-optimization fusion is conducted based on predefined spectral characteristic weights to obtain a two-dimensional spectral fusion image. Further, disparity calculation is performed on the binocular visible light image pairs transmitted from the left and right eye cameras to construct an endoscopic three-dimensional point cloud coordinate set. Then, the two-dimensional spectral fusion image space is mapped to the endoscopic three-dimensional point cloud coordinate set based on a back-projection subpixel mapping algorithm to obtain a three-dimensional spectral point cloud. Finally, Poisson fusion-driven inter-block stitching fusion is performed on the three-dimensional spectral point cloud to output an interactive three-dimensional navigation model. These technologies collectively address the technical problems of insufficient spectral image fusion accuracy in existing 3D endoscopic image processing methods, which leads to loss of image details and inaccurate 3D reconstruction. They achieve the technical effect of high-precision registration and reconstruction of 3D spectral point clouds through multimodal image collaborative correction and optimization fusion, thereby improving the accuracy and real-time performance of navigation in endoscopes.
[0055] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0056] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0057] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.
Claims
1. A spectral image processing method for three-dimensional endoscopes, characterized in that, The method includes: Receive multiple single-band spectral images synchronously acquired and transmitted from multiple sensors at the endoscope end; After lens distortion correction is performed on the multiple single-band spectral images, the multiple single-band spectral images are divided into multiple groups of single-spectral image blocks based on a preset inter-block overlap rate. After performing block-level cross-spectral registration on the multiple sets of single-spectral image blocks, multi-level cross-optimization fusion of the multiple sets of single-spectral image blocks is performed based on predefined spectral characteristic weights to obtain a two-dimensional spectral fusion image. By calculating the parallax of the binocular visible light images transmitted from the left and right cameras, an endoscopic 3D point cloud coordinate set is constructed. The two-dimensional spectral fusion image space is mapped to the coordinate set of the endoscopic three-dimensional point cloud based on the back projection subpixel mapping algorithm to obtain the three-dimensional spectral point cloud. The three-dimensional spectral point cloud is subjected to Poisson fusion-driven inter-block stitching and fusion to output an interactive three-dimensional navigation model.
2. The spectral image processing method for a three-dimensional endoscope as described in claim 1, characterized in that, The method further includes: Lens distortion parameters of a three-dimensional endoscope are obtained using a multi-scale calibration plate, wherein the multi-scale calibration plate includes a multi-scale sub-pixel level corner point array; By performing nonlinear optimization on the lens distortion parameters, the radial distortion coefficient and the tangential distortion coefficient are output. Based on the radial and tangential distortion coefficients, reverse mapping distortion correction is performed on the multiple single-band spectra to obtain multiple corrected spectra.
3. The spectral image processing method for a three-dimensional endoscope as described in claim 2, characterized in that, The method further includes segmenting the multiple single-band spectral images into multiple groups of single-spectral image blocks based on a preset inter-block overlap rate. Predefined block size and inter-block overlap ratio; After spatially aligning the multiple corrected spectra, the multiple sets of monospectral image blocks are obtained by simultaneously segmenting the multiple corrected spectra based on the block size and inter-block overlap rate using a sliding window mechanism. Each group of monospectral image blocks is assigned a block space index.
4. The spectral image processing method for a three-dimensional endoscope as described in claim 3, characterized in that, The method for performing block-level cross-spectral registration on the multiple sets of single-spectral image blocks includes: Based on the block spatial index, cross-spectral image block registration and acquisition are performed on the multiple sets of monospectral image blocks to obtain multiple sets of cross-spectral image blocks; Curvature registration is performed on the multiple sets of transspectral image blocks to obtain multiple sets of geometrically aligned transspectral blocks; By restoring the multiple sets of geometrically aligned transspectral blocks, multiple sets of single-spectral registration blocks are obtained.
5. The spectral image processing method for a three-dimensional endoscope as described in claim 4, characterized in that, Based on predefined spectral characteristic weights, multi-level cross-fusion of the multiple sets of single-spectral image blocks is performed to obtain a two-dimensional spectral fusion image. The method further includes: The weights of the spectral characteristics are obtained interactively; Using the spectral characteristic weights as constraints, multi-level cross-fusion of the multiple sets of single-spectral registration blocks is performed based on the block spatial index mapping to output an initial edge-enhanced fusion block; The weighted average value of the overlapping areas of the initial edge enhancement fusion block is used for seam fusion, and the two-dimensional spectral fusion image is output.
6. The spectral image processing method for a three-dimensional endoscope as described in claim 5, characterized in that, The method further includes: performing curvature registration on the multiple sets of transspectral image patches to obtain multiple sets of geometrically aligned transspectral patches; and performing curvature registration on the multiple sets of transspectral image patches to obtain multiple sets of geometrically aligned transspectral patches. The first group of transspectral image blocks is enumerated into multiple groups of mapped image blocks; Based on improved SIFT feature extraction, bidirectional matching closed-loop verification of the multiple sets of mapped image blocks is performed to filter out multiple sets of cross-spectral feature point pairs; The first set of transspectral feature points is obtained by spatially aligning the multiple sets of transspectral feature point pairs; Based on the first set of cross-spectral feature points, a surface is fitted to obtain a first local radius of curvature. Based on the first local radius of curvature, the first set of cross-spectral image blocks is registered to obtain a first set of geometrically aligned cross-spectral blocks. By analogy, the registration of the multiple sets of transspectral image blocks is performed to obtain the multiple sets of geometrically aligned transspectral blocks.
7. The spectral image processing method for a three-dimensional endoscope as described in claim 1, characterized in that, The method further includes constructing an endoscopic 3D point cloud coordinate set by calculating the disparity between binocular visible light images transmitted from the left and right eye cameras. Perform semi-global stereo matching on the binocular visible light images and output an optimized disparity map; By performing global depth value transformation on the optimized disparity map, a registration depth map is output; Based on the intrinsic parameters of the 3D endoscope camera obtained from calibration, a pixel-by-pixel 3D spatial coordinate transformation is performed on the registration depth map to output the endoscopic 3D point cloud coordinate set.
8. The spectral image processing method for a three-dimensional endoscope as described in claim 7, characterized in that, The method further includes mapping the two-dimensional spectral fusion image space to the endoscopic three-dimensional point cloud coordinate set based on the back projection subpixel mapping algorithm to obtain the three-dimensional spectral point cloud. Traverse the endoscopic 3D point cloud coordinate set and calculate the corresponding 2D pixel coordinate set by back-projection of the camera intrinsic parameters; Bilinear interpolation is used to extract a sub-pixel level spectral feature value set corresponding to the two-dimensional pixel coordinate set from the two-dimensional spectral fusion image mapping; Based on the mapping relationship between the two-dimensional pixel coordinate set and the endoscopic three-dimensional point cloud coordinate set, the sub-pixel level spectral feature value set is bound to the endoscopic three-dimensional point cloud coordinate set to obtain the three-dimensional spectral point cloud.
9. The spectral image processing method for a three-dimensional endoscope as described in claim 3, characterized in that, The inter-block overlap rate is 20% to 40%.
10. The spectral image processing method for a three-dimensional endoscope as described in claim 1, characterized in that, The method further includes performing Poisson fusion-driven inter-block stitching fusion on the three-dimensional spectral point cloud to output an interactive three-dimensional navigation model. Based on the block spatial index, the overlapping area of cloud blocks is located in the three-dimensional spectral point cloud index; Calculate the point cloud normal field and spectral gradient field of the overlapping cloud region; Based on the point cloud normal vector field and spectral gradient field, seams are eliminated by energy minimization Poisson fusion, and the interactive 3D navigation model is output. The interactive 3D navigation model integrates real-time sectionalizing function of arbitrary plane.
Citation Information
Cited By
Multi-view hyperspectral three-dimensional reconstruction and spectral mapping method
CN122391519A
Multi-view hyperspectral three-dimensional reconstruction and spectral mapping method
CN122391519B