Geometric consistency of endoscope and instrument anchoring three-dimensional reconstruction and polyp measurement method

By projecting image features into a Riemannian manifold space during colonoscopy and using biopsy forceps as instrument anchors, combined with latent diffusion models and Bayesian regression, the accuracy and safety issues of polyp measurement under monocular endoscopy were solved, achieving sub-millimeter-level measurement accuracy and reliability.

CN122391183APending Publication Date: 2026-07-14SUZHOU ENTROPTONG INTELLIGENT TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU ENTROPTONG INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-05-15
Publication Date
2026-07-14

Smart Images

  • Figure CN122391183A_ABST
    Figure CN122391183A_ABST
Patent Text Reader

Abstract

The application discloses a geometric consistency and instrument anchoring three-dimensional reconstruction and polyp measurement method of an endoscope. The method first projects monocular image features to a Riemannian manifold space through a geometric consistency perception module, generates a high-fidelity relative depth map by using a manifold-constrained latent diffusion model, and maintains cross-domain structural consistency. Secondly, by using a scale estimation module based on instrument anchoring, biopsy forceps in a field of view are automatically identified and key points are extracted, a PnP solver and a Bayesian regression are used to solve six degrees of freedom of the instrument, and finally, geometric constraint equations are constructed based on the known physical size of the biopsy forceps, a global absolute scale factor is solved, and the relative depth map is converted into a metric three-dimensional point cloud. A few-sample linear calibration mechanism is introduced, and linear deviation caused by domain offset can be eliminated by using a small amount of clinical frames. According to the application, sub-millimeter polyp size measurement can be realized under monocular endoscopy by using conventional surgical instruments without additional hardware.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing technology, and in particular to a method for geometrical consistency and instrument anchoring three-dimensional reconstruction of endoscopes and polyp measurement. Background Technology

[0002] The prevention and early treatment of colorectal cancer (CRC) heavily rely on colonoscopy. In clinical practice, accurate morphological characterization of precancerous lesions (adenomas), particularly measuring their maximum diameter and volume density, is crucial for assessing malignancy risk and determining surgical options (such as EMR / ESD). For the classification of high-risk adenomas, clinical guidelines recommend controlling the size measurement error to within 1.2 mm as a reference standard for making informed clinical decisions. However, current clinical practice struggles to achieve this level of precision.

[0003] The root cause of this problem lies in the fact that most conventional white light endoscopes widely used in clinical practice are monocular devices. Their imaging principle involves projecting a three-dimensional intestinal lumen scene into a two-dimensional image, a process that loses depth information. Therefore, the true physical size of an object cannot be determined from a single image, a phenomenon technically known as "scale blurring." Furthermore, in the intestinal environment, issues are particularly difficult to resolve due to tissue surface reflections, mucus interference, and uneven illumination.

[0004] Existing technologies for addressing the above problems suffer from the following drawbacks, which significantly limit the accuracy and widespread applicability of their clinical applications:

[0005] First, it relies on the doctor's subjective estimation. Currently, the most common method is for doctors to visually inspect the lesion or compare it with surgical instruments (such as biopsy forceps), estimating the polyp size based on experience. This method is too subjective, and the judgments of different doctors vary greatly. Studies have shown that even experienced endoscopists often have an estimation error exceeding 20%. This qualitative rather than quantitative assessment method is insufficient to meet the sub-millimeter precision required for classifying high-risk adenomas.

[0006] Second, traditional geometric solutions are unstable. Although some studies have attempted to use surgical instruments (such as biopsy forceps) as references and employ algebraic geometry methods to solve pose, these methods are usually extremely sensitive to image segmentation noise. Theoretical analysis shows that under non-ideal observation conditions, these rigid geometric solutions often fail to reach the Cramér-Rao lower bound (CRB) for scale estimation, leading to unstable measurement results.

[0007] Third, deep learning-based methods suffer from severe domain shift. In recent years, while deep learning-based monocular depth estimation (MDE) models (such as Marigold and UniDepth) have performed excellently on natural scene images, they face severe "distribution mismatch" when directly applied to endoscopic images. Firstly, these models are mostly trained on natural internet images or synthetic datasets, while the real intestinal environment has unique high-frequency mucosal textures, fluid interference, and non-Lambertian highly reflective surfaces. This texture difference causes the model to fail to correctly perceive geometric structures when transferring from the "physical phantom domain" to the "real clinical domain," resulting in relative depth maps that often exhibit structural degradation or excessive smoothing. Secondly, existing MDE models typically output affine-invariant relative depths, i.e. Lacking absolute physical units, depth priors cannot be directly used to guide clinical surgical resections. Furthermore, due to differences in physical properties between domains, depth priors are often accompanied by non-random systematic scale shifts, which limit the measurement accuracy in zero-deployment scenarios.

[0008] Fourth, there is a lack of safety assessment for observation degradation scenarios. In monocular vision calculation, when the instrument is occluded or the posture is not ideal, the output measurement value often has extremely high uncertainty. Existing technologies lack a mechanism for identifying and intercepting such high-risk observation frames. Summary of the Invention

[0009] The purpose of this invention is to overcome the shortcomings of the prior art by providing a method for geometric consistency and instrument anchoring three-dimensional reconstruction of endoscopes and polyp measurement, which solves the problems of scale ambiguity, cross-domain structural distortion, systematic deviation of depth prior, and lack of risk assessment for observation degradation scenarios.

[0010] To achieve the above objectives, the technical solution adopted by this invention is: a method for three-dimensional reconstruction of endoscope geometric consistency and instrument anchoring, and for polyp measurement, the method comprising the following steps:

[0011] Step S1: Acquire video frame images under a monocular endoscope. and the intrinsic parameter matrix of the endoscope camera ;

[0012] Step S2, using the Geometric Consistency Awareness (GCP) module to process the image The image features are extracted from a pre-trained feature encoder and then projected onto a symmetric positive definite (SPD) Riemannian manifold space through a manifold projection layer to obtain the manifold feature embedding. ;

[0013] Step S3, utilizing the embedding of the manifold features The latent diffusion model (LDM) with conditions predicts the image through a manifold-constrained inverse stochastic differential equation (SDE) generation process. High-fidelity true depth map ;

[0014] Step S4, for the image Semantic segmentation and keypoint detection are performed to identify the biopsy forceps region and extract the two-dimensional keypoint coordinates of the biopsy forceps;

[0015] Step S5: Construct an instrument-anchored scale estimation model (IASD). Based on the coordinates of the two-dimensional key points and the known physical dimensions of the biopsy forceps, calculate the pose of the biopsy forceps relative to the camera, and calculate the global metric scale factor according to geometric projection constraints. ;

[0016] Step S6, the relative depth map described in step S3 Compared with the global metric scaling factor described in step S5 Merge to generate an absolute metric depth map Based on this, a three-dimensional point cloud is reconstructed. In 3D point cloud In this process, the polyp region is automatically located using a segmentation mask, and the Euclidean distance between any two points within the polyp region is calculated to measure the physical size of the polyp. The maximum value is then taken as the physical diameter of the polyp. .

[0017] Further, in step S2, the process of projecting image features onto a symmetric positive definite (SPD) Riemannian manifold space includes: converting the feature map output by the feature encoder... Construct as a regularized covariance matrix The formula is as follows:

[0018] in, For spatial sample number, It is the identity matrix. To prevent small positive numbers from exhibiting singularity; the manifold feature is embedded. Located in Riemannian manifold superior.

[0019] Furthermore, the Riemannian manifold space employs the affine invariant Riemannian metric (AIRM). During the model training phase, the Riemann contrastive loss function is used to align features from different domains to eliminate domain offset between synthetic data and real clinical data.

[0020] Furthermore, in step S3, the latent diffusion model includes a manifold-aware adaptive group normalization layer (MA-AdaGN) for embedding manifold features. Injection denoising process And using the inverse SDE solver with manifold constraints to perform reverse SDE sampling, noise is reduced. Gradually restore to depth .

[0021] Furthermore, in step S5, the known physical dimensions of the biopsy forceps include the length of the forceps head. The calculation of the global metric scaling factor The formula is:

[0022] in, For the camera's focal length, This represents the projected pixel length of the biopsy forceps in the image. The angle between the biopsy forceps axis and the camera optical axis. relative depth map The average relative depth value within the biopsy forceps area.

[0023] Furthermore, the included angle The results were obtained by combining the three-dimensional model of the biopsy forceps with the two-dimensional key points of the image using the PnP (Perspective-n-Point) solver.

[0024] Furthermore, the scale estimation model is implemented using a Bayesian regressor, whose estimated variance approximates the Cramér-Rao lower bound (CRB) for scale discovery under observable conditions, in order to achieve sub-millimeter-level measurement accuracy.

[0025] Furthermore, step S5 also includes an uncertainty gating mechanism: using a Bayesian regressor to estimate the biopsy clamp pose parameters. Posterior distribution uncertainty measure , recorded as And calculate the cognitive uncertainty index of the posterior distribution; if the cognitive uncertainty index is less than a preset threshold If the scale factor is calculated in step (5), then the scale factor is calculated. If the scale factor is greater than the preset threshold, the observation conditions are determined to be degraded, the scale factor is not calculated, and only the relative depth map is output.

[0026] Furthermore, step S5 also includes a scale factor linear calibration step: for the global metric scale factor of the initial solution. The existing systematic bias is addressed using a pre-learned linear correction function. After correction, the calibrated scale factor is obtained. Among them, the correction coefficient and The depth estimation error was obtained by randomly selecting a very small number of samples (k≤6) from non-homologous clinical sequences and performing least squares fitting, in order to eliminate the depth estimation error caused by the difference in distribution between the training domain and the clinical domain.

[0027] Furthermore, in step S6, the reconstruction of the three-dimensional point cloud The specific calculation is as follows: for each pixel coordinate in the image Its corresponding three-dimensional spatial coordinates for:

[0028] in, Homogeneous coordinates , For pixels The absolute depth at that location.

[0029] The beneficial effects of the above-mentioned structure of this invention are as follows: This invention provides a method for geometric consistency and instrument anchoring three-dimensional reconstruction and polyp measurement of endoscopes. This invention constrains features to Riemannian manifold space through a geometric consistency perception (GCP) module and aligns geometric structures using second-order statistics (covariance), significantly improving the structural fidelity of the model in complex clinical environments and solving the problem of cross-domain structural perception. This invention innovatively uses conventional biopsy forceps as "in-situ physical anchors" and recovers absolute scale from monocular video through the IASD module. Combined with a few-sample linear calibration mechanism, only 6 frames of reference data are needed to reduce the mean diameter error (MDE) from 3.18mm to 0.39mm, achieving sub-millimeter level clinical precision. This realizes sub-millimeter level quantization under monocular equipment and achieves accurate measurement without hardware modification. At the same time, the designed uncertainty gating mechanism can automatically identify high-risk frames and provide risk warnings, effectively preventing the output of misleading measurement results, improving surgical safety, and possessing high clinical-grade reliability. Attached Figure Description

[0030] The technical solution of the present invention will be further described below with reference to the accompanying drawings:

[0031] Figure 1 This is a schematic diagram of the overall architecture of the endoscope's geometric consistency and instrument anchoring three-dimensional reconstruction and polyp measurement method described in this invention.

[0032] Figure 2 This is a geometric derivation diagram of the instrument-anchored scale estimation (IASD) module described in this invention.

[0033] Figure 3 This is a flowchart illustrating the geometric consistency and instrument anchoring three-dimensional reconstruction and polyp measurement method of the endoscope described in this invention.

[0034] Figure 4This is a comparison chart of the depth estimation effects of the embodiments of the present invention and the prior art on clinical data, where the left side is a depth comparison chart and the right side is a surface normal vector comparison chart. Detailed Implementation

[0035] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0036] like Figures 1-4 As shown, a method for 3D reconstruction of endoscope geometric consistency and instrument anchoring, and for polyp measurement, is described, comprising the following steps:

[0037] Step S1: Acquire video frame images under a monocular endoscope. and the intrinsic parameter matrix of the endoscope camera ;

[0038] Step S2, using the Geometric Consistency Awareness (GCP) module to process the image The image features are extracted from a pre-trained feature encoder and then projected onto a symmetric positive definite (SPD) Riemannian manifold space through a manifold projection layer to obtain the manifold feature embedding. ;

[0039] Step S3, utilizing the embedding of the manifold features The latent diffusion model (LDM) with conditions predicts the image through a manifold-constrained inverse stochastic differential equation (SDE) generation process. High-fidelity true depth map ;

[0040] Step S4, for the image Semantic segmentation and keypoint detection are performed to identify the biopsy forceps region and extract the two-dimensional keypoint coordinates of the biopsy forceps;

[0041] Step S5: Construct an instrument-anchored scale estimation model (IASD). Based on the coordinates of the two-dimensional key points and the known physical dimensions of the biopsy forceps, calculate the pose of the biopsy forceps relative to the camera, and calculate the global metric scale factor according to geometric projection constraints. ;

[0042] Step S6, the relative depth map described in step S3 Compared with the global metric scaling factor described in step S5 Merge to generate an absolute metric depth map Based on this, a three-dimensional point cloud is reconstructed. In 3D point cloud In this process, the polyp region is automatically located using a segmentation mask, and the Euclidean distance between any two points within the polyp region is calculated to measure the physical size of the polyp. The maximum value is then taken as the physical diameter of the polyp. .

[0043] In step S2, the process of projecting image features onto a symmetric positive definite (SPD) Riemannian manifold space includes: converting the feature map output by the feature encoder... Construct as a regularized covariance matrix The formula is as follows:

[0044] in, For spatial sample number, It is the identity matrix. To prevent small positive numbers from exhibiting singularity; the manifold feature is embedded. Located in Riemannian manifold superior.

[0045] The Riemannian manifold space employs the affine invariant Riemannian metric (AIRM) for two tangent vectors in the tangent space. Its inner product is defined as

[0046] in, Let a point be on the manifold. The trace of the matrix is ​​represented; during the model training phase, the Riemann contrastive loss function is used to align features in different domains to eliminate domain offset between synthetic data and real clinical data.

[0047] In step S3, the latent diffusion model includes a manifold-aware adaptive group normalization layer (MA-AdaGN) for embedding manifold features. Injection denoising process The calculation formula is as follows:

[0048] in, These are the intermediate activation values ​​of the denoising network. and These are the mean and standard deviation, respectively. and The modulation parameters are generated from manifold features; and reverse SDE sampling is performed using a manifold-constrained inverse stochastic differential equation (SDE) solver, with the inverse time SDE formula being:

[0049] in, As latent variables; This is the drift coefficient function for inverse SDE, used to describe latent variables. In time The deterministic reverse evolution trend; The diffusion coefficient is used to control the intensity of noise injection during the reverse process; To embed with the manifold features The conditional log probability density with respect to the latent variables The gradient, i.e. the conditional fractional function, is approximately estimated by the denoising network based on the manifold features injected by the MA-AdaGN; This is the increment of the reverse-time standard Wiener process.

[0050] This inverse SDE method uses latent noise variables Stepwise reduction to deep latent variables Generate a high-fidelity true depth map This information is used for subsequent steps S4 to S6 to estimate the instrument anchoring scale and reconstruct the absolute measurement.

[0051] In step S5, the known physical dimensions of the biopsy forceps include the length of the forceps head. The calculation of the global metric scaling factor The formula is:

[0052] in, For the camera's focal length, This represents the projected pixel length of the biopsy forceps in the image. The angle between the biopsy forceps axis and the camera optical axis. relative depth map The average relative depth value within the biopsy forceps area. The included angle. The scale estimation model is obtained by combining the three-dimensional model of the biopsy forceps with the two-dimensional key points of the image using a PnP (Perspective-n-Point) solver. The scale estimation model is implemented using a Bayesian regressor, and its estimated variance approximates the Cramér-Rao Bound (CRB) for scale discovery under observable instrument conditions, thereby achieving sub-millimeter-level measurement accuracy.

[0053] Step S5 also includes an uncertainty gating mechanism: using a Bayesian regressor to estimate the biopsy clamp pose parameters. Posterior distribution uncertainty measure , recorded as And calculate the cognitive uncertainty index of the posterior distribution; if the cognitive uncertainty index is less than a preset threshold If the scale factor is calculated in step (5), then the scale factor is calculated. If the scale factor is greater than the preset threshold, the observation conditions are determined to be degraded, the scale factor is not calculated, and only the relative depth map is output.

[0054] Step S5 also includes a scale factor linear calibration step: for the global metric scale factor of the initial solution. The existing systematic bias is addressed using a pre-learned linear correction function. After correction, the calibrated scale factor is obtained. Among them, the correction coefficient and The depth estimation error was obtained by randomly selecting a very small number of samples (k≤6) from non-homologous clinical sequences and performing least squares fitting, in order to eliminate the depth estimation error caused by the difference in distribution between the training domain and the clinical domain.

[0055] In step S6, the reconstruction of the 3D point cloud The specific calculation is as follows:

[0056] For each pixel coordinate in the image Its corresponding three-dimensional spatial coordinates for:

[0057] in, Homogeneous coordinates , For pixels The absolute depth at that location.

[0058] To achieve the above process, this embodiment is deployed based on a deep learning workstation. The overall implementation process covers the construction and training of the network model, the inference and execution of the core algorithm, and the verification and analysis of clinical accuracy.

[0059] The specific implementation details for each stage are as follows:

[0060] (1) Construction and parameter configuration of network model

[0061] The deep learning model of this invention mainly consists of a feature extraction encoder, a manifold projection layer, and a diffusion model decoder. The specific configuration is as follows:

[0062] (1.1) Feature Encoder Architecture

[0063] (i) Swin-Transformer was selected as the backbone network to extract image features.

[0064] (ii) The encoder consists of 4 stages, and the number of module blocks in each stage is configured as [2,2,6,2].

[0065] (iii) Set the feature embedding dimension .

[0066] (iv) The output of the encoder is mapped through a fully connected layer to a 512-dimensional bottleneck layer, i.e., the latent space feature dimension. .

[0067] (1.2) Manifold projection layer

[0068] (i) To achieve Riemannian manifold learning, a Riemann projection layer is connected at the end of the encoder. .

[0069] (ii) This layer uses a geodesic stochastic gradient descent (SGD) optimizer to force the 512-dimensional latent features to be constrained to a symmetric positive definite (SPD) matrix manifold.

[0070] (1.3) Configuration of diffusion model

[0071] (i) The Latent Diffusion Model (LDM) is used as the depth generator.

[0072] (ii) The denoising network adopts the U-Net architecture, whose backbone is based on ResNet and embeds a cross-attention layer to fuse features from the manifold projection layer. .

[0073] (iii) Training phase: Set the total number of diffusion steps The steps are taken to learn the complete noise distribution.

[0074] (iv) Inference phase: To meet the real-time requirements of clinical practice, the DDIM (Denoising Diffusion Implicit Models) sampling algorithm is used for acceleration, reducing the number of sampling steps to [missing information]. step.

[0075] (2) Training strategy and optimization details of the model

[0076] To resolve conflicts in multi-task learning and ensure model convergence, this embodiment employs the following specific training strategy:

[0077] (2.1) Dataset preparation

[0078] (i) Construct a hybrid training dataset containing the "Physical Tier" (sample size). ) and "Clinical Video Domain" (Clinical Tier, sample size) ).

[0079] (ii) For physical phantom data, high-precision laser scan ground truth is provided for supervised scale learning; for clinical data, pseudo-labels generated by multi-view geometry (SfM) are used for weak supervision.

[0080] (2.2) Weights of the multi-task loss function

[0081] Construct the total loss function ,in For manifold contrast loss function; The geometric consistency loss function; Generate a loss function for diffusion; The loss function is the scale consistency function; the weight parameters for each sub-item are configured as follows:

[0082] (i) Manifold contrast loss weights ;

[0083] (ii) Geometric consistency loss weights ;

[0084] (iii) Diffusion generation of loss weights ;

[0085] (iv) Scaling estimation loss weights .

[0086] (2.3) Optimizer and Course Learning

[0087] (i) The AdamW optimizer is used, and the base learning rate is set to The weight decay is set to 0.05.

[0088] (ii) Implement a learning strategy: In the first 20 training cycles, only the manifold loss is activated (i.e., set...). (Other values ​​are 0), prioritizing feature alignment of the stable manifold space; then, geometric, diffusion, and scale losses are gradually added for joint optimization.

[0089] (2.4) Gradient surgery

[0090] To mitigate gradient interference in multi-task learning, the PCGrad algorithm is applied during the optimization process. When the gradient directions of different loss functions conflict (i.e., the cosine similarity is negative), the gradient is projected onto the normal plane to ensure that all tasks can converge collaboratively.

[0091] (2.5) Computational efficiency

[0092] Based on the above configuration, this invention processes a single image on an NVIDIA RTX 3090 GPU optimized with TensorRT. The average inference time for the high-resolution image is 32.5ms (approximately 30 FPS), which meets the requirements for real-time intraoperative assistance.

[0093] (3) Input data acquisition and initialization

[0094] (3.1) Image acquisition: The system acquires video frame images from the monocular endoscope in real time. In this embodiment, the input image is adjusted to... The resolution of pixels is adjusted to match the input requirements of deep learning models.

[0095] (3.2) Parameter loading: Load the intrinsic parameter matrix of the endoscope camera. (including focal length) and the known physical dimensions of biopsy forceps. (In this embodiment, the standard length is set) ).

[0096] (4) Geometric consistency perception and relative depth generation

[0097] This step aims to generate a structurally consistent but dimensionless relative depth map using the GCP module and diffusion model.

[0098] (4.1) Manifold feature embedding

[0099] (i) Image Input the trained Swin-Transformer encoder to extract feature maps.

[0100] (ii) Through the Riemann projection layer The feature map is mapped to a symmetric positive definite (SPD) covariance matrix. And based on this, we obtain the manifold feature embeddings located on the Riemannian manifold. This feature Located in Riemannian manifold The above method not only removes texture interference but also preserves second-order statistical information related to anatomical structures.

[0101] (4.2) Diffusion generation of manifold constraints

[0102] (i) Initialization: Sample random noise from a standard Gaussian distribution .

[0103] (ii) DDIM sampling: using the inverse SDE solver of manifold constraints (SDE-Solver) to measure manifold features As a condition, a denoising process is performed. To meet clinical real-time requirements (30 FPS), this embodiment uses the DDIM sampling algorithm, setting the number of generation steps to 5 ( ).

[0104] (iii) Output: A high-fidelity relative depth map is output after decoding. At this point, the depth value only represents a relative distance and does not have physical units.

[0105] (5) Instrument-anchored scale estimation (IASD) and gating mechanism

[0106] This step aims to use biopsy forceps, commonly seen in endoscopic images, as "in-situ physical anchors" to recover the global metric scaling factor. .

[0107] (5.1) Instrument perception

[0108] (i) Segmentation: A lightweight semantic segmentation network is used to infer the pixel-level mask region of the biopsy clamp from image I. .

[0109] (ii) Key point detection: Within the mask area, two key anatomical points of the biopsy forceps, namely the distal end and the proximal end, are located using a key point detection network (such as HRNet).

[0110] (iii) Calculation of projection length: Calculate the Euclidean distance between the two key points on the image plane to obtain the projection length of the biopsy forceps in the image. .

[0111] (5.2) Bayesian pose posterior estimation

[0112] Using a Bayesian regressor, combined with camera intrinsics And segmentation mask, to estimate the device pose parameters (Including projection length) and tilt angle posterior probability distribution This distribution reflects the model's confidence in its judgment of the current instrument posture.

[0113] (5.3) Uncertainty Gating Judgment

[0114] (i) Calculate the uncertainty measure of pose estimation , recorded as .

[0115] (ii) Set a safety threshold (Out-of-Distribution Threshold). Based on experimental verification, in order to balance usability and accuracy, this embodiment sets the threshold to [value missing]. .

[0116] (iii) Judgment logic:

[0117] like : If the observation conditions are deemed favorable, proceed to (5.4) for precise measurement.

[0118] like If the observation conditions are deemed degraded (e.g., the instrument is parallel to the optical axis or is blocked), enter the (5.5) safe retreat mode.

[0119] (5.4) Geometric Scale Solution

[0120] (i) Constructing a geometric projection model: such as Figure 2 As shown on the left, assuming the endoscopic camera follows a pinhole imaging model, the focal length is... The biopsy forceps, as a known physical anchor point, have a physical jaw length of [length missing]. (Standard length in this embodiment) Let the coordinates of the tool's pivot point in three-dimensional space be... Its corresponding absolute physical depth is .

[0121] (ii) Establish projection constraint equations: such as Figure 2 As shown on the right, the projection length of the tool on the image imaging plane Subject to physical depth and tilt angle The common constraints. According to the principle of perspective projection, the following formula is satisfied:

[0122]

[0123] Meanwhile, in the depth generation model of this invention, the predicted relative depth map With absolute physical depth There exists a linear proportional relationship that needs to be solved:

[0124] in, relative depth map In the biopsy forceps mask area The average depth value within.

[0125] (iii) Calculate the global metric scaling factor: Substitute the depth relation into the projection formula and rearrange it to obtain the global metric scaling factor. The calculation formula is as follows:

[0126]

[0127] The system will obtain the data from step (5.1). The results obtained in step (5.2) And network prediction Substituting into the above formula, the unique absolute scale of the current frame can be calculated. .

[0128] (5.5) Safety rollback mode

[0129] (i) The system determines that the current frame is an “out-of-distribution sample” (OOD) or a “high variance frame”.

[0130] (ii) Return a high variance error alert, stop the calculation of the absolute scale, and output only the relative depth map to prevent misleading doctors.

[0131] (6) Three-dimensional reconstruction and morphological quantification

[0132] (6.1) Measurement scaling: In precision measurement mode, utilizing the calculated... relative depth map Global scaling to absolute metric depth map (Unit: millimeters)

[0133] (6.2) Point cloud back projection: using the camera intrinsic inverse matrix The coordinates of each pixel in the image plane Back-projecting to 3D space generates a dense metric-level point cloud. :

[0134]

[0135] (6.3) Polyp diameter measurement

[0136] (i) In a 3D point cloud In this process, the polyp area is automatically located using a segmentation mask.

[0137] (ii) Calculate the Euclidean distance between any two points within the polyp region, and take the maximum value as the physical diameter of the polyp. :

[0138]

[0139] (6.4) Output Results: The final output of the system is a 3D point cloud containing absolute scale. The largest diameter of polyps (mm) and the scale uncertainty of this measurement. .

[0140] (7) Error decomposition and few-sample linear calibration

[0141] This invention reveals that the scale bias generated by monocular depth estimation in cross-domain deployment (from the physical phantom domain to the clinical application domain) is not random noise, but rather a shift with linear systematic characteristics. To eliminate this shift and achieve sub-millimeter measurement accuracy, this invention implements the following calibration procedure:

[0142] (7.1) Physical decomposition of absolute depth error

[0143] According to the geometric derivation of the present invention, the absolute depth error output by the system It consists of three parts:

[0144]

[0145] in, To learn depth bias; This is geometric observation noise; This represents the internal parameter calibration error.

[0146] (7.2) Few-sample linear calibration steps

[0147] (i) Sample collection: In the early stages of clinical application, a very small number of sample frames are randomly selected from non-homologous clinical sequences (e.g. (frame), to obtain the prediction scale factor of biopsy forceps in the image. and reference scales based on physical truth calculations .

[0148] (ii) Fitting the correction function: The least squares method is used to fit the linear equation. Solve for the correction coefficients and In this embodiment, the fitting result is approximately... , .

[0149] (iii) Online real-time correction: In the subsequent real-time reconstruction process, the system will use the scale factor calculated in step (5.4) to correct the scale factor. Substituting into the correction function, we obtain the final metric scaling factor. This mechanism can eliminate the effects of domain offset with extremely low data cost, enabling sub-millimeter level measurements.

[0150] (8) Accuracy verification and effect analysis

[0151] To verify the accuracy and effectiveness of the above-mentioned monocular endoscope 3D reconstruction and polyp measurement method based on geometric consistency perception and instrument anchoring, the PolypSense3D dataset (containing physical phantoms and real clinical data) was used for testing and verification.

[0152] (8.1) Verification process

[0153] (i) Data acquisition: Select clinical colonoscopy video sequences, using standard biopsy forceps (physical length) that were used and fully opened during the procedure. Using this as a reference, the true value of the polyp diameter is obtained, with the error controlled within [a certain range]. Within.

[0154] (ii) Reasoning and reconstruction: The video frames are input into the GCP-IASD system, and manifold feature extraction, relative depth generation, Bayesian pose estimation and scale calculation are performed in sequence to finally generate a 3D point cloud.

[0155] (iii) Performance evaluation: predicting the diameter The mean diameter error (MDE) is calculated by comparing it with the true value; at the same time, the relative geometric error (AbsRel) between the generated depth map and the real structure is evaluated.

[0156] (8.2) Results Comparison and Analysis

[0157] The measurement results of multiple sets of clinical data were calculated and statistically analyzed using the above method, and compared with existing mainstream monocular depth estimation methods (such as Marigold and UniDepth). The results show that:

[0158] (i) Qualitative comparative analysis of structural fidelity: such as Figure 4 As shown, when processing intestinal images with strong reflectivity and mucus interference, existing technologies (such as the Marigold model) generate depth maps and surface normals that exhibit significant structural distortion, artifacts, and over-smoothing, resulting in low surface fidelity. In contrast, this invention constrains features to a Riemannian manifold space using a GCP module, generating depth maps with sharper edges and surface normal maps that finely depict the minute undulations of the mucosal surface with higher geometric consistency. This demonstrates that this invention possesses stronger geometric structure preservation capabilities in complex clinical environments.

[0159] (ii) Quantitative comparative analysis of measurement accuracy: After introducing the IASD module and combining it with the small-sample calibration strategy described in step (7.2), the measurement error (MDE) of polyp diameter in this invention is significantly reduced to [missing value]. This indicator is far superior to the clinical classification recommendations for high-risk adenomas. The threshold is higher than that of traditional doctor's visual estimation.

[0160] (iii) Security and Uncertainty Gating Validation: For observation degradation scenarios in monocular vision (such as instruments parallel to the optical axis), the Bayesian gating mechanism of this invention can effectively evaluate pose uncertainty. Tests show that when At the same time, the system can effectively intercept high-risk frames by calculating the cognitive uncertainty index and automatically revert to the mode of outputting only relative depth, effectively preventing erroneous measurement data from interfering with medical decisions.

[0161] The beneficial effects of the above-mentioned structure of this invention are as follows: This invention provides a method for geometric consistency and instrument anchoring three-dimensional reconstruction and polyp measurement of endoscopes. This invention constrains features to Riemannian manifold space through a geometric consistency perception (GCP) module and aligns geometric structures using second-order statistics (covariance), significantly improving the structural fidelity of the model in complex clinical environments and solving the problem of cross-domain structural perception. This invention innovatively uses conventional biopsy forceps as "in-situ physical anchors" and recovers absolute scale from monocular video through the IASD module. Combined with a few-sample linear calibration mechanism, only 6 frames of reference data are needed to reduce the mean diameter error (MDE) from 3.18mm to 0.39mm, achieving sub-millimeter level clinical precision. This realizes sub-millimeter level quantization under monocular equipment and achieves accurate measurement without hardware modification. At the same time, the designed uncertainty gating mechanism can automatically identify high-risk frames and provide risk warnings, effectively preventing the output of misleading measurement results, improving surgical safety, and possessing high clinical-grade reliability.

[0162] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for geometrical consistency and instrument anchoring three-dimensional reconstruction and polyp measurement of endoscopes, characterized in that, The method includes the following steps: Step S1: Acquire video frame images under a monocular endoscope. and the intrinsic parameter matrix of the endoscope camera ; Step S2, using the Geometric Consistency Awareness (GCP) module to process the image The image features are extracted from a pre-trained feature encoder and then projected onto a symmetric positive definite (SPD) Riemannian manifold space through a manifold projection layer to obtain the manifold feature embedding. ; Step S3, utilizing the embedding of the manifold features The latent diffusion model (LDM) with conditions predicts the image through a manifold-constrained inverse stochastic differential equation (SDE) generation process. High-fidelity true depth map ; Step S4, for the image Semantic segmentation and keypoint detection are performed to identify the biopsy forceps region and extract the two-dimensional keypoint coordinates of the biopsy forceps; Step S5: Construct an instrument-anchored scale estimation model (IASD). Based on the coordinates of the two-dimensional key points and the known physical dimensions of the biopsy forceps, calculate the pose of the biopsy forceps relative to the camera, and calculate the global metric scale factor according to geometric projection constraints. ; Step S6, the relative depth map described in step S3 Compared with the global metric scaling factor described in step S5 Merge to generate an absolute metric depth map Based on this, a three-dimensional point cloud is reconstructed. In 3D point cloud In this process, the polyp region is automatically located using a segmentation mask, and the Euclidean distance between any two points within the polyp region is calculated to measure the physical size of the polyp. The maximum value is then taken as the physical diameter of the polyp. .

2. The method for geometric consistency and instrument anchoring three-dimensional reconstruction and polyp measurement of endoscopes according to claim 1, characterized in that: In step S2, the process of projecting image features onto a symmetric positive definite (SPD) Riemannian manifold space includes: converting the feature map output by the feature encoder... Construct as a regularized covariance matrix The formula is as follows: in, For spatial sample number, It is the identity matrix. To prevent small positive numbers from exhibiting singularity; the manifold feature is embedded. Located in Riemannian manifold superior.

3. The method for geometric consistency and instrument anchoring three-dimensional reconstruction and polyp measurement of endoscopes according to claim 2, characterized in that: The Riemannian manifold space employs the affine invariant Riemannian metric (AIRM). During the model training phase, the Riemann contrastive loss function is used to align features from different domains to eliminate domain offsets between synthetic data and real clinical data.

4. The method for geometric consistency and instrument anchoring three-dimensional reconstruction and polyp measurement of endoscopes according to claim 1, characterized in that: In step S3, the latent diffusion model includes a manifold-aware adaptive group normalization layer (MA-AdaGN) for embedding manifold features. Injection denoising process And using the inverse SDE solver with manifold constraints to perform reverse SDE sampling, noise is reduced. Gradually restore to depth .

5. The method for geometric consistency and instrument anchoring three-dimensional reconstruction and polyp measurement of endoscopes according to claim 1, characterized in that: In step S5, the known physical dimensions of the biopsy forceps include the length of the forceps head. The calculation of the global metric scaling factor The formula is: in, For the camera's focal length, This represents the projected pixel length of the biopsy forceps in the image. The angle between the biopsy forceps axis and the camera optical axis. relative depth map The average relative depth value within the biopsy forceps area.

6. The method for geometric consistency and instrument anchoring three-dimensional reconstruction and polyp measurement of endoscopes according to claim 5, characterized in that: The included angle The results were obtained by combining the three-dimensional model of the biopsy forceps with the two-dimensional key points of the image using the PnP (Perspective-n-Point) solver.

7. The method for geometric consistency and instrument anchoring three-dimensional reconstruction and polyp measurement of endoscopes according to claim 5, characterized in that: The scale estimation model is implemented using a Bayesian regressor, whose estimated variance approximates the Cramér-Rao Bound (CRB) for scale discovery under observable conditions, in order to achieve sub-millimeter-level measurement accuracy.

8. The method for geometric consistency and instrument anchoring three-dimensional reconstruction and polyp measurement of endoscopes according to claim 1, characterized in that: Step S5 also includes an uncertainty gating mechanism: using a Bayesian regressor to estimate the biopsy clamp pose parameters. Posterior distribution uncertainty measure , recorded as And calculate the cognitive uncertainty index of the posterior distribution; if the cognitive uncertainty index is less than a preset threshold If the scale factor is calculated in step (5), then the scale factor is calculated. If the scale factor is greater than the preset threshold, the observation conditions are determined to be degraded, the scale factor is not calculated, and only the relative depth map is output.

9. The method for geometric consistency and instrument anchoring three-dimensional reconstruction and polyp measurement of endoscopes according to claim 1, characterized in that: Step S5 also includes a scale factor linear calibration step: for the global metric scale factor of the initial solution. The existing systematic bias is addressed using a pre-learned linear correction function. After correction, the calibrated scale factor is obtained. Among them, the correction coefficient and The depth estimation error was obtained by randomly selecting a very small number of samples (k≤6) from non-homologous clinical sequences and performing least squares fitting, in order to eliminate the depth estimation error caused by the difference in distribution between the training domain and the clinical domain.

10. The method for geometric consistency and instrument anchoring three-dimensional reconstruction and polyp measurement of endoscopes according to claim 1, characterized in that: In step S6, the reconstruction of the 3D point cloud The specific calculation is as follows: for each pixel coordinate in the image Its corresponding three-dimensional spatial coordinates for: in, Homogeneous coordinates , For pixels The absolute depth at that location.