An endoscope image dynamic reconstruction method based on point cloud compensation and dual-domain morphing

The dynamic reconstruction method of endoscopic images based on point cloud compensation and dual-domain deformation solves the problems of insufficient point cloud initialization and low efficiency of dynamic deformation modeling caused by surgical tool occlusion, and achieves high-quality and fast endoscopic image reconstruction, which is suitable for augmented reality/virtual reality assisted diagnosis, surgical training and surgical navigation.

CN120510302BActive Publication Date: 2025-10-10NANCHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510969231.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-10
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

In existing endoscopic image dynamic reconstruction technology, surgical tool occlusion causes insufficient point cloud initialization quality and low efficiency of dynamic deformation modeling, making it difficult to meet the real-time requirements of clinical applications.

Method used

A method based on point cloud compensation and dual-domain deformation is adopted to generate a uniformly distributed initial point cloud by compensating the point cloud in the occluded area. A two-stage training strategy is used to combine the Gaussian ellipsoid and Fourier series models for dynamic deformation modeling, separating static optimization from dynamic optimization.

Benefits of technology

It effectively fills the voids in the point cloud, improves the prior information of the initialized point cloud, shortens the training time, achieves high-quality dynamic reconstruction, and meets the timeliness requirements of clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510302B_ABST
    Figure CN120510302B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of computer vision and three-dimensional reconstruction, and discloses an endoscope image dynamic reconstruction method based on point cloud compensation and double-domain deformation, which comprises the following steps: firstly, performing point cloud compensation on the occluded area, generating a reference frame point cloud for a reference frame from which a surgical tool is removed, and generating a compensation point cloud by using the visible part of the reference frame in other frames, and then combining the two to obtain an initialization point cloud with sufficient distribution prior; then, constructing a Gaussian ellipsoid set based on the initialization point cloud, and performing two-stage double-domain deformation training, in the first stage, performing Gaussian preheating training to keep the position and rotation of the Gaussian ellipsoid unchanged over time; in the second stage, using the time-domain Gaussian distribution function and the frequency-domain Fourier series to cooperatively fit the deformation variable of the position and rotation of the Gaussian ellipsoid over time. The application optimizes the initialization quality through point cloud compensation, and realizes fast and high-quality dynamic reconstruction of endoscope images through the two-stage double-domain deformation method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and three-dimensional reconstruction, and in particular to a method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation. Background Art

[0002] Endoscopy is an indispensable diagnostic and therapeutic tool in modern clinical medicine. Three-dimensional dynamic reconstruction of endoscopic images provides doctors with a comprehensive stereoscopic view of the surgical area, which is of great value for applications such as augmented reality (AR) / virtual reality (VR)-assisted diagnosis, surgical training, and surgical navigation.

[0003] Currently, dynamic scene reconstruction technology based on Gaussian splatting has attracted much attention due to its efficient rendering speed and high-quality reconstruction results. However, when this technology is applied to dynamic reconstruction of endoscopic images, it still faces two major challenges:

[0004] First, the initialization quality of the Gaussian point cloud is insufficient. In endoscopic surgery videos, surgical tools frequently appear and obscure tissue. These tools often need to be removed during 3D reconstruction. This results in point cloud holes in the areas removed by the tools during point cloud initialization. This leads to an uneven distribution of the initialized Gaussian point cloud and insufficient prior information, which in turn affects the final reconstruction quality.

[0005] Second, dynamic deformation modeling is inefficient. The dynamic characteristics of human tissue (such as deformation caused by heartbeat and breathing) complicate reconstruction. Existing dynamic reconstruction methods typically use complex neural network structures, such as decoders based on multi-layer perceptrons (MLPs) and time-consuming feature plane interpolation, to model the deformation field. While this approach can represent complex deformations, its significant computational overhead results in extremely long model training times, making it difficult to meet the real-time or near-real-time requirements of clinical applications.

[0006] Therefore, how to solve the problem of missing initialization point cloud caused by occlusion of surgical tools and design an efficient dynamic deformation modeling mechanism to significantly shorten the training time while ensuring reconstruction quality is a technical problem that needs to be solved urgently in this field. Summary of the Invention

[0007] The main purpose of the present invention is to provide a method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation, so as to solve the problems mentioned in the background technology of insufficient prior knowledge of the initialization point cloud distribution and long dynamic reconstruction training time.

[0008] In a first aspect, the present invention provides a method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation, comprising the following steps:

[0009] The occluded area point cloud compensation step includes: obtaining an endoscopic image sequence comprising multiple frames of images, wherein the endoscopic image sequence includes a color image, a depth map, and a surgical tool mask map for each frame; generating an initialization point cloud with sufficient distribution prior based on the endoscopic image sequence, wherein the initialization point cloud is formed by merging the reference frame point cloud and the compensation point cloud;

[0010] Two-stage dual-domain deformation steps: construct a set of Gaussian ellipsoids based on the initialized point cloud; perform two-stage training on the Gaussian ellipsoid set to reconstruct the dynamic scene; wherein, the first stage of the two-stage training is Gaussian warm-up training, which keeps the position and rotation properties of each Gaussian ellipsoid unchanged over time; the second stage is dual-domain deformation training, which collaboratively fits the deformation variables of the position and rotation properties of each Gaussian ellipsoid over time by jointly using the time domain model and the frequency domain model.

[0011] As an optional implementation of the first aspect of the present application, the occluded area point cloud compensation step specifically includes: based on the selected reference frame, using the surgical tool mask map to remove the surgical tool from the color image, and combining the depth map and the internal and external parameters of the camera to generate the reference frame point cloud; identifying the tissue area occluded by the surgical tool in the reference frame but visible in one or more other frames in the image sequence, and generating a surgical tool compensation mask map; the surgical tool compensation mask map is obtained by calculating the set difference operation of the surgical tool mask map of the reference frame and the surgical tool mask maps of other frames; based on the surgical tool compensation mask map, the color images, depth maps and internal and external parameters of the camera of the other frames, generating the compensated point cloud; merging the reference frame point cloud and the compensated point cloud, and sampling a predetermined number of points therefrom to form the final initialized point cloud.

[0012] As an optional implementation of the first aspect of the present application, the first-stage Gaussian warm-up training in the two-stage dual-domain deformation step aims to optimize the initial static properties of the Gaussian ellipsoid; the training process keeps the position and rotation quaternion of each Gaussian ellipsoid constant in the time dimension, and only optimizes the initial position, initial rotation, scaling matrix, spherical harmonic function coefficients and opacity of the Gaussian ellipsoid for a predetermined number of iterations to generate a more precisely positioned static Gaussian ellipsoid set around the initialized point cloud.

[0013] As an optional implementation of the first aspect of the present application, the second stage dual-domain deformation training in the two-stage dual-domain deformation step models the position and rotation properties of each Gaussian ellipsoid at any time as the superposition of its basic properties at the reference time and a deformation variable that changes with time; the deformation variable is decomposed into the sum of a time domain deformation variable and a frequency domain deformation variable.

[0014] As an optional implementation of the first aspect of the present application, the fitting method of the time domain deformation variable and the frequency domain deformation variable is: the time domain deformation variable is represented by the weighted sum of one or more Gaussian distribution functions, and the center, variance and weight of the Gaussian distribution function are all learnable parameters for capturing the non-periodic fine deformation of the tissue; the frequency domain deformation variable is represented by a Fourier series, and the sine term coefficient and cosine term coefficient of the Fourier series are all learnable parameters for capturing the periodic or large-scale rough deformation of the tissue.

[0015] As an optional implementation scheme of the first aspect of the present application, the training process of the two-stage dual-domain deformation step is driven by minimizing a total loss function; the total loss function includes a weighted sum of color loss, depth loss and spatial loss; the color loss and depth loss are respectively used to constrain the consistency between the rendered color image and depth map and the true value, and are only calculated in the non-surgical tool mask area; the spatial loss is to apply total variation regularization to the rendered color image and depth map to reduce visual artifacts that may appear in areas occluded by surgical tools and ensure spatial smoothness.

[0016] As an optional implementation of the first aspect of the present application, after generating the compensated point cloud, a preset percentage of point clouds are randomly extracted from the compensated point cloud; after merging the reference frame point cloud and the extracted compensated point cloud, a preset number of points are randomly selected from the merged point cloud as the final initialization point cloud.

[0017] In a second aspect, an embodiment of the present application provides an endoscopic image dynamic reconstruction system based on point cloud compensation and dual-domain deformation, comprising:

[0018] an occlusion region point cloud compensation module configured to acquire an endoscopic image sequence comprising multiple frames, the endoscopic image sequence including a color image, a depth map, and a surgical tool mask map for each frame; and generate an initialization point cloud with sufficient distribution prior based on the endoscopic image sequence, the initialization point cloud being formed by merging a reference frame point cloud and a compensation point cloud;

[0019] A two-stage dual-domain deformation module is configured to construct a set of Gaussian ellipsoids based on the initialization point cloud; the Gaussian ellipsoid set is trained in two stages to reconstruct a dynamic scene; wherein the first stage of the two-stage training is Gaussian warm-up training, in which the position and rotation properties of each Gaussian ellipsoid are kept constant over time; and the second stage is dual-domain deformation training, in which the position and rotation properties of each Gaussian ellipsoid are collaboratively fitted with the deformation variables over time by jointly using a time domain model and a frequency domain model.

[0020] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein when the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0021] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0022] Compared with the existing technology, the present invention proposes a method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation, which has the following beneficial effects:

[0023] 1. The occluded area point cloud compensation mechanism effectively fills the point cloud holes caused by surgical tool removal, obtaining a more evenly distributed initial point cloud with more sufficient prior information, laying a solid foundation for high-quality reconstruction.

[0024] 2. A two-stage training strategy separates static optimization from dynamic optimization, improving training stability and efficiency.

[0025] 3. The innovative use of a dual-domain deformation model that combines Gaussian distribution function and Fourier series avoids the time-consuming feature interpolation and MLP decoding process, greatly shortens the model training time, and can achieve high-quality dynamic reconstruction in a short time, meeting the timeliness requirements of clinical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is a schematic diagram of the overall architecture of an endoscopic image dynamic reconstruction method based on point cloud compensation and dual-domain deformation proposed by the present invention;

[0027] Figure 2 Schematic diagram of the process of compensating the occluded area point cloud in the method of the present invention;

[0028] Figure 3 Schematic diagram of the process of the two-stage dual-domain deformation step in the method of the present invention;

[0029] Figure 4 Schematic diagram of the structure of an endoscopic image dynamic reconstruction system based on point cloud compensation and dual-domain deformation proposed in the present invention. DETAILED DESCRIPTION

[0030] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0031] The terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the application can be implemented in a sequence other than those illustrated or described here. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated before and after are in a kind of "or" relationship. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically limited.

[0032] Example 1

[0033] See also Figure 1 , which is a schematic diagram of the overall architecture of an endoscopic image dynamic reconstruction method based on point cloud compensation and dual-domain deformation provided by an embodiment of the present invention, which mainly includes two steps: point cloud compensation of occluded area and two-stage dual-domain deformation.

[0034] The occluded area point cloud compensation step includes: obtaining an endoscopic image sequence comprising multiple frames of images, wherein the endoscopic image sequence includes a color image, a depth map, and a surgical tool mask map for each frame; generating an initialization point cloud with sufficient distribution prior based on the endoscopic image sequence, wherein the initialization point cloud is formed by merging the reference frame point cloud and the compensation point cloud;

[0035] Two-stage dual-domain deformation steps: construct a set of Gaussian ellipsoids based on the initialized point cloud; perform two-stage training on the Gaussian ellipsoid set to reconstruct the dynamic scene; wherein, the first stage of the two-stage training is Gaussian warm-up training, which keeps the position and rotation properties of each Gaussian ellipsoid unchanged over time; the second stage is dual-domain deformation training, which collaboratively fits the deformation variables of the position and rotation properties of each Gaussian ellipsoid over time by jointly using the time domain model and the frequency domain model.

[0036] Specifically, in order to remove surgical tools from endoscopic images and quickly obtain a uniformly distributed and sufficient Gaussian point cloud, the present invention proposes a point cloud compensation process for the occluded area as follows: Figure 2 shown.

[0037] In the process of Gaussian point cloud initialization, the point cloud compensation is performed on the surgical tool occlusion area in the reference frame, the point cloud initialization is performed on the image of the reference frame (0th frame) in which the surgical tool is removed, and the reference frame point cloud is generated . Due to the removal of the surgical tool in the reference frame, the reference frame point cloud has point cloud loss in the surgical tool removal area. In order to make up for this part of the missing point cloud, the visible part of the surgical tool occlusion area in the reference frame is further searched in the remaining frames, and the point cloud initialization is performed on the area. Then, 1% of the point cloud generated from the remaining frames is randomly extracted as the remaining frame compensation point cloud . Finally, the reference frame point cloud and the remaining frame compensation point cloud are combined, and 30,000 points are randomly selected therefrom as the final initialization point cloud .

[0038] For a given endoscopic color image , depth map and surgical tool mask image . First, the reference frame color image is removed from the corresponding surgical tool in the reference frame, and is projected into a three-dimensional space to obtain the generated point cloud in the reference frame (reference frame point cloud ). The calculation of the reference frame point cloud can be represented by formula (1):

[0039]

[0040] wherein, represents a camera intrinsic matrix, the parameters of the matrix are only related to the camera itself; represents an extrinsic matrix corresponding to the reference frame image, the parameters of the matrix are only related to the shooting pose of the camera; represents the reference frame endoscopic color image, represents the depth map corresponding to the reference frame endoscopic color image, represents the surgical tool mask image corresponding to the reference frame endoscopic color image, represents the point multiplication of the pixel element.

[0041] Due to the removal of the surgical tool in the reference frame , the reference frame point cloud has point cloud loss in the surgical tool removal area. However, during the movement of the surgical tool, the tissue occluded by the surgical tool in the reference frame can be visible in other frames. Based on this observation, next, the surgical tool mask image , and collect the masks corresponding to the new tissue pixels appearing in the remaining frames to obtain the surgical tool compensation mask map . Surgical tool compensation mask map The calculation of is shown in formula (2):

[0042]

[0043] in, and Represent the surgical tool mask images corresponding to the endoscopic images in the reference frame and the remaining frames, Represents element intersection.

[0044] Then, the surgical tool compensation mask obtained above is Project into 3D space to obtain the residual frame compensated point cloud . The reference frame point cloud Compensate point cloud with remaining frames Merge to get the final initialized point cloud . Remaining frame compensated point cloud And the final initialized point cloud It can be expressed by formula (3) and formula (4) respectively:

[0045]

[0046]

[0047] in, represents the intrinsic parameter matrix of the camera, represents the extrinsic parameter matrix corresponding to the endoscopic color image in the remaining frames, represents the color image of the endoscope in the remaining frames, represents the depth map corresponding to each endoscopic color image in the remaining frames, represents the surgical tool compensation mask obtained by formula (4), represents the dot product of each pixel element, and The reference frame initialization point cloud and the remaining frame compensated point cloud The random sampling coefficient of the present invention is set , .

[0048] Finally, the initialized point cloud With each point as the center, a set of Gaussian ellipsoids is created. The process can be expressed by formula (5) and formula (6):

[0049]

[0050]

[0051] in, represents the position of the Gaussian ellipsoid, represents the rotation matrix, represents the scaling matrix, Indicates the center of the Gaussian ellipsoid position.

[0052] Specifically, in order to achieve high-quality deformable tissue reconstruction in a relatively short time, the present invention proposes a two-stage dual-domain deformation to handle the deformation problem during the dynamic reconstruction of endoscopic images. The two-stage dual-domain deformation process proposed by the present invention is as follows: Figure 3 As shown. This method is divided into two stages of training: in the first stage of training, 1000 iterations of Gaussian warm-up training are performed to keep the position and rotation of the Gaussian ellipsoid unchanged over time. The purpose of this is to generate some Gaussian ellipsoids with more precise positions near the initialized Gaussian ellipsoid; in the second stage of training, dual-domain deformation is used, that is, Gaussian distribution function is used in the time domain and Fourier series is used in the frequency domain to collaboratively approximate the changes in the position and rotation of the Gaussian ellipsoid over time. In the time domain, the Gaussian distribution function can provide more flexibility for the basis function to adapt to the subtle deformation of some tissues in endoscopic images, and is good at capturing fine vascular structures. In addition, in the frequency domain, the Fourier series is good at capturing larger deformations or tissue appearance contours associated with intense movement. The combination of the two can achieve more accurate and rapid dynamic reconstruction.

[0053] In the first stage of training, Gaussian warm-up training is performed to keep the position and rotation of the Gaussian ellipsoid constant over time. In the second stage of training, the position and rotation properties of the Gaussian ellipsoid that change over time are , , the Gaussian distribution function and Fourier series are used to collaboratively calculate the change of each Gaussian attribute over time. Specifically, the present invention calculates the position and rotation attributes of each Gaussian ellipsoid , Changes over time, described as their value at frame 0 Basic properties , With time The change in attribute By using the Gaussian distribution function in the time domain and the Fourier series in the frequency domain, the position and rotation properties of the Gaussian ellipsoid are fitted to change with time. The calculation is shown in formula (7):

[0054]

[0055] in, , Represents the Gaussian ellipsoid at time The position and rotation properties, Represents Gaussian properties over time The amount of change, Depend on and It consists of two parts, namely .in, represents the time domain variation generated by the Gaussian distribution function, Represents the frequency domain variation produced by the Fourier series.

[0056] By using the Gaussian distribution function in the time domain and in the frequency domain using Fourier series To calculate the change of Gaussian attributes over time, it can be expressed by formula (8) and formula (9) respectively:

[0057]

[0058]

[0059] in, represents a learnable center, represents the variance of the Gaussian distribution function, Represents the position and rotation properties of the Gaussian ellipsoid, that is , Represents the learnable weight value assigned to each attribute of the Gaussian ellipsoid, and represent the coefficients of the Fourier series sine and cosine functions, respectively. represents the number of terms in the Fourier series. In the experiment of the present invention, .

[0060] In order to speed up the training of the model, the present invention experiments use color loss and depth loss Supervise the rendered color and depth images using spatial loss to constrain the predicted color and depth maps.

[0061] Calculating color loss and depth loss Previously, the method of the present invention used a method of accumulating pixel color and depth to calculate the color and depth of a pixel. This method is used to calculate the color of a pixel. and depth As shown in formula (10) and formula (11) respectively:

[0062]

[0063]

[0064] in, It is by The color is calculated from the SH coefficient of a Gaussian ellipsoid, It is Gaussian ellipsoid Values ​​can be obtained through the covariance matrix Multiplies the opacity of the Gaussian ellipsoid by get.

[0065] Color loss for network training and depth loss The calculations are shown in formula (12) and formula (13):

[0066]

[0067]

[0068] in, Represents the rendered color and depth maps, Represent the true color and depth map respectively, Indicates the surgical tool mask corresponding to each frame of endoscopic image, Represents 2D space coordinates.

[0069] The experiment of this invention adopts space loss To avoid black artifacts in areas occluded by surgical tools, spatial loss Using total variation loss ( To regularize the rendering results, The calculation method is shown in formula (14):

[0070]

[0071] In each iteration, the total loss of optimization is calculated as shown in formula (15):

[0072]

[0073] in, Indicates the weight used to balance the loss during training. In the experiment of this invention, , , .

[0074] In this embodiment, the datasets used in the present invention are two datasets commonly used in the current endoscopic image dynamic reconstruction task, namely, the EndoNeRF dataset and the Hamlyn dataset.

[0075] The EndoNeRF dataset consists of color images extracted from frames extracted from endoscopic videos captured from a single viewpoint during prostatectomy surgery using the da Vinci robot. The dataset also includes depth maps and surgical tool masks, aiming to reconstruct surgical scenes involving non-rigid deformations and tool occlusions. It includes six sequences with an image resolution of 512 × 640 pixels. Two of these sequences are publicly available: one for tissue cutting, providing 156 frames, and the other for soft tissue stretching, providing 63 frames. The experiments in this paper were conducted on these two publicly available sequences, and the division of the EndoNeRF dataset remains consistent with the original EndoNeRF paper.

[0076] The Hamlyn dataset is a collection of cardiac and in-vivo sequences captured by the da Vinci robot during surgery. The dataset contains rectified images, stereo depth maps, and camera intrinsics. The dataset used in this experiment is derived from seven challenge sequences provided by ForPlane, which exhibit weak textures, deformations, reflections, surgical tool occlusions, and illumination changes. Each sequence contains 301 frames with an image resolution of 480×640 pixels. ForPlane generates a corresponding depth map and surgical tool mask for each sequence and divides the Hamlyn dataset into a training set of 151 frames and a test set of the remaining 150 frames. This division ensures sufficient training data while providing a robust set for evaluation. The division of the Hamlyn dataset in this experiment is consistent with the original ForPlane paper.

[0077] Evaluation indicators: In the task of dynamic reconstruction of endoscopic images, commonly used evaluation indicators include peak signal-to-noise ratio, structural similarity index, and learning-perceived image block similarity.

[0078] (1) Peak Signal-to-Noise Ratio (PSNR) is often used to compare the difference between the input image and the rendered image. The higher the PSNR value, the higher the similarity between the quality of the rendered image and the input image. PSNR represents the spatial structural difference between the images by calculating the mean squared error (MSE) between the input image and the rendered image. The calculation of PSNR is shown in formula (16):

[0079]

[0080] Where MAX represents the maximum possible value of a pixel in the image, and the calculation formula of MSE is shown in formula (17):

[0081]

[0082] in, The coordinates on the clean image are elements, and The coordinates on the noise image are The element of the image is , Indicates the length of the image, .

[0083] (2) The Structural Similarity Index (SSIM) evaluates from the perspective of image structure, defining structural information as properties that reflect the structure of objects in the scene. SSIM measures image distortion by brightness, contrast, and structure, and can also be used to assess image similarity. A larger SSIM value indicates a higher degree of image similarity, and generally better image quality.

[0084] (3) Learned Perceptual Image Patch Similarity (LPIPS) simulates the human visual system's perception of images and uses a neural network model to approximate the image similarity perceived by the human visual system. The smaller the LPIPS value, the smaller the difference between the two images and the better the quality of the synthesized image.

[0085] It should be noted that the parameters of the present invention are set as follows: the image resolution of the EndoNeRF dataset input is 640×512 pixels, and the image resolution of the Hamlyn dataset input is 640×480 pixels. In the initialization point cloud stage, 1% of the compensation point cloud generated by the remaining frames is randomly selected as the compensation point cloud of the remaining frames. , randomly select 30,000 points from the overall initialization point cloud as the final initialization point cloud Adam is used as the optimizer, and the initial learning rate is 1.6×10 -3 The Gaussian warm-up was performed for 1000 iterations, and the optimization of the entire model framework was performed for 3000 iterations.

[0086] Example 2

[0087] See also Figure 4 , shown is a schematic structural diagram of an endoscopic image dynamic reconstruction system based on point cloud compensation and dual-domain deformation proposed in the second embodiment of the present application. The system includes the following key modules:

[0088] The occlusion region point cloud compensation module 100 is configured to acquire an endoscopic image sequence comprising multiple frames, wherein the endoscopic image sequence includes a color image, a depth map, and a surgical tool mask map for each frame; generate an initialization point cloud with sufficient distribution prior based on the endoscopic image sequence, wherein the initialization point cloud is formed by merging the reference frame point cloud and the compensation point cloud;

[0089] The two-stage dual-domain deformation module 200 is configured to construct a set of Gaussian ellipsoids based on the initialization point cloud; the Gaussian ellipsoid set is trained in two stages to reconstruct the dynamic scene; wherein the first stage of the two-stage training is Gaussian warm-up training, in which the position and rotation attributes of each Gaussian ellipsoid are kept unchanged over time; the second stage is dual-domain deformation training, which collaboratively fits the deformation variables of the position and rotation attributes of each Gaussian ellipsoid over time by jointly using the time domain model and the frequency domain model.

[0090] In this embodiment, the present invention addresses the issues of insufficient Gaussian initialization point cloud distribution priors and long training times in existing dynamic endoscopic image reconstruction methods based on Gaussian splattering. An occlusion region point cloud compensation module and a two-stage dual-domain deformation module are proposed to alleviate these issues. First, the occlusion region point cloud compensation module initializes the point cloud of a reference frame. Then, by querying other frames, it compensates for the point cloud of areas occluded by surgical tools in the reference frame, thereby obtaining an initial point cloud with a relatively sufficient distribution prior. Second, a two-stage dual-domain deformation module is proposed. In the first stage of training, the position and rotation of the Gaussian ellipsoid are kept constant over time for warm-up training. In the second stage of training, the Gaussian distribution function and Fourier series are used to collaboratively approximate the temporal deformation of the position and rotation of the Gaussian ellipsoid. This deformation calculation method eliminates the time-consuming feature plane interpolation and MLP decoding, thereby achieving faster reconstruction speed. The proposed method outperforms existing methods, reducing training time and GPU usage, and achieving higher reconstruction quality with a training time of approximately one minute.

[0091] In an embodiment of the present application, a dynamic endoscopic image reconstruction system based on point cloud compensation and dual-domain deformation can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a network attached storage (NAS), a personal computer (PC), etc., which is not specifically limited in the embodiment of the present application.

[0092] In the embodiments of the present application, a dynamic endoscopic image reconstruction system based on point cloud compensation and dual-domain deformation can be a device having an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.

[0093] The embodiment of the present application provides an endoscopic image dynamic reconstruction system based on point cloud compensation and dual-domain deformation, which can achieve Figure 1 In order to avoid repetition, the various processes of the method embodiment of the dynamic reconstruction method of endoscopic images based on point cloud compensation and dual-domain deformation are not described here.

[0094] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the various processes of the above-mentioned embodiment of the method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation are implemented, and the same technical effect can be achieved. To avoid repetition, they will not be described here.

[0095] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the embodiment of the above-mentioned method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0096] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0097] It should be noted that in this application, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or apparatus including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or apparatus. Without more limitations, the element defined by the statement "including a" does not exclude the presence of additional identical elements in the process, method, article or apparatus including the element. In addition, it should be pointed out that the scope of the method and device in the embodiments of the present application is not limited to the order of performing the functions shown or discussed, but can also include performing the functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method can be performed in an order different from that described, and various steps can also be added, omitted or combined. In addition, the features described with reference to certain examples can be combined in other examples.

[0098] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment method can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for making a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) execute the method described in each embodiment of the present application.

[0099] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the above specific embodiments, and the above specific embodiments are only illustrative, not limiting, and those skilled in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims.

Claims

1. A method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation, characterized in that: The following steps are involved: The occluded area point cloud compensation step includes: obtaining an endoscopic image sequence comprising multiple frames of images, wherein the endoscopic image sequence includes a color image, a depth map, and a surgical tool mask map for each frame; generating an initialization point cloud with uniform distribution and sufficient prior information based on the endoscopic image sequence, wherein the initialization point cloud is formed by merging a reference frame point cloud and a compensation point cloud, specifically comprising: based on a selected reference frame, removing the surgical tool from the color image using the surgical tool mask map, and generating the reference frame point cloud in combination with the depth map and camera internal and external parameters; identifying tissue areas that are occluded by the surgical tool in the reference frame but visible in one or more other frames in the image sequence, and generating a surgical tool compensation mask map; the surgical tool compensation mask map is obtained by calculating a set difference operation between the surgical tool mask map of the reference frame and the surgical tool mask maps of other frames; generating the compensation point cloud based on the surgical tool compensation mask map, the color images, depth maps, and camera internal and external parameters of the other frames; merging the reference frame point cloud and the compensation point cloud, and sampling a predetermined number of points therefrom to form the final initialization point cloud; Two-stage dual-domain deformation steps: construct a Gaussian ellipsoid set based on the initialized point cloud; perform two-stage training on the Gaussian ellipsoid set to reconstruct the dynamic scene; wherein, the first stage of the two-stage training is Gaussian warm-up training, which keeps the position and rotation properties of each Gaussian ellipsoid unchanged over time; the second stage is dual-domain deformation training, which collaboratively fits the deformation variables of the position and rotation properties of each Gaussian ellipsoid over time by jointly using the time domain model and the frequency domain model.

2. The method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation according to claim 1, characterized in that: The first stage of Gaussian warm-up training in the two-stage dual-domain deformation step aims to optimize the initial static properties of the Gaussian ellipsoid; The training process keeps the position and rotation quaternion of each Gaussian ellipsoid constant in the time dimension, and only optimizes the initial position, initial rotation, scaling matrix, spherical harmonic function coefficients and opacity of the Gaussian ellipsoid for a predetermined number of iterations to generate a static Gaussian ellipsoid set with more accurate positions around the initialization point cloud.

3. The method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation according to claim 1 or 2, characterized in that: The second stage of dual-domain deformation training in the two-stage dual-domain deformation step models the position and rotation properties of each Gaussian ellipsoid at any time as the superposition of its basic properties at the reference time and a deformation variable that changes with time; the deformation variable is decomposed into the sum of a time domain deformation variable and a frequency domain deformation variable.

4. The method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation according to claim 3, characterized in that: The fitting method of the time domain shape variable and the frequency domain shape variable is: The temporal deformation variable is represented by a weighted sum of one or more Gaussian distribution functions, wherein the center, variance, and weight of the Gaussian distribution function are all learnable parameters for capturing the non-periodic fine deformation of the tissue; The frequency domain deformation variable is represented by a Fourier series, and the sine term coefficient and the cosine term coefficient of the Fourier series are both learnable parameters for capturing the periodicity or large-scale rough deformation of the tissue.

5. The method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation according to claim 1, characterized in that: The training process of the two-stage dual-domain deformation step is driven by minimizing a total loss function; the total loss function includes a weighted sum of color loss, depth loss and spatial loss; The color loss and depth loss are respectively used to constrain the consistency between the color image and depth map generated by rendering and the true value, and are only calculated in the non-surgical tool mask area; The spatial loss is to apply total variation regularization to the rendered color image and depth map to reduce visual artifacts that may appear in areas occluded by surgical tools and ensure spatial smoothness.

6. An endoscopic image dynamic reconstruction system based on point cloud compensation and dual-domain deformation, characterized in that: include: The occlusion area point cloud compensation module is configured to acquire an endoscopic image sequence comprising multiple frames of images, wherein the endoscopic image sequence includes a color image, a depth map, and a surgical tool mask map for each frame; generate an initialization point cloud with uniform distribution and sufficient prior information based on the endoscopic image sequence, wherein the initialization point cloud is formed by merging a reference frame point cloud and a compensation point cloud, specifically comprising: based on a selected reference frame, removing the surgical tool from the color image using the surgical tool mask map, and generating the reference frame point cloud in combination with the depth map and camera intrinsic and extrinsic parameters; identifying tissue areas that are occluded by the surgical tool in the reference frame but visible in one or more other frames in the image sequence, and generating a surgical tool compensation mask map; the surgical tool compensation mask map is obtained by calculating a set difference operation between the surgical tool mask map of the reference frame and the surgical tool mask maps of the other frames; generating the compensation point cloud based on the surgical tool compensation mask map, the color images, depth maps, and camera intrinsic and extrinsic parameters of the other frames; merging the reference frame point cloud and the compensation point cloud, and sampling a predetermined number of points therefrom to form the final initialization point cloud; A two-stage dual-domain deformation module is configured to construct a Gaussian ellipsoid set based on the initialized point cloud; the Gaussian ellipsoid set is trained in two stages to reconstruct a dynamic scene; wherein the first stage of the two-stage training is Gaussian warm-up training, in which the position and rotation properties of each Gaussian ellipsoid are kept constant over time; and the second stage is dual-domain deformation training, in which the position and rotation properties of each Gaussian ellipsoid are collaboratively fitted with the deformation variables over time by jointly using a time domain model and a frequency domain model.

7. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation as described in any one of claims 1 to 5 are implemented.

8. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Dynamic scene reconstruction method and device based on multi-scale Gaussian sphere

    CN119991973A

  • Intelligent jitter identification and compensation method and system for endoscopic surgery image

    CN120075619A