Endoscope image dynamic reconstruction method based on point cloud compensation and double-domain deformation

Through the occlusion area point cloud compensation and two-stage dual-domain deformation method, the problem of uneven initial point cloud and low dynamic deformation modeling efficiency in dynamic reconstruction of endoscopic images is solved, and high-quality and rapid reconstruction effects are achieved.

CN120510302AActive Publication Date: 2025-08-19NANCHANG UNIV

Patent Information

Application Number
CN202510969231.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-08-19
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

In the existing endoscopic image dynamic reconstruction technology, the problems of uneven distribution of initial point clouds and low efficiency in dynamic deformation modeling caused by surgical tool occlusion are difficult to meet the real-time requirements of clinical applications.

Method used

The occlusion area point cloud compensation and two-stage two-domain deformation method are used to generate an evenly distributed initialized point cloud and dynamically reconstruct using a deformation model combining Gaussian ellipsoids and Fourier series.

Benefits of technology

It improves the quality and reconstruction efficiency of initialized point clouds, shortens training time, and meets the real-time requirements of clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510302A_ABST
    Figure CN120510302A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision and three-dimensional reconstruction, and discloses an endoscope image dynamic reconstruction method based on point cloud compensation and double-domain deformation, and the method comprises the steps: firstly, carrying out the point cloud compensation of a shielding region, generating a reference frame point cloud for a reference frame from which a surgical tool is removed, and carrying out the point cloud compensation; generating compensation point clouds by utilizing parts, visible to the shielding area of the reference frame, in other frames, and combining the compensation point clouds to obtain initialized point clouds with sufficient distribution priori; then, a Gaussian ellipsoid set is constructed based on the initialized point cloud, and two-stage double-domain deformation training is executed, in the first stage, Gaussian preheating training is carried out, and the position and rotation of a Gaussian ellipsoid are kept not to change along with time; in the second stage, the Gaussian distribution function of the time domain and the Fourier series of the frequency domain are adopted to cooperatively fit the position of the Gaussian ellipsoid and the deformation quantity of rotation along with time. According to the method, the initialization quality is optimized through point cloud compensation, and rapid and high-quality endoscope image dynamic reconstruction is realized through a two-stage double-domain deformation method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and three-dimensional reconstruction, and in particular to a method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation. Background Art

[0002] Endoscopy is an indispensable diagnostic and therapeutic tool in modern clinical medicine. Three-dimensional dynamic reconstruction of endoscopic images provides doctors with a comprehensive stereoscopic view of the surgical area, which is of great value for applications such as augmented reality (AR) / virtual reality (VR)-assisted diagnosis, surgical training, and surgical navigation.

[0003] Currently, dynamic scene reconstruction technology based on Gaussian splatting has attracted much attention due to its efficient rendering speed and high-quality reconstruction results. However, when this technology is applied to dynamic reconstruction of endoscopic images, it still faces two major challenges: First, the initialization quality of the Gaussian point cloud is insufficient. In endoscopic surgery videos, surgical tools frequently appear and obscure tissue. These tools often need to be removed during 3D reconstruction. This results in point cloud holes in the areas removed by the tools during point cloud initialization. This leads to an uneven distribution of the initialized Gaussian point cloud and insufficient prior information, which in turn affects the final reconstruction quality.

[0004] Second, dynamic deformation modeling is inefficient. The dynamic characteristics of human tissue (such as deformation caused by heartbeat and breathing) complicate reconstruction. Existing dynamic reconstruction methods typically use complex neural network structures, such as decoders based on multi-layer perceptrons (MLPs) and time-consuming feature plane interpolation, to model the deformation field. While this approach can represent complex deformations, its significant computational overhead results in extremely long model training times, making it difficult to meet the real-time or near-real-time requirements of clinical applications.

[0005] Therefore, how to solve the problem of missing initialization point cloud caused by occlusion of surgical tools and design an efficient dynamic deformation modeling mechanism to significantly shorten the training time while ensuring reconstruction quality is a technical problem that needs to be solved urgently in this field. Summary of the Invention

[0006] The main purpose of the present invention is to provide a method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation, so as to solve the problems mentioned in the background technology of insufficient prior knowledge of the initialization point cloud distribution and long dynamic reconstruction training time.

[0007] In a first aspect, the present invention provides a method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation, comprising the following steps: The occluded area point cloud compensation step includes: obtaining an endoscopic image sequence comprising multiple frames of images, wherein the endoscopic image sequence includes a color image, a depth map, and a surgical tool mask map for each frame; generating an initialization point cloud with sufficient distribution prior based on the endoscopic image sequence, wherein the initialization point cloud is formed by merging the reference frame point cloud and the compensation point cloud; Two-stage dual-domain deformation steps: construct a set of Gaussian ellipsoids based on the initialized point cloud; perform two-stage training on the Gaussian ellipsoid set to reconstruct the dynamic scene; wherein, the first stage of the two-stage training is Gaussian warm-up training, which keeps the position and rotation properties of each Gaussian ellipsoid unchanged over time; the second stage is dual-domain deformation training, which collaboratively fits the deformation variables of the position and rotation properties of each Gaussian ellipsoid over time by jointly using the time domain model and the frequency domain model.

[0008] As an optional implementation of the first aspect of the present application, the occluded area point cloud compensation step specifically includes: based on the selected reference frame, using the surgical tool mask map to remove the surgical tool from the color image, and combining the depth map and the internal and external parameters of the camera to generate the reference frame point cloud; identifying the tissue area occluded by the surgical tool in the reference frame but visible in one or more other frames in the image sequence, and generating a surgical tool compensation mask map; the surgical tool compensation mask map is obtained by calculating the set difference operation of the surgical tool mask map of the reference frame and the surgical tool mask maps of other frames; based on the surgical tool compensation mask map, the color images, depth maps and internal and external parameters of the camera of the other frames, generating the compensated point cloud; merging the reference frame point cloud and the compensated point cloud, and sampling a predetermined number of points therefrom to form the final initialized point cloud.

[0009] As an optional implementation of the first aspect of the present application, the first-stage Gaussian warm-up training in the two-stage dual-domain deformation step aims to optimize the initial static properties of the Gaussian ellipsoid; the training process keeps the position and rotation quaternion of each Gaussian ellipsoid constant in the time dimension, and only optimizes the initial position, initial rotation, scaling matrix, spherical harmonic function coefficients and opacity of the Gaussian ellipsoid for a predetermined number of iterations to generate a more precisely positioned static Gaussian ellipsoid set around the initialized point cloud.

[0010] As an optional implementation of the first aspect of the present application, the second stage dual-domain deformation training in the two-stage dual-domain deformation step models the position and rotation properties of each Gaussian ellipsoid at any time as the superposition of its basic properties at the reference time and a deformation variable that changes with time; the deformation variable is decomposed into the sum of a time domain deformation variable and a frequency domain deformation variable.

[0011] As an optional implementation of the first aspect of the present application, the fitting method of the time domain deformation variable and the frequency domain deformation variable is: the time domain deformation variable is represented by the weighted sum of one or more Gaussian distribution functions, and the center, variance and weight of the Gaussian distribution function are all learnable parameters for capturing the non-periodic fine deformation of the tissue; the frequency domain deformation variable is represented by a Fourier series, and the sine term coefficient and cosine term coefficient of the Fourier series are all learnable parameters for capturing the periodic or large-scale rough deformation of the tissue.

[0012] As an optional implementation scheme of the first aspect of the present application, the training process of the two-stage dual-domain deformation step is driven by minimizing a total loss function; the total loss function includes a weighted sum of color loss, depth loss and spatial loss; the color loss and depth loss are respectively used to constrain the consistency between the rendered color image and depth map and the true value, and are only calculated in the non-surgical tool mask area; the spatial loss is to apply total variation regularization to the rendered color image and depth map to reduce visual artifacts that may appear in areas occluded by surgical tools and ensure spatial smoothness.

[0013] As an optional implementation of the first aspect of the present application, after generating the compensated point cloud, a preset percentage of point clouds are randomly extracted from the compensated point cloud; after merging the reference frame point cloud and the extracted compensated point cloud, a preset number of points are randomly selected from the merged point cloud as the final initialization point cloud.

[0014] In a second aspect, an embodiment of the present application provides an endoscopic image dynamic reconstruction system based on point cloud compensation and dual-domain deformation, comprising: an occlusion region point cloud compensation module configured to acquire an endoscopic image sequence comprising multiple frames, the endoscopic image sequence including a color image, a depth map, and a surgical tool mask map for each frame; and generate an initialization point cloud with sufficient distribution prior based on the endoscopic image sequence, the initialization point cloud being formed by merging a reference frame point cloud and a compensation point cloud; A two-stage dual-domain deformation module is configured to construct a set of Gaussian ellipsoids based on the initialization point cloud; the Gaussian ellipsoid set is trained in two stages to reconstruct a dynamic scene; wherein the first stage of the two-stage training is Gaussian warm-up training, in which the position and rotation properties of each Gaussian ellipsoid are kept constant over time; and the second stage is dual-domain deformation training, in which the position and rotation properties of each Gaussian ellipsoid are collaboratively fitted with the deformation variables over time by jointly using a time domain model and a frequency domain model.

[0015] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein when the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0016] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0017] Compared with the existing technology, the present invention proposes a method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation, which has the following beneficial effects: 1. The occluded area point cloud compensation mechanism effectively fills the point cloud holes caused by surgical tool removal, obtaining a more evenly distributed initial point cloud with more sufficient prior information, laying a solid foundation for high-quality reconstruction. 2. A two-stage training strategy separates static optimization from dynamic optimization, improving training stability and efficiency. 3. The innovative use of a dual-domain deformation model that combines Gaussian distribution function and Fourier series avoids the time-consuming feature interpolation and MLP decoding process, greatly shortens the model training time, and can achieve high-quality dynamic reconstruction in a short time, meeting the timeliness requirements of clinical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a schematic diagram of the overall architecture of an endoscopic image dynamic reconstruction method based on point cloud compensation and dual-domain deformation proposed by the present invention; Figure 2 Schematic diagram of the process of compensating the occluded area point cloud in the method of the present invention; Figure 3 Schematic diagram of the process of the two-stage dual-domain deformation step in the method of the present invention; Figure 4 Schematic diagram of the structure of an endoscopic image dynamic reconstruction system based on point cloud compensation and dual-domain deformation proposed in the present invention. DETAILED DESCRIPTION

[0019] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0020] The terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the application can be implemented in a sequence other than those illustrated or described here. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated before and after are in a kind of "or" relationship. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically limited.

[0021] Example 1 See also Figure 1 , which is a schematic diagram of the overall architecture of an endoscopic image dynamic reconstruction method based on point cloud compensation and dual-domain deformation provided by an embodiment of the present invention, which mainly includes two steps: point cloud compensation of occluded area and two-stage dual-domain deformation.

[0022] The occluded area point cloud compensation step includes: obtaining an endoscopic image sequence comprising multiple frames of images, wherein the endoscopic image sequence includes a color image, a depth map, and a surgical tool mask map for each frame; generating an initialization point cloud with sufficient distribution prior based on the endoscopic image sequence, wherein the initialization point cloud is formed by merging the reference frame point cloud and the compensation point cloud; Two-stage dual-domain deformation steps: construct a set of Gaussian ellipsoids based on the initialized point cloud; perform two-stage training on the Gaussian ellipsoid set to reconstruct the dynamic scene; wherein, the first stage of the two-stage training is Gaussian warm-up training, which keeps the position and rotation properties of each Gaussian ellipsoid unchanged over time; the second stage is dual-domain deformation training, which collaboratively fits the deformation variables of the position and rotation properties of each Gaussian ellipsoid over time by jointly using the time domain model and the frequency domain model.

[0023] Specifically, in order to remove surgical tools from endoscopic images and quickly obtain a uniformly distributed and sufficient Gaussian point cloud, the present invention proposes a point cloud compensation process for the occluded area as follows: Figure 2 shown.

[0024] During the Gaussian point cloud initialization process, point cloud compensation is performed on the surgical tool occlusion area in the reference frame, and the point cloud is initialized on the reference frame (frame 0) without the surgical tool image to generate the reference frame point cloud. Due to the removal of surgical tools in the reference frame, the reference frame point cloud There is a point cloud missing in the area where the surgical tool was removed. To compensate for this missing point cloud, we further search the remaining frames for the visible part of the area blocked by the surgical tool in the reference frame and initialize the point cloud for this area. Then, we randomly extract 1% from the point cloud generated by the remaining frames as the compensation point cloud for the remaining frames. Finally, the reference frame point cloud Compensate point cloud with remaining frames Combine and randomly select 30,000 points as the final initialization point cloud .

[0025] For a given Frame endoscope color imaging , Depth Map and surgical tools mask illustration First, the base frame color image Remove the corresponding surgical tools in the reference frame and project them into three-dimensional space to obtain the point cloud generated in the reference frame (Benchmark frame point cloud ). Baseline frame point cloud The calculation of can be expressed by formula (1): in, Represents the camera intrinsic parameter matrix, the parameters of which are only related to the camera itself; Represents the extrinsic parameter matrix corresponding to the reference frame image. The parameters of this matrix are only related to the shooting pose of the camera; Indicates the reference frame endoscopic color image, represents the depth map corresponding to the reference frame endoscopic color image, The surgical tool mask corresponding to the reference frame endoscopic color image is represented. Represents the element-wise dot product.

[0026] Since the base frame The removal of surgical tools results in the reference frame point cloud There is a point cloud missing in the area where the surgical tool is removed. However, during the movement of the surgical tool, the tissue occluded by the surgical tool in the reference frame may be visible in other frames. Based on this observation, we next query the surgical tool mask map of the remaining frames. , and collect the masks corresponding to the new tissue pixels appearing in the remaining frames to obtain the surgical tool compensation mask map . Surgical tool compensation mask map The calculation of is shown in formula (2): in, and Represent the surgical tool mask images corresponding to the endoscopic images in the reference frame and the remaining frames, Represents element intersection.

[0027] Then, the surgical tool compensation mask obtained above is Project into 3D space to obtain the residual frame compensated point cloud . The reference frame point cloud Compensate point cloud with remaining frames Merge to get the final initialized point cloud . Remaining frame compensated point cloud And the final initialized point cloud It can be expressed by formula (3) and formula (4) respectively: in, represents the intrinsic parameter matrix of the camera, represents the extrinsic parameter matrix corresponding to the endoscopic color image in the remaining frames, represents the color image of the endoscope in the remaining frames, represents the depth map corresponding to each endoscopic color image in the remaining frames, represents the surgical tool compensation mask obtained by formula (4), represents the dot product of each pixel element, and The base frame initialization point cloud and the remaining frame compensated point cloud The random sampling coefficient of the present invention is set , .

[0028] Finally, the initialized point cloud With each point as the center, a set of Gaussian ellipsoids is created. The process can be expressed by formula (5) and formula (6): in, represents the position of the Gaussian ellipsoid, represents the rotation matrix, represents the scaling matrix, Indicates the center of the Gaussian ellipsoid position.

[0029] Specifically, in order to achieve high-quality deformable tissue reconstruction in a relatively short time, the present invention proposes a two-stage dual-domain deformation to handle the deformation problem during the dynamic reconstruction of endoscopic images. The two-stage dual-domain deformation process proposed by the present invention is as follows: Figure 3As shown. This method is divided into two stages of training: in the first stage of training, 1000 iterations of Gaussian warm-up training are performed to keep the position and rotation of the Gaussian ellipsoid unchanged over time. The purpose of this is to generate some Gaussian ellipsoids with more precise positions near the initialized Gaussian ellipsoid; in the second stage of training, dual-domain deformation is used, that is, Gaussian distribution function is used in the time domain and Fourier series is used in the frequency domain to collaboratively approximate the changes in the position and rotation of the Gaussian ellipsoid over time. In the time domain, the Gaussian distribution function can provide more flexibility for the basis function to adapt to the subtle deformation of some tissues in endoscopic images, and is good at capturing fine vascular structures. In addition, in the frequency domain, the Fourier series is good at capturing larger deformations or tissue appearance contours associated with intense movement. The combination of the two can achieve more accurate and rapid dynamic reconstruction.

[0030] In the first stage of training, Gaussian warm-up training is performed to keep the position and rotation of the Gaussian ellipsoid constant over time. In the second stage of training, the position and rotation properties of the Gaussian ellipsoid that change over time are , , the Gaussian distribution function and Fourier series are used to collaboratively calculate the change of each Gaussian attribute over time. Specifically, the present invention calculates the position and rotation attributes of each Gaussian ellipsoid , Changes over time, described as their value at frame 0 Basic properties , With time The change in attribute By using the Gaussian distribution function in the time domain and the Fourier series in the frequency domain, the position and rotation properties of the Gaussian ellipsoid are fitted to change with time. The calculation is shown in formula (7): in, , Represents the Gaussian ellipsoid at time The position and rotation properties, Represents Gaussian properties over time The amount of change, Depend on and It consists of two parts, namely .in, represents the time domain variation generated by the Gaussian distribution function, Represents the frequency domain variation produced by the Fourier series.

[0031] By using the Gaussian distribution function in the time domain and in the frequency domain using Fourier series To calculate the change of Gaussian attributes over time, it can be expressed by formula (8) and formula (9) respectively: in, represents a learnable center, represents the variance of the Gaussian distribution function, Represents the position and rotation properties of the Gaussian ellipsoid, that is , Represents the learnable weight value assigned to each attribute of the Gaussian ellipsoid, and represent the coefficients of the Fourier series sine and cosine functions, respectively. represents the number of terms in the Fourier series. In the experiment of the present invention, .

[0032] In order to speed up the training of the model, the present invention experiments use color loss and depth loss Supervise the rendered color and depth images using spatial loss to constrain the predicted color and depth maps.

[0033] Calculating color loss and depth loss Previously, the method of the present invention used a method of accumulating pixel color and depth to calculate the color and depth of a pixel. and depth As shown in formula (10) and formula (11) respectively: in, It is by The color is calculated from the SH coefficient of a Gaussian ellipsoid, It is Gaussian ellipsoid Values can be obtained through the covariance matrix Multiply the opacity of the Gaussian ellipsoid by get.

[0034] Color loss for network training and depth loss The calculations are shown in formula (12) and formula (13): in, Represents the rendered color and depth maps, Represent the true color and depth map respectively, Indicates the surgical tool mask corresponding to each frame of endoscopic image, Represents 2D space coordinates.

[0035] The experiment of the present invention adopts space loss To avoid black artifacts in areas occluded by surgical tools, spatial loss Using total variation loss ( To regularize the rendering results, The calculation method is shown in formula (14): In each iteration, the total loss of optimization is calculated as shown in formula (15): in, Indicates the weight used to balance the loss during training. In the experiment of this invention, , , .

[0036] In this embodiment, the datasets used in the present invention are two datasets commonly used in the current endoscopic image dynamic reconstruction task, namely, the EndoNeRF dataset and the Hamlyn dataset.

[0037] The EndoNeRF dataset consists of color images extracted from frames extracted from endoscopic videos captured from a single viewpoint during prostatectomy surgery using the da Vinci robot. The dataset also includes depth maps and surgical tool masks, aiming to reconstruct surgical scenes involving non-rigid deformations and tool occlusions. It includes six sequences with an image resolution of 512 × 640 pixels. Two of these sequences are publicly available: one for tissue cutting, providing 156 frames, and the other for soft tissue stretching, providing 63 frames. The experiments in this paper were conducted on these two publicly available sequences, and the division of the EndoNeRF dataset remains consistent with the original EndoNeRF paper.

[0038] The Hamlyn dataset is a collection of cardiac and in-vivo sequences captured by the da Vinci robot during surgery. The dataset contains rectified images, stereo depth maps, and camera intrinsics. The dataset used in this experiment is derived from seven challenge sequences provided by ForPlane, which exhibit weak textures, deformations, reflections, surgical tool occlusions, and illumination changes. Each sequence contains 301 frames with an image resolution of 480×640 pixels. ForPlane generates a corresponding depth map and surgical tool mask for each sequence and divides the Hamlyn dataset into a training set of 151 frames and a test set of the remaining 150 frames. This division ensures sufficient training data while providing a robust set for evaluation. The division of the Hamlyn dataset in this experiment is consistent with the original ForPlane paper.

[0039] Evaluation indicators: In the task of dynamic reconstruction of endoscopic images, commonly used evaluation indicators include peak signal-to-noise ratio, structural similarity index, and learning-perceived image block similarity.

[0040] (1) Peak Signal-to-Noise Ratio (PSNR) is often used to compare the difference between the input image and the rendered image. The higher the PSNR value, the higher the similarity between the quality of the rendered image and the input image. PSNR represents the spatial structural difference between the images by calculating the mean squared error (MSE) between the input image and the rendered image. The calculation of PSNR is shown in formula (16): Where MAX represents the maximum possible value of a pixel in the image, and the calculation formula of MSE is shown in formula (17): in, The coordinates on the clean image are elements, and The coordinates on the noise image are The element of the image is , Indicates the length of the image, .

[0041] (2) The Structural Similarity Index (SSIM) evaluates from the perspective of image structure, defining structural information as properties that reflect the structure of objects in the scene. SSIM measures image distortion by brightness, contrast, and structure, and can also be used to assess image similarity. A larger SSIM value indicates a higher degree of image similarity, and generally better image quality.

[0042] (3) Learned Perceptual Image Patch Similarity (LPIPS) simulates the human visual system's perception of images and uses a neural network model to approximate the image similarity perceived by the human visual system. The smaller the LPIPS value, the smaller the difference between the two images and the better the quality of the synthesized image.

[0043] It should be noted that the parameters of the present invention are set as follows: the image resolution of the EndoNeRF dataset input is 640×512 pixels, and the image resolution of the Hamlyn dataset input is 640×480 pixels. In the initialization point cloud stage, 1% of the compensation point cloud generated by the remaining frames is randomly selected as the compensation point cloud of the remaining frames. , randomly select 30,000 points from the overall initialization point cloud as the final initialization point cloud Adam is used as the optimizer, and the initial learning rate is 1.6×10 -3 The Gaussian warm-up was performed for 1000 iterations, and the optimization of the entire model framework was performed for 3000 iterations.

[0044] Example 2 See also Figure 4 , shown is a schematic structural diagram of an endoscopic image dynamic reconstruction system based on point cloud compensation and dual-domain deformation proposed in the second embodiment of the present application. The system includes the following key modules: The occlusion region point cloud compensation module 100 is configured to acquire an endoscopic image sequence comprising multiple frames, wherein the endoscopic image sequence includes a color image, a depth map, and a surgical tool mask map for each frame; generate an initialization point cloud with sufficient distribution prior based on the endoscopic image sequence, wherein the initialization point cloud is formed by merging the reference frame point cloud and the compensation point cloud; The two-stage dual-domain deformation module 200 is configured to construct a set of Gaussian ellipsoids based on the initialization point cloud; the Gaussian ellipsoid set is trained in two stages to reconstruct the dynamic scene; wherein the first stage of the two-stage training is Gaussian warm-up training, in which the position and rotation attributes of each Gaussian ellipsoid are kept unchanged over time; the second stage is dual-domain deformation training, which collaboratively fits the deformation variables of the position and rotation attributes of each Gaussian ellipsoid over time by jointly using the time domain model and the frequency domain model.

[0045] In this embodiment, the present invention addresses the issues of insufficient Gaussian initialization point cloud distribution priors and long training times in existing dynamic endoscopic image reconstruction methods based on Gaussian splattering. An occlusion region point cloud compensation module and a two-stage dual-domain deformation module are proposed to alleviate these issues. First, the occlusion region point cloud compensation module initializes the point cloud of a reference frame. Then, by querying other frames, it compensates for the point cloud of areas occluded by surgical tools in the reference frame, thereby obtaining an initial point cloud with a relatively sufficient distribution prior. Second, a two-stage dual-domain deformation module is proposed. In the first stage of training, the position and rotation of the Gaussian ellipsoid are kept constant over time for warm-up training. In the second stage of training, the Gaussian distribution function and Fourier series are used to collaboratively approximate the temporal deformation of the position and rotation of the Gaussian ellipsoid. This deformation calculation method eliminates the time-consuming feature plane interpolation and MLP decoding, thereby achieving faster reconstruction speed. The proposed method outperforms existing methods, reducing training time and GPU usage, and achieving higher reconstruction quality with a training time of approximately one minute.

[0046] In an embodiment of the present application, a dynamic endoscopic image reconstruction system based on point cloud compensation and dual-domain deformation can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a network attached storage (NAS), a personal computer (PC), etc., which is not specifically limited in the embodiment of the present application.

[0047] In the embodiments of the present application, a dynamic endoscopic image reconstruction system based on point cloud compensation and dual-domain deformation can be a device having an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.

[0048] The embodiment of the present application provides an endoscopic image dynamic reconstruction system based on point cloud compensation and dual-domain deformation, which can achieve Figure 1 In order to avoid repetition, the various processes of the method embodiment of the dynamic reconstruction method of endoscopic images based on point cloud compensation and dual-domain deformation are not described here.

[0049] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the various processes of the above-mentioned embodiment of the method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation are implemented, and the same technical effect can be achieved. To avoid repetition, they will not be described here.

[0050] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the embodiment of the above-mentioned method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0051] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.

[0052] It should be noted that, in the present invention, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0053] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a more preferred embodiment. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of this application.

[0054] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation, characterized in that: The following steps are involved: The occluded area point cloud compensation step includes: obtaining an endoscopic image sequence comprising multiple frames of images, wherein the endoscopic image sequence includes a color image, a depth map, and a surgical tool mask map for each frame; generating an initialization point cloud with sufficient distribution prior based on the endoscopic image sequence, wherein the initialization point cloud is formed by merging the reference frame point cloud and the compensation point cloud; Two-stage dual-domain deformation steps: construct a set of Gaussian ellipsoids based on the initialized point cloud; perform two-stage training on the Gaussian ellipsoid set to reconstruct the dynamic scene; wherein, the first stage of the two-stage training is Gaussian warm-up training, which keeps the position and rotation properties of each Gaussian ellipsoid unchanged over time; the second stage is dual-domain deformation training, which collaboratively fits the deformation variables of the position and rotation properties of each Gaussian ellipsoid over time by jointly using the time domain model and the frequency domain model.

2. The method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation according to claim 1, characterized in that: The occluded area point cloud compensation step specifically includes: Based on the selected reference frame, the surgical tool is removed from the color image using the surgical tool mask map, and the reference frame point cloud is generated by combining the depth map and camera internal and external parameters; Identifying tissue regions obscured by the surgical tool in the reference frame but visible in one or more other frames in the image sequence, and generating a surgical tool compensation mask; the surgical tool compensation mask is obtained by calculating a set difference operation between the surgical tool mask of the reference frame and the surgical tool masks of the other frames; generating the compensated point cloud based on the surgical tool compensation mask image, the color images of the other frames, the depth map, and the internal and external parameters of the camera; The reference frame point cloud and the compensated point cloud are merged, and a predetermined number of points are sampled therefrom to form the final initialization point cloud.

3. The method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation according to claim 1, characterized in that: The first stage of Gaussian warm-up training in the two-stage dual-domain deformation step aims to optimize the initial static properties of the Gaussian ellipsoid; The training process keeps the position and rotation quaternion of each Gaussian ellipsoid constant in the time dimension, and only optimizes the initial position, initial rotation, scaling matrix, spherical harmonic function coefficients and opacity of the Gaussian ellipsoid for a predetermined number of iterations to generate a static Gaussian ellipsoid set with more accurate positions around the initialization point cloud.

4. The method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation according to claim 1 or 3, characterized in that: The second stage of dual-domain deformation training in the two-stage dual-domain deformation step models the position and rotation properties of each Gaussian ellipsoid at any time as the superposition of its basic properties at the reference time and a deformation variable that changes with time; the deformation variable is decomposed into the sum of a time domain deformation variable and a frequency domain deformation variable.

5. The method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation according to claim 4, characterized in that: The fitting method of the time domain shape variable and the frequency domain shape variable is: The temporal deformation variable is represented by a weighted sum of one or more Gaussian distribution functions, wherein the center, variance, and weight of the Gaussian distribution function are all learnable parameters for capturing the non-periodic fine deformation of the tissue; The frequency domain deformation variable is represented by a Fourier series, and the sine term coefficient and the cosine term coefficient of the Fourier series are both learnable parameters for capturing the periodicity or large-scale rough deformation of the tissue.

6. The method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation according to claim 1, characterized in that: The training process of the two-stage dual-domain deformation step is driven by minimizing a total loss function; the total loss function includes the weighted sum of color loss, depth loss and spatial loss; The color loss and depth loss are respectively used to constrain the consistency between the color image and depth map generated by rendering and the true value, and are only calculated in the non-surgical tool mask area; The spatial loss is to apply total variation regularization to the rendered color image and depth map to reduce visual artifacts that may appear in areas occluded by surgical tools and ensure spatial smoothness.

7. The method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation according to claim 1, characterized in that: After generating the compensation point cloud, randomly extracting a preset percentage of point clouds from the compensation point cloud; After merging the reference frame point cloud and the extracted compensated point cloud, a preset number of points are randomly selected from the merged point cloud as the final initialization point cloud.

8. An endoscopic image dynamic reconstruction system based on point cloud compensation and dual-domain deformation, characterized in that: include: an occlusion region point cloud compensation module configured to acquire an endoscopic image sequence comprising multiple frames, the endoscopic image sequence including a color image, a depth map, and a surgical tool mask map for each frame; and generate an initialization point cloud with sufficient distribution prior based on the endoscopic image sequence, the initialization point cloud being formed by merging a reference frame point cloud and a compensation point cloud; a two-stage dual-domain deformation module configured to construct a set of Gaussian ellipsoids based on the initialization point cloud; The Gaussian ellipsoid set is trained in two stages to reconstruct dynamic scenes. The first stage of the two-stage training is Gaussian warm-up training, which keeps the position and rotation properties of each Gaussian ellipsoid unchanged over time. The second stage is dual-domain deformation training, which uses a time domain model and a frequency domain model to collaboratively fit the deformation of the position and rotation properties of each Gaussian ellipsoid over time.

9. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation as described in any one of claims 1 to 7 are implemented.

10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the method for dynamic reconstruction of endoscopic images based on point cloud compensation and dual-domain deformation as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Visual compensation method and device, electronic equipment and storage medium

    CN118822910A

  • Scene data optimization reconstruction method, system and product

    CN119169212A

  • Semantic vision SLAM (Simultaneous Localization and Mapping) method based on depth mask segmentation in dynamic environment

    CN119206203A

  • Dynamic scene reconstruction method and device based on multi-scale Gaussian sphere

    CN119991973A

  • Endoscopic surgery scene real-time reconstruction method based on 3D Gaussian

    CN120070755A

Cited By

  • Visual tracking method for processing flexible linear object with partial shielding

    CN121236107A