Methods, systems, media, and equipment for endoscopic pose estimation based on style transfer

By employing a style transfer-based endoscopic pose estimation method, coordinate transformation and virtual camera settings are performed using 3D medical images and bronchial segmentation and reconstruction data. The style transfer network model is then trained, which solves the problem of artifacts in endoscopic pose estimation and improves the accuracy of pose estimation and surgical efficiency.

CN122415448APending Publication Date: 2026-07-17PRECISION ROBOTICS (SHANGHAI) LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PRECISION ROBOTICS (SHANGHAI) LTD
Filing Date
2026-03-23
Publication Date
2026-07-17

Smart Images

  • Figure CN122415448A_ABST
    Figure CN122415448A_ABST
Patent Text Reader

Abstract

This application provides a style transfer-based method, system, medium, and device for endoscopic pose estimation. The method includes: determining a first dataset, a second dataset, and a third dataset. The first dataset includes three-dimensional medical images containing bronchi and three-dimensional segmentation and reconstruction data of the bronchi. The third dataset includes real endoscopic images. A pre-defined style transfer network model M1 is jointly trained using the second and third datasets. Style transfer is performed from the third dataset to the second dataset. Similarity is measured between the second dataset and the style-transferred third dataset. The most similar image in the second dataset to the third dataset is identified, and the depth and pose estimates corresponding to the real endoscopic images are determined. This application reduces artifact effects, improves the accuracy of endoscopic pose estimation, and enhances the accuracy and efficiency of surgical intervention to reach the target lesion area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and more specifically, to a method, system, medium, and device for endoscopic pose estimation based on style transfer. Background Technology

[0002] Inspired by the success of autonomous driving, vision-based technologies for surgical navigation have gained clinical attention. These technologies eliminate the need for additional electromagnetic or optical tracking instruments and facilitate further miniaturization of endoscopic devices to reach narrow, distal cavities. In endoscopic interventions, a key application is the combination of preoperative images (such as CT) with intraoperative endoscopic images for position and orientation estimation. In this context, preoperative multi-slice CT or MRI is used to reconstruct a 3D virtual scene for surgical planning and intraoperative navigation. This is achieved by registering real-time 2D endoscopic video with a 3D map, enabling accurate restoration of the instrument tip's position and orientation while taking into account physiological motion and tissue deformation. However, alignment between the two imaging domains presents significant challenges, particularly in the absence of readily apparent surface textures or anatomical landmarks, due to artifacts such as fluid, bleeding, motion blur, and foreign bodies.

[0003] In the patent “CN117671012B; A method, apparatus and device for calculating the absolute and relative pose of an endoscope during surgery”, the patient's in-body image data captured by the endoscope system during the RMIS process is collected, and the endoscope pose estimation model is trained using a preprocessed virtual dataset and a real-world-simulated dataset. The patient's in-body image data is then input into the trained endoscope pose estimation model to obtain the endoscope absolute pose data corresponding to the real-time image. Based on the endoscope absolute pose, the real-time relative pose of the endoscope in the real-time image is calculated. However, the problem of image artifacts still exists, which affects the accuracy of absolute pose estimation when the endoscope pose estimation model is used to estimate the absolute pose of the real-time image.

[0004] To address this challenge, deep learning-based methods, especially domain adaptation techniques, have made significant progress in recent years. Existing methods mainly focus on feature-level domain adaptation and image-level domain adaptation. Feature-level domain adaptation involves mapping images of different modalities to a shared latent space to extract domain-invariant features for subsequent analysis tasks; while image-level domain adaptation, or image translation, transfers data from a given source domain to a target domain.

[0005] In Generative Adversarial Networks (GANs), the generator produces images that resemble the target from a source domain image, while the discriminator measures the similarity between the generated and target domains. Unlike GANs, Energy Models (EBMs) directly model the probability distribution of the data. EBMs learn an energy function that assigns lower energy to reasonable images and higher energy to unreasonable images.

[0006] However, in real-world surgical scenarios, style transfer is heavily influenced by complex artifacts. Style transfer of endoscopic images is extremely difficult due to unpredictable artifacts such as blood, mucus, and bubbles. These unpredictable artifacts complicate the direct mapping of virtual images to real-world scenes. Summary of the Invention

[0007] In view of the deficiencies in the prior art, the purpose of this application is to provide an endoscope pose estimation method, system, medium and device based on style transfer.

[0008] A first aspect of this application provides a style transfer-based method for estimating endoscopic pose, comprising: The first part of the dataset is obtained by acquiring three-dimensional medical images containing bronchi and three-dimensional segmentation and reconstruction data of bronchi. Based on the first part of the dataset, perform physical information-based coordinate transformation and virtual camera settings to determine the second part of the dataset; Acquire real endoscopic images of the bronchi as the third part of the dataset; A preset style transfer network model M1 is trained using the second part of the data and the third part of the dataset, and the style transfer network model M1 trained by the model is used to transfer the style from the third part of the dataset to the second part of the data. A similarity measure is performed on the second part of the dataset and the third dataset after style transfer. The image most similar to the third part of the dataset is determined in the second part of the dataset, and the depth and pose information of the most similar image is determined as the depth and pose estimates corresponding to the real endoscope.

[0009] Optionally, the acquisition of three-dimensional medical images containing the bronchi and three-dimensional segmentation and reconstruction data of the bronchi, as the first part of the dataset, includes: Acquire three-dimensional medical images containing bronchi; The three-dimensional medical image containing the bronchi is segmented to determine the three-dimensional segmentation and reconstruction data of the bronchi. The three-dimensional medical image containing the bronchi and the three-dimensional segmentation and reconstruction data of the bronchi are used as the first part of the dataset.

[0010] Optionally, the second part of the dataset includes a virtual image of the bronchus, a virtual depth corresponding to the virtual image of the bronchus, and a virtual pose of the virtual camera.

[0011] Optionally, determining the second dataset based on the first dataset by performing physical information-based coordinate transformation and virtual camera settings includes: The three-dimensional medical image containing the bronchi is segmented using a preset segmentation model to determine the bronchus segmentation model. The bronchial segmentation model is rendered using a rendering isometry algorithm to generate three-dimensional mesh data; The three-dimensional mesh data is transformed based on physical information to convert it to the world coordinate system, thereby determining the three-dimensional mesh data in the world coordinate system. Centerlines are extracted from the three-dimensional mesh data of the world coordinate system to determine the three-dimensional world coordinates of each centerline; A virtual camera is set at each sampling point of each centerline; The virtual depth map of the bronchus, the virtual depth corresponding to the virtual image of the bronchus, and the virtual pose of the virtual camera are obtained by using a virtual camera at each sampling point of each centerline, and are used as the second part of the dataset.

[0012] Optionally, the preset style transfer network model M1 includes a local denoising network model N1 and a global style transfer model N2. The local denoising model N1 includes a first encoder and a first decoder, and the global style transfer model N2 includes a second encoder and a second decoder.

[0013] Optionally, the step of using the second part of the data and the third part of the dataset to train a preset style transfer network model M1, and determining the trained style transfer network model M1, includes: The first encoder is used to extract features from the real endoscopic images of the bronchus in the third part of the dataset to determine the first implicit image features; The first decoder is used to decode the first implicit image features to determine the real endoscopic image of the bronchus without artifacts, and to determine the local denoising network model N1 trained by the model. The second encoder is used to extract features from the artifact-free real endoscopic image of the bronchus to determine the second implicit image features; The second decoder is used to decode the second implicit image features to determine the real endoscopic image of the bronchus corresponding to the style of the second part of the dataset, and to determine the global style transfer model N2 trained by the model. Based on the local denoising network model N1 trained by the model and the global style transfer model N2 trained by the model, the style transfer network model M1 trained by the model is determined.

[0014] Optionally, the step of training a preset style transfer network model M1 using the second part of the data and the third part of the dataset to determine the trained style transfer network model M1 further includes: Determine pseudo-labels for training the local denoising network model N1.

[0015] Optionally, determining the pseudo-labels for training the local denoising network model N1 includes: The pre-trained feature extraction network M2 is used to extract features from the second part of the dataset and the third part of the dataset after style transfer, to determine the first implicit feature and the second implicit feature. The first implicit feature and the second implicit feature are similar to each other by cosine similarity to determine the virtual image of the bronchus, the depth of the virtual image of the bronchus, and the virtual pose of the virtual camera in the second part of the dataset. The depth corresponding to the virtual image of the bronchus in the second part of the dataset corresponding to the second implicit feature is determined as the depth corresponding to the real endoscopic image of the bronchus in the third part of the dataset, and the virtual pose of the virtual camera corresponding to the second implicit feature in the second part of the dataset is determined as the pose estimate of the real endoscope.

[0016] Optionally, the step of training a preset style transfer network model M1 using the second part of the data and the third part of the dataset to determine the trained style transfer network model M1 further includes: The style transfer network model M1, after training, is optimized using an overall loss function, which is: in, Represents the overall loss function, This represents the objective function for adversarial training optimization of the transferred image. This represents the adversarial training constraints imposed on the inverse generation network. This represents a consistency constraint on the forward-generating network. This represents the robust feature extraction loss function guided by virtual images.

[0017] Optionally, the step of measuring the similarity between the second dataset and the style-transferred third dataset, determining the image in the second dataset that is most similar to the third dataset, and determining the depth and pose information of the most similar image as the depth and pose estimates corresponding to the real endoscope, includes: The pre-trained feature extraction network M2 is used to extract features from the second part of the dataset and the third part of the dataset after style transfer, to determine the first implicit feature and the second implicit feature. The first implicit feature and the second implicit feature are used to measure feature similarity using cosine similarity to determine the virtual image of the bronchus, the depth corresponding to the virtual image of the bronchus, and the virtual pose of the virtual camera in the second part of the dataset corresponding to the second implicit feature. The depth corresponding to the virtual image of the bronchus in the second part of the dataset corresponding to the second implicit feature is determined as the depth corresponding to the real endoscopic image of the bronchus in the third part of the dataset, and the virtual pose of the virtual camera corresponding to the second implicit feature in the second part of the dataset is determined as the pose estimate of the real endoscope.

[0018] A second aspect of this application provides an endoscope pose estimation system based on style transfer, comprising: The first part of the dataset determination module is used to acquire three-dimensional medical images containing bronchi and three-dimensional segmentation and reconstruction data of bronchi, which are used as the first part of the dataset. The second part of the dataset determination module is used to determine the second part of the dataset by performing coordinate transformation based on physical information and setting up virtual cameras according to the first part of the dataset. The third part of the dataset determination module is used to acquire the actual endoscopic images of the bronchus as the third part of the dataset. The style transfer module is used to train a preset style transfer network model M1 using the second part of the data and the third part of the dataset, and to use the style transfer network model M1 trained by the model to perform style transfer from the third part of the dataset to the second part of the data. The endoscope pose estimation module is used to measure the similarity between the second part of the dataset and the style-transferred third dataset, identify the image in the second part of the dataset that is most similar to the third part of the dataset, and determine the depth and pose information of the most similar image as the depth and pose estimates corresponding to the real endoscope.

[0019] A third aspect of this application provides a non-transitory computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method provided in the first aspect of this application.

[0020] A fourth aspect of this application provides an electronic device comprising: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method provided in the first aspect of this application.

[0021] Compared with the prior art, the embodiments of this application have at least one of the following beneficial effects: This application provides a style transfer-based endoscopic pose estimation method. Based on preoperative 3D medical images and bronchial 3D segmentation and reconstruction data, it performs physical information-based coordinate transformation and virtual camera settings to obtain a second dataset. A real endoscopic image is then acquired as a third dataset. A pre-defined style transfer network trained on the second and third datasets is used to perform style transfer from the real endoscopic image in the third dataset to the second dataset. This references the virtual image of the bronchus to estimate the endoscopic pose during surgery. This method can suppress artifacts, establish a connection between preoperative 3D medical images and real endoscopic images, improve the accuracy of endoscopic pose estimation, and enhance the efficiency and accuracy of surgeons reaching the target lesion area during existing surgical procedures. Attached Figure Description

[0022] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart illustrating a style transfer-based endoscopic pose estimation method according to an exemplary embodiment.

[0023] Figure 2 This is a schematic diagram illustrating the working process of a style transfer network model M1 trained according to an exemplary embodiment.

[0024] Figure 3 This is a schematic diagram illustrating a comparison of the effects of style transfer and pose estimation according to an exemplary embodiment.

[0025] Figure 4 This is a schematic diagram illustrating the structure of an endoscope pose estimation system based on style transfer, according to an exemplary embodiment. Detailed Implementation

[0026] The present application will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any way. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present application. These all fall within the protection scope of this disclosure and the present application.

[0027] The terms "comprising" and "having," and any variations thereof, in the embodiments of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or devices.

[0028] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature.

[0029] It should be noted that all information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with relevant regulations.

[0030] In surgical navigation, to register real-time intraoperative endoscopic images and preoperative images, such as CT scans, and achieve style transfer, existing style transfer methods include deep learning-based methods, generative adversarial network methods, and energy model methods. However, in actual surgical scenarios, style transfer is strongly affected by complex artifacts, making style transfer of endoscopic images difficult and affecting the registration of intraoperative endoscopic images and preoperative images. To address these issues, this application provides an endoscopic pose estimation method based on style transfer to solve the aforementioned problems.

[0031] Figure 1 This is a flowchart illustrating a style transfer-based endoscopic pose estimation method according to an exemplary embodiment. Figure 2 This is a schematic diagram illustrating the working process of a style transfer network model M1 trained according to an exemplary embodiment.

[0032] Reference Figure 1 , Figure 2 As shown, this application provides an endoscope pose estimation method based on style transfer, including S11 to S15.

[0033] S11, acquire three-dimensional medical images containing bronchi and three-dimensional segmentation and reconstruction data of bronchi, as the first part of the dataset.

[0034] Among them, the three-dimensional medical images containing the bronchi can be obtained from CT scan image data containing the bronchi.

[0035] The first part of the data includes three-dimensional medical images containing the bronchi and bronchial segmentation data, namely CT scan images containing the bronchi and bronchial segmentation data.

[0036] S12, perform coordinate transformation based on physical information and virtual camera settings based on the first part of the dataset to determine the second part of the dataset.

[0037] The second part of the dataset includes virtual images of the bronchi, the depth corresponding to the virtual images of the bronchi, and the virtual pose of the virtual camera.

[0038] Each virtual camera captures a virtual image. The virtual depth map of the bronchus is the virtual image captured by the virtual camera, and the depth corresponding to the virtual image of the bronchus is the depth at which the virtual camera is located.

[0039] S13, Obtain real endoscopic images of the bronchi as the third part of the dataset.

[0040] Among them, the actual endoscopic images of the bronchus are image data collected by the endoscope during actual surgery.

[0041] S14, a pre-set style transfer network model M1 is trained using the second part of the data and the third part of the data set, and the style transfer network model M1 trained by the model is used to transfer the style from the third part of the data set to the second part of the data set.

[0042] S15, perform similarity measurement on the second part of the dataset and the style-transferred third dataset, identify the image in the second part of the dataset that is most similar to the third part of the dataset, and determine the depth and pose information of the most similar image as the depth and pose estimates corresponding to the real endoscope.

[0043] The above technical solution, based on preoperative 3D medical images and bronchial 3D segmentation and reconstruction data, performs physical information-based coordinate transformation and virtual camera settings to obtain a second dataset. A real endoscopic image is then acquired as the third dataset. A pre-defined style transfer network, trained using the second and third datasets, is used to transfer style from the real endoscopic image in the third dataset to the second dataset. This references the virtual image of the bronchus to estimate the endoscope's posture during surgery. This approach suppresses artifacts, establishes a connection between preoperative 3D medical images and real endoscopic images, improves the accuracy of endoscopic posture estimation, and enhances the efficiency and accuracy of surgeons reaching the target lesion area during current surgical procedures.

[0044] In order to obtain the first part of the dataset, in some specific embodiments of this application, for S11, three-dimensional medical images containing bronchi and three-dimensional segmentation and reconstruction data of bronchi are obtained as the first part of the dataset, which can be implemented by S111 to S113.

[0045] S111, acquire a three-dimensional medical image containing the bronchi.

[0046] Three-dimensional medical images containing the bronchi can be obtained using CT scan image data.

[0047] S112, perform segmentation processing on the three-dimensional medical image containing the bronchi to determine the three-dimensional segmentation and reconstruction data of the bronchi.

[0048] Specifically, the segmentation process can employ a pre-trained segmentation model. For example, this method uses a pre-trained nnUnet as the bronchus segmentation model.

[0049] Bronchial 3D segmentation and reconstruction data represents the 3D data of the bronchi segmented from a 3D medical image containing the bronchi, and the reconstruction of the 3D data.

[0050] S113, the three-dimensional medical images containing the bronchi and the three-dimensional segmentation and reconstruction data of the bronchi are used as the first part of the dataset.

[0051] In the above embodiments of this application, a first part of the dataset is obtained, which includes a three-dimensional medical image containing the bronchus and three-dimensional segmentation and reconstruction data of the bronchus, as a preoperative three-dimensional medical image.

[0052] In some specific embodiments of this application, the second part of the dataset includes virtual images of the bronchi, virtual depth corresponding to the virtual images of the bronchi, and virtual pose of the virtual camera.

[0053] In order to obtain a virtual image of the bronchi, in some specific embodiments of this application, for S12, the second part of the dataset is determined by performing coordinate transformation based on physical information and setting up a virtual camera according to the first part of the dataset, which can be done using S121 to S126.

[0054] S121, a preset segmentation model is used to segment the three-dimensional medical image containing the bronchi to determine the bronchus segmentation model.

[0055] Specifically, the preset segmentation model can use nnUnet as the bronchial segmentation model. S122 uses a rendering isometry algorithm to render the bronchial segmentation model and generate three-dimensional mesh data.

[0056] S123 performs a coordinate transformation based on physical information on the 3D mesh data, converting the 3D mesh data to the world coordinate system, and determining the 3D mesh data in the world coordinate system.

[0057] Specifically, the physical information of coordinate transformation based on physical information includes the voxel origin, voxel spacing, and voxel arrangement direction of the pixel.

[0058] In this embodiment, the image coordinates corresponding to the three-dimensional mesh data are converted into world coordinates by using the voxel origin, voxel spacing and voxel arrangement direction of the pixels in the three-dimensional medical image containing the bronchi. This converts the three-dimensional mesh data into the world coordinate system.

[0059] S124 extracts the centerlines from the 3D grid data in the world coordinate system and determines the 3D world coordinates of each centerline.

[0060] Among them, the three-dimensional world coordinates of the centerline of the three-dimensional mesh data in the world coordinate system can be extracted using the preset segmentation model mentioned above.

[0061] S125, a virtual camera is set at each sampling point on each centerline.

[0062] In this context, the position vector of each virtual camera represents the three-dimensional world coordinates of each centerline, and the direction vector of the virtual camera points to the next centerline sampling point on its own centerline.

[0063] S126. The virtual depth map of the bronchus, the virtual depth corresponding to the virtual image of the bronchus, and the virtual pose of the virtual camera are obtained by using a virtual camera at each sampling point of each center line, as the second part of the dataset.

[0064] For each virtual camera, a virtual image is captured, and the virtual depth and virtual pose of the virtual camera corresponding to the virtual image are obtained as the second part of the dataset.

[0065] In the above embodiments of this application, by performing coordinate transformation based on physical information and setting up a virtual camera on the first part of the dataset, virtual images, virtual depth corresponding to the virtual images, and virtual pose of the virtual camera are obtained as intermediate quantities for style transfer between preoperative three-dimensional medical images and real endoscopic images of the bronchus during surgery.

[0066] Reference Figure 2 As shown in some specific embodiments of this application, the preset style transfer network model M1 includes a local denoising network model N1 and a global style transfer model N2.

[0067] The local denoising network model N1 includes a first encoder E1 and a first decoder G1.

[0068] The global style transfer model N2 includes a second encoder E2 and a second decoder G2.

[0069] Specifically, the local denoising network model N1 represents the local encoder. The first decoder G1 represents the local decoder. The second encoder, E2, represents the global encoder. The second decoder G2 represents the global decoder. .

[0070] In order to train the preset style transfer network model M1, in some specific embodiments of this application, S14, the preset style transfer network model M1 is trained using the second part of the data and the third part of the dataset, and the style transfer network model M1 after model training is determined, which may include S141 to S145.

[0071] S141, the first encoder E1 is used to extract features from the real endoscopic images of the bronchus in the third part of the dataset to determine the first implicit image features.

[0072] S142, the first decoder G1 is used to decode the first implicit image features to determine the real endoscopic image of the bronchus without artifacts, and to determine the local denoising network model N1 trained by the model.

[0073] S143, the second encoder E2 is used to extract features from the real endoscopic image of the bronchus without artifacts to determine the second implicit image features.

[0074] S144, the second decoder G2 is used to decode the second implicit image features to determine the real endoscopic image of the bronchus corresponding to the style of the second part of the dataset, and to determine the global style transfer model N2 trained by the model.

[0075] Among them, the real endoscopic images of the bronchi corresponding to the style of the second part of the dataset are real endoscopic images of the bronchi with a style similar to that of the second part of the dataset.

[0076] S145. Based on the local denoising network model N1 trained by the model and the global style transfer model N2 trained by the model, determine the style transfer network model M1 trained by the model.

[0077] The style transfer network model M1 trained by the model in this application is represented as follows: in, This represents the style transfer network model trained by the model. This represents a local denoising network model. This represents a global style transfer model. Indicates the input image. This represents a composite function mapping.

[0078] Reference Figure 2 As shown, the working process of the style transfer network model M1 trained by this application is as follows: Input the disturbed image into the local encoder Perform feature extraction to determine the features ; Special Input Local Decoder Decoding is performed to determine a clean and realistic image; Input a clean, realistic image into the global encoder Feature extraction is performed to determine the characteristics. ; Special Input Global Decoder Decode the image to determine a clean virtual image; Input a clean virtual image into the global encoder Feature extraction is performed to determine the characteristics. ; Special Input Global Decoder Decoding is performed to determine a clean and realistic image; Input a clean, realistic image into the local encoder Perform feature extraction to determine the features ; Special Input Local Decoder The image is decoded to identify the interfered image.

[0079] In order to train the preset style transfer network model M1, in some specific embodiments of this application, S14, the preset style transfer network model M1 is trained using the second part of the data and the third part of the dataset, and may also include S146.

[0080] S146, determine the pseudo-labels to be used for training the local denoising network model N1.

[0081] In this application, step S146 is performed before step S141.

[0082] Specifically, in S146, pseudo-labels are determined for training the local denoising network model N1, which may include S101 to S103.

[0083] S101, the second part of the dataset is sampled to determine the virtual image of the sampled bronchus, the depth corresponding to the virtual image of the bronchus, and the virtual pose of the virtual camera.

[0084] S102, Gaussian sputtering model is modeled based on the sampled virtual image of the bronchus, the depth corresponding to the virtual image of the bronchus, and the virtual pose of the virtual camera, and the Gaussian sputtering model is determined.

[0085] Specifically, the parameters of the Gaussian sputtering model include position. Scale factor Rotation Quaternions ,transparency and spherical harmonic coefficients .

[0086] S103 performs feature transfer processing on the Gaussian sputtering model, optimizes the spherical harmonic coefficients of the Gaussian sputtering model, and determines the pseudo-labels used for training the local denoising network model N1.

[0087] In the process of optimizing the Gaussian sputtering model through feature transfer processing, only the spherical harmonic coefficients of the Gaussian sputtering model are optimized, while other features are retained.

[0088] In order to train the preset style transfer network model M1, in some specific embodiments of this application, the preset style transfer network model M1 is trained using the second part of the data and the third part of the dataset, and the style transfer network model M1 after model training is determined, and may also include S147.

[0089] S147 uses the overall loss function to optimize the style transfer network model M1 after model training.

[0090] The overall loss function is: in, Represents the overall loss function. This represents the objective function for adversarial training optimization of the transferred image. This represents the adversarial training constraints on the inverse generator network. This represents a consistency constraint on the forward-generating network. This represents the robust feature extraction loss function guided by virtual images.

[0091] In this application, The images generated by the style transfer network model M1 are similar in style to those in the second part of the dataset.

[0092] Adversarial training to optimize the objective function of transfer images Represented as: in, Let E represent the objective function for adversarial training optimization of the transferred image, and G represent the encoder and decoder, respectively. This represents a virtual depth map of the bronchi in the second part of the image dataset. S This represents the second part of the dataset. t This represents the actual endoscopic images of the bronchi in the third part of the dataset. T This represents the third part of the dataset.

[0093] For the adversarial training constraints of the inverse generative network, the constraint is the reverse generation process, which takes the third part of the dataset T as the input target.

[0094] Consistency constraints on forward-generating networks Represented as: For consistency constraints The forward and backward generation networks produce the same image. Therefore, the third part of the dataset generates the same image result after going through both the forward and backward generation networks, thus constraining consistency. The requirement is to provide an example image that, after passing through both the forward and inverse generation networks, closely approximates the input image.

[0095] Robust Feature Extraction Loss Function Guided by Virtual Images It is used to remove artifacts from intraoperative images. However, artifact features from noisy images can lead to suboptimal style transfer results. We use the clean and clear data features from the second part of the dataset to guide feature extraction of real endoscopic images of the bronchus from the third part of the dataset, which contains artifacts.

[0096] In this application, a contrastive learning strategy is adopted, using the noise-free features of virtual images in the second part of the dataset to extract noise-resistant robust features from noisy input images in the third part of the dataset.

[0097] For noisy input images and corresponding virtual images First, extract M small patches from the two images, each patch having a size of M. Secondly, implicit features are extracted using their respective encoder backbones. Since the patch locations are the same in the virtual image and the noisy input image, the structural and semantic information of corresponding patches in the two images remains consistent. Therefore, any given noise patch features... Patch features at the same location as the corresponding virtual image The same, but with patch features at other locations in the noisy image. They are different. Represented as: in, This represents hyperparameters.

[0098] To test the style transfer network model M1 trained by the model, in some specific embodiments of this application, a style transfer-based endoscope pose estimation method may further include: The style transfer network model M1 trained by the model was tested using the test set of the third part of the dataset.

[0099] In some specific embodiments of this application, this application uses a dataset of 30,562 frames collected in a clinical study involving 25 patients as the dataset for training, validating, and testing the preset style transfer network model M1, and divides it into a training set, a validation set, and a test set. The training set includes 20,000 frames of data, the validation set includes 5,000 frames of data, and the test set includes 5,562 frames of data.

[0100] For the true pose values ​​of the endoscopes in the dataset, a coarse relative pose estimation algorithm is applied to obtain a coarse pose, and fine adjustments are made by manually aligning the structure between the real image and the virtual rendered image.

[0101] The size of the endoscope image to be tested is 3 400 400 pixels.

[0102] The preset style transfer network model M1 was built on the PyTorch 1.12.0 platform, and the computing hardware consisted of one Nvidia GTX 3090 GPU.

[0103] The pre-defined style transfer network model M1 was trained using the training set. During the training process, the optimizer was Adam, the initial learning rate was 1e-4, the batch size was set to 8, and the training process lasted for 100 generations.

[0104] The style transfer network model M1 trained by the model was validated using a validation set.

[0105] The style transfer network model M1 trained on the model was tested using a test set.

[0106] To achieve endoscope pose estimation, in some specific embodiments of this application, for S15, a similarity measurement is performed on the second part of the dataset and the third dataset after style transfer, the image most similar to the third part of the dataset is determined in the second part of the dataset, and the depth and pose information of the most similar image is determined as the depth and pose estimation value corresponding to the real endoscope, which can be implemented using S151 to S153.

[0107] S151, a pre-trained feature extraction network M2 is used to extract features from the second part of the dataset and the third part of the dataset after style transfer, and the first implicit feature F2 and the second implicit feature F3 are determined.

[0108] S152, the first implicit feature F2 and the second implicit feature F3 are similar to each other by cosine similarity to determine the virtual image of the bronchus corresponding to the second implicit feature F2 in the second part of the dataset, the depth corresponding to the virtual image of the bronchus, and the virtual pose of the virtual camera.

[0109] Specifically, the feature similarity of the first implicit feature F2 and the second implicit feature F3 can be measured by using cosine similarity. The data most similar to each data in the third part of the dataset is searched in the second part of the dataset. The virtual image of the bronchus, the depth corresponding to the virtual image of the bronchus, and the virtual pose of the virtual camera are determined as the virtual features of the second implicit feature in the second part of the dataset.

[0110] S153, the depth corresponding to the virtual image of the bronchus in the second implicit feature F3 in the second part of the dataset is determined as the depth corresponding to the real endoscope image of the bronchus in the third part of the dataset, and the virtual pose of the virtual camera corresponding to the second implicit feature in the second part of the dataset is determined as the pose estimate of the real endoscope.

[0111] Based on steps S151 to S153, for example, the images in the third part of the dataset are real endoscopic images I1 of the bronchus with noise; the pre-trained feature extraction network M2 can employ the encoder of the image retrieval network R2former. .

[0112] The encoder using the R2former image retrieval network All virtual images in the second part of the dataset Perform feature extraction to determine the first implicit feature. The encoder of the image retrieval network R2former is used. Feature extraction is performed on the current input image I1 in the third part of the dataset to determine the second implicit features. .

[0113] The similarity between the first implicit feature and the second implicit feature is calculated using cosine similarity: In the second part of the dataset, the i-th virtual image with the highest similarity is selected as the target data, thereby determining the virtual depth map of the bronchus corresponding to the current input image I1 in the third part of the dataset, the depth corresponding to the virtual image of the bronchus, and the virtual pose of the virtual camera.

[0114] The preferred features in the above embodiments can be used individually in any embodiment, or in any combination thereof, provided they do not conflict with each other. Furthermore, parts not described in detail in the embodiments can be implemented using existing technologies.

[0115] The following examples and comparative examples will be used to further illustrate this application in order to better understand the above-mentioned technical solutions. It should be understood that the following are only some examples and are not intended to limit this application.

[0116] Figure 3 This is a schematic diagram illustrating a comparison of the effects of style transfer and pose estimation according to an exemplary embodiment.

[0117] The style transfer network model M1 trained by the model was tested using the test set mentioned above.

[0118] Reference Figure 3 As shown, under the same experimental conditions, the style transfer network model M1 provided in this application, along with CycleGAN, CEP, CUT, AI-copilot, and UNSB, were used to perform style transfer and endoscope pose estimation on the data in the test set.

[0119] Reference Figure 3 As shown, by using the style transfer network model M1 provided in this application and the style transfer-based endoscope pose estimation method of this application, images that are closer to the target style can be generated better, and lower pose errors can be obtained.

[0120] Figure 3In the image (a), we see virtual images generated by different style transfer methods and the matching results obtained through the registration algorithm, with artifact interference. Figure 3 (b) in the figure represents a comparison of the average trajectory errors of different style transfer methods on the dataset. Figure 3 In the figure, (c) represents the error distribution between the method provided in this application and CUT at each time point on the two real trajectories.

[0121] In another possible embodiment, the style transfer network model M1 provided in this application, along with CycleGAN, CEP, CUT, AI-copilot, and UNSB, are used to perform quantitative error analysis on the data in the test set. The evaluation metrics are Fréchet Inception Distance (FID) and Kernel Inception Distance (KID) to quantify the quality difference between the generated images and the real images. In this application, these metrics can be used to quantify the quality difference between the virtual images and the real images in pose estimation.

[0122] The FID (Fixed-Indicator) metric aims to measure the difference in distribution between generated and real images in feature space; in this application, it refers to the difference in pose distribution between virtual and real images. Its calculation process is as follows: Feature extraction: A pre-trained Inception v3 network is used to extract features from both real and generated images. Typically, feature vectors are extracted in the penultimate layer of the model, and these feature vectors can effectively represent the visual features of the image.

[0123] Distribution modeling: Calculate the mean μ and covariance matrix Σ of the extracted feature vectors to obtain the feature distributions of the real image and the generated image, respectively. Let the mean and covariance of the real image be... The mean and covariance of the generated images are .

[0124] FID distance calculation: in, This represents the FID indicator. This represents the squared difference between the mean of the real image and the generated image. Represents the trace operation of a matrix. The square root of the covariance matrix of the real image and the generated image.

[0125] The lower the FID value, the higher the quality of the generated image.

[0126] The KID metric aims to evaluate the quality of generated images using unbiased estimation and kernel methods, thereby improving the robustness and accuracy of the evaluation. Its calculation process is as follows: Feature extraction: A pre-trained Inception v3 network is used to extract features from both real and generated images. Typically, feature vectors are extracted in the penultimate layer of the model, and these feature vectors can effectively represent the visual features of the image.

[0127] Kernel method: Employs a kernel-based metric to calculate the distance between features of the generated image and the real image.

[0128] KID value calculation: in, This refers to the KID metric. Represents the kernel function. Features representing a true image This represents the features of the generated image.

[0129] The lower the KID value, the higher the similarity between the generated image and the real image, and the better the quality of the generated image.

[0130] Table 1 As shown in Table 1, the style transfer-based endoscope pose estimation method provided in this application achieves good transfer performance across all style transfer metrics, achieving an FID of 141.983 and a KID of 0.118 on the test set. Compared to other methods, the style transfer network model M1 provided in this application exhibits better robustness and stability in real-world scenarios. In terms of the pose estimation metric Translation ATE, the style transfer-based endoscope pose estimation method proposed in this application demonstrates higher accuracy, achieving a mean regularity error of 8.87 mm.

[0131] Figure 4 This is a block diagram illustrating a style transfer-based endoscopic pose estimation system according to an exemplary embodiment.

[0132] Based on the same concept, this application also provides an endoscope pose estimation system 1000 based on style transfer, referring to... Figure 4 As shown, it includes: a first part dataset 1100, a second part dataset determination module 1200, a third part dataset determination module 1300, a style transfer module 1400, and an endoscope pose estimation module 1500.

[0133] The first part of the dataset determination module 1100 is used to acquire three-dimensional medical images containing bronchi and three-dimensional segmentation and reconstruction data of bronchi as the first part of the dataset. The second part of the dataset determination module 1200 is used to determine the second part of the dataset by performing coordinate transformation based on physical information and setting up virtual cameras according to the first part of the dataset. The third part of the dataset determination module 1300 is used to acquire real endoscopic images of the bronchus as the third part of the dataset. Style transfer module 1400 is used to train a preset style transfer network model M1 using the second part of the data and the third part of the dataset, and to use the style transfer network model M1 trained by the model to perform style transfer from the third part of the dataset to the second part of the data. The endoscope pose estimation module 1500 is used to measure the similarity between the second part of the dataset and the style-transferred third dataset. It identifies the image in the second part of the dataset that is most similar to the third part of the dataset and uses the depth and pose information of the most similar image as the depth and pose estimates corresponding to the real endoscope.

[0134] The embodiments described above in this application, based on preoperative 3D medical images and bronchial 3D segmentation and reconstruction data, perform physical information-based coordinate transformation and virtual camera settings to obtain a second dataset, and obtain real endoscopic images as a third dataset. A pre-set style transfer network trained on the second and third datasets is used to perform style transfer from the real endoscopic images in the third dataset to the second dataset. This references the virtual images of the bronchus to estimate the posture of the endoscope during the operation, which can suppress images caused by artifacts, establish a connection between preoperative 3D medical images and real endoscopic images, improve the accuracy of endoscopic posture estimation, and improve the efficiency and accuracy of doctors and experts reaching the target lesion area during existing surgical procedures.

[0135] Regarding the embodiments of the above system, the specific ways in which each module performs operations have been described in detail in the embodiments of the method, and will not be elaborated here.

[0136] Based on the same technical concept, in some specific embodiments of this application, a terminal includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and a method that the processor can use to execute when executing the program.

[0137] Based on the same technical concept, in some specific embodiments of this application, a computer-readable storage medium is provided on which a computer program is stored, which can be used to execute a method when the program is executed by a processor.

[0138] Optionally, the memory is used to store programs; the memory may include volatile memory, such as random-access memory (RAM), such as static random-access memory (SRAM), double data rate synchronous dynamic random-access memory (DDR SDRAM), etc.; the memory may also include non-volatile memory, such as flash memory. The memory is used to store computer programs (such as application programs and functional modules that implement the above methods), computer instructions, etc., and the aforementioned computer programs and computer instructions can be partitioned and stored in one or more memories. Furthermore, the aforementioned computer programs, computer instructions, data, etc., can be accessed by the processor.

[0139] The aforementioned computer programs, computer instructions, etc., can be stored in partitions within one or more memory locations. Furthermore, the aforementioned computer programs, computer instructions, data, etc., can be accessed by a processor.

[0140] A processor is used to execute a computer program stored in memory to implement the various steps of the methods involved in the above embodiments. For details, please refer to the relevant descriptions in the preceding method embodiments.

[0141] The processor and memory can be separate structures or integrated structures. When the processor and memory are separate structures, they can be coupled together via a bus.

[0142] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0143] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0144] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0145] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0146] The foregoing has described some specific embodiments of this application. It should be understood that this application is not limited to the specific embodiments described above, and those skilled in the art can make various modifications or variations within the scope of the claims, which do not affect the substantive content of this application. The above-described preferred features can be used in any combination without conflict.

Claims

1. A method for estimating endoscopic pose based on style transfer, characterized in that, include: The first part of the dataset is obtained by acquiring three-dimensional medical images containing bronchi and three-dimensional segmentation and reconstruction data of bronchi. Based on the first part of the dataset, perform physical information-based coordinate transformation and virtual camera settings to determine the second part of the dataset; Acquire real endoscopic images of the bronchi as the third part of the dataset; A preset style transfer network model M1 is trained using the second part of the data and the third part of the dataset, and the style transfer network model M1 trained by the model is used to transfer the style from the third part of the dataset to the second part of the data. A similarity measure is performed on the second part of the dataset and the third dataset after style transfer. The image most similar to the third part of the dataset is determined in the second part of the dataset, and the depth and pose information of the most similar image is determined as the depth and pose estimates corresponding to the real endoscope.

2. The method according to claim 1, characterized in that, The acquisition of three-dimensional medical images containing bronchi and three-dimensional segmentation and reconstruction data of bronchi, as the first part of the dataset, includes: Acquire three-dimensional medical images containing bronchi; The three-dimensional medical image containing the bronchi is segmented to determine the three-dimensional segmentation and reconstruction data of the bronchi. The three-dimensional medical image containing the bronchi and the three-dimensional segmentation and reconstruction data of the bronchi are used as the first part of the dataset.

3. The method according to claim 1, characterized in that, The second part of the dataset includes virtual images of the bronchus, virtual depth corresponding to the virtual images of the bronchus, and virtual pose of the virtual camera. The step of performing physical information-based coordinate transformation and virtual camera settings based on the first part of the dataset to determine the second part of the dataset includes: The three-dimensional medical image containing the bronchi is segmented using a preset segmentation model to determine the bronchus segmentation model. The bronchial segmentation model is rendered using a rendering isometry algorithm to generate three-dimensional mesh data; The three-dimensional mesh data is transformed based on physical information to convert it to the world coordinate system, thereby determining the three-dimensional mesh data in the world coordinate system. Centerlines are extracted from the three-dimensional mesh data of the world coordinate system to determine the three-dimensional world coordinates of each centerline; A virtual camera is set at each sampling point of each centerline; The virtual image of the bronchus, the virtual depth corresponding to the virtual image of the bronchus, and the virtual pose of the virtual camera are acquired using a virtual camera at each sampling point of each centerline, and are used as the second part of the dataset.

4. The method according to claim 1, characterized in that, The preset style transfer network model M1 includes a local denoising network model N1 and a global style transfer model N2. The local denoising model N1 includes a first encoder and a first decoder, and the global style transfer model N2 includes a second encoder and a second decoder. The step of training a preset style transfer network model M1 using the second part of the data and the third part of the dataset, and determining the trained style transfer network model M1, includes: The first encoder is used to extract features from the real endoscopic images of the bronchus in the third part of the dataset to determine the first implicit image features; The first decoder is used to decode the first implicit image features to determine the real endoscopic image of the bronchus without artifacts, and to determine the local denoising network model N1 trained by the model. The second encoder is used to extract features from the artifact-free real endoscopic image of the bronchus to determine the second implicit image features; The second decoder is used to decode the second implicit image features to determine the real endoscopic image of the bronchus corresponding to the style of the second part of the dataset, and to determine the global style transfer model N2 trained by the model. Based on the local denoising network model N1 trained by the model and the global style transfer model N2 trained by the model, the style transfer network model M1 trained by the model is determined.

5. The method according to claim 4, characterized in that, The step of training a preset style transfer network model M1 using the second part of the data and the third part of the dataset, and determining the trained style transfer network model M1, further includes: Determine pseudo-labels for training the local denoising network model N1; The determination of pseudo-labels for training the local denoising network model N1 includes: The second part of the dataset is sampled to determine the virtual image of the bronchus, the depth corresponding to the virtual image of the bronchus, and the virtual pose of the virtual camera. Gaussian sputtering model is modeled based on the sampled virtual image of the bronchus, the depth corresponding to the virtual image of the bronchus, and the virtual pose of the virtual camera. The Gaussian sputtering model is determined, and the parameters of the Gaussian sputtering model include position, scale factor, rotation quaternion, transparency, and spherical harmonic coefficient. The Gaussian sputtering model is subjected to feature transfer processing to optimize the spherical harmonic coefficients of the Gaussian sputtering model, and the pseudo-labels used for training the local denoising network model N1 are determined.

6. The method according to claim 4, characterized in that, The step of training a preset style transfer network model M1 using the second part of the data and the third part of the dataset, and determining the trained style transfer network model M1, further includes: The style transfer network model M1, after training, is optimized using an overall loss function, which is: ; in, Represents the overall loss function, This represents the objective function for adversarial training optimization of the transferred image. This represents the adversarial training constraints on the inverse generator network. This represents a consistency constraint on the forward-generating network. This represents the robust feature extraction loss function guided by virtual images.

7. The method according to claim 1, characterized in that, The step of measuring the similarity between the second part of the dataset and the style-transferred third dataset, determining the image in the second part of the dataset that is most similar to the third part of the dataset, and determining the depth and pose information of the most similar image as the depth and pose estimates corresponding to the real endoscope, includes: The pre-trained feature extraction network M2 is used to extract features from the second part of the dataset and the third part of the dataset after style transfer, to determine the first implicit feature and the second implicit feature. The first implicit feature and the second implicit feature are similar to each other by cosine similarity to determine the virtual image of the bronchus, the depth of the virtual image of the bronchus, and the virtual pose of the virtual camera in the second part of the dataset. The depth corresponding to the virtual image of the bronchus in the second part of the dataset corresponding to the second implicit feature is determined as the depth corresponding to the real endoscopic image of the bronchus in the third part of the dataset, and the virtual pose of the virtual camera corresponding to the second implicit feature in the second part of the dataset is determined as the pose estimate of the real endoscope.

8. An endoscope pose estimation system based on style transfer, characterized in that, include: The first part of the dataset determination module is used to acquire three-dimensional medical images containing bronchi and three-dimensional segmentation and reconstruction data of bronchi, which are used as the first part of the dataset. The second part of the dataset determination module is used to determine the second part of the dataset by performing coordinate transformation based on physical information and setting up virtual cameras according to the first part of the dataset. The third part of the dataset determination module is used to acquire the actual endoscopic images of the bronchus as the third part of the dataset. The style transfer module is used to train a preset style transfer network model M1 using the second part of the data and the third part of the dataset, and to use the style transfer network model M1 trained by the model to perform style transfer from the third part of the dataset to the second part of the data. The endoscope pose estimation module is used to measure the similarity between the second part of the dataset and the style-transferred third dataset, identify the image in the second part of the dataset that is most similar to the third part of the dataset, and determine the depth and pose information of the most similar image as the depth and pose estimates corresponding to the real endoscope.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-7.

10. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-7.