Laparoscopic surgery mixed reality navigation method based on deep learning and dynamic point tracking
By employing a mixed reality navigation method that combines deep learning and dynamic point tracking, the problem of large registration errors in fusing preoperative 3D models and intraoperative laparoscopic video views in navigation systems has been solved, thereby improving navigation accuracy and enhancing the spatial perception and safety of surgery.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTHEAST UNIV
- Filing Date
- 2025-06-05
- Publication Date
- 2026-08-04
AI Technical Summary
In existing navigation systems, the fusion of preoperative 3D models and intraoperative laparoscopic videos results in large registration errors and low navigation accuracy, making it difficult to locate anatomical structures such as the renal artery and tumors, increasing the difficulty of surgical procedures and easily damaging the integrity of the tumor.
A mixed reality navigation method based on deep learning and dynamic point tracking is adopted. The three-dimensional model of the integrated kidney structure of the object to be detected is aligned with the initial frame of the laparoscopic video using mixed reality technology. The camera pose estimation module is initialized, the feature points in the laparoscopic video are tracked in real time, and the camera pose parameters are calculated according to the feature point selection strategy to generate a photo with adaptive transparency to overlay the current frame.
It significantly enhances the spatial perception of kidney anatomy, reduces the registration error between the three-dimensional integrated kidney structure model and laparoscopic video, improves navigation accuracy, and enhances the precision and safety of surgery.
Smart Images

Figure CN120894390B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of medical image processing and mixed reality technology, and in particular relates to a mixed reality navigation method for laparoscopic surgery based on deep learning and dynamic point tracking. Background Technology
[0002] Laparoscopic partial nephrectomy is the mainstream minimally invasive surgical procedure for treating renal cell carcinoma. Its core challenge lies in the real-time intraoperative localization of the three-dimensional spatial relationships between the renal artery, vein, tumor, and surrounding tissues. Traditional surgery relies on the surgeon's experience, memory of preoperative CTA images, and intraoperative two-dimensional laparoscopic video for locating the renal artery and tumor. Surgeons must visually identify the complex three-dimensional anatomical structures of the kidney in the two-dimensional laparoscopic video, but two-dimensional images are insufficient to represent these complex anatomical relationships, especially regarding the anatomical positions of small vascular branches and tumor boundaries, as well as the relationships between the kidney and tumor, tumor and arteries, and tumor and collecting systems. Therefore, intraoperative localization of the tumor and the arteries supplying it is slow, making the surgery difficult and prone to damaging the tumor's integrity. Clinically, there is an urgent need for a mixed reality navigation system that integrates preoperative CTA image views with intraoperative laparoscopic video views to improve surgical accuracy and safety. Furthermore, existing preoperative CTA image visualization technology has already achieved the reconstruction of a three-dimensional integrated renal structural model based on CTA.
[0003] However, current navigation systems suffer from large registration errors and low navigation accuracy due to the fusion of preoperative 3D models and intraoperative laparoscopic videos. Summary of the Invention
[0004] This application provides a mixed reality navigation method for laparoscopic surgery based on deep learning and dynamic point tracking, which can solve the problems of large registration errors and low navigation accuracy in the current navigation system's view fusion of preoperative 3D model and intraoperative laparoscopic video.
[0005] In a first aspect, embodiments of this application provide a mixed reality navigation method for laparoscopic surgery based on deep learning and dynamic point tracking, comprising the following steps: Step P1, aligning the 3D model of the integrated kidney structure of the detection object with the 2D kidney in the initial frame of the laparoscopic video using mixed reality (MR), and initializing the camera pose estimation module (CPEM); the initialization of the CPEM includes initializing the camera pose parameters (T... o ,θ o ), Initialize the 2D feature point set P2 D Initialize the 3D feature point set P3 D Step P2: The point tracker PT, which uses multi-feature point joint tracking, tracks P2 in the laparoscopic video in real time. D And dynamically update P2 at preset time intervals. D and P2D Corresponding P3 D Step P3: Based on the feature point selection strategy FSS, in each frame of laparoscopic video, at P2... D With P3 D Four pairs of feature points are selected, and the coordinates of these four pairs of feature points are used to calculate the camera pose parameters (T) corresponding to the current frame of the laparoscopic video in real time. t ,θ t According to (T) t ,θ t ) Render and photograph a 3D model of the integrated kidney structure, and generate an image with adaptive transparency to overlay the current frame.
[0006] In one possible implementation of the first aspect, step P1 involves aligning a 3D model of the integrated kidney structure of the object being detected with a 2D kidney in the initial frame of the laparoscopic video using MR, specifically including the following steps:
[0007] P101. Load the 3D model integrating the kidney structure into the MR system, and construct a time-series-based deep neural network F. The input of neural network F is the time-series image sequence of adjacent τ frames in the laparoscopic video [I] t-τ ,...,I t ]∈R^{τ×H×W×3}, the output of neural network F is the dynamic anchor box B of the tip of the surgical forceps in each frame of the time sequence image. t = {x, y, w, h}, where x and y are the coordinates of the center point of the dynamic anchor frame, and w and h are the width and height of the dynamic anchor frame;
[0008] P102, Based on B t The trajectory Δp of the extraction forceps tip t = (Δx, Δy), which is mapped to the translation T of the 3D model in the camera coordinate system. t =λ·Δp t λ is the preset translation scaling factor; α is the rotation angle of the forceps body about its major axis. t After Kalman filtering, it is mapped to the Z-axis rotation θ of the 3D model around the camera coordinate system. t =k·α t k is the preset rotation ratio coefficient; d is the opening and closing distance of the surgical forceps. t Scaling factor s of the 3D model mapped by bilinear interpolation t =1+μ·d t μ is the preset scaling factor; β is the tilt angle of the surgical forceps. t Mapped to transparency parameter ρ t =σ·β t σ is the preset transparency coefficient;
[0009] P103. Model the time-series image sequence using a neural network F, and process the current frame I. t Simultaneously fuse contextual features from the previous τ frames to generate a motion-continuous detection result B. t According to B t Generate three-dimensional pose control signal {T t ,θ t ,s t ,ρ t}, according to {T t ,θ t , st, ρ t The 3D model integrating the kidney structure is driven to update its pose accordingly;
[0010] P104. Lock the video frame of the laparoscopic video at the current moment as the initial frame, record the pose parameters of the 3D model integrating the kidney structure at the current moment, and freeze the detection result of the neural network F at the current moment as the origin of the reference coordinate system.
[0011] Optionally, in another possible implementation of the first aspect, the initialization of CPEM in step P1 above specifically includes the following steps:
[0012] Initialize camera pose parameters (T) o ,θ o ): The camera position parameter T when aligning the 3D model integrating the kidney structure in the camera scene with the 2D kidney in the laparoscopic video frame. o With rotation angle parameter θ o ;
[0013] Initialize the 2D feature point set P2 D Based on the initial frame of the laparoscopic video and a semantic segmentation mask of an integrated 3D model of the kidney structure aligned with the initial frame, multiple feature points are selected in the kidney and tumor regions of the mask image. Neighborhood detection is performed on each feature point using connected component analysis and spatial uniformity detection. Feature points located at the edges of anatomical structures or in densely distributed areas within the kidney and tumor regions are removed, and feature points satisfying the neighborhood coverage condition are retained as P2. D and initialize P2 D The neighborhood coverage condition is that among the rasterized coordinate points with a distance of 20 pixels in the 8 discrete neighborhood directions around a feature point, at least 6 rasterized coordinate points have been occupied by other feature points.
[0014] Initialize the 3D feature point set P3 D : P2 DThe 2D coordinates are converted into ray direction vectors in the camera coordinate system. The coordinates of the 3D feature points corresponding to the 2D feature points are calculated by the intersection of the ray direction vectors and the 3D model of the integrated kidney structure. Invalid points that do not hit the anatomical structure areas of the kidney and tumor regions are filtered out from the 3D feature point set. The 3D feature point set after removing invalid points is used as P3. D and initialize P3 D .
[0015] Optionally, in another possible implementation of the first aspect, in step P2 above, the point tracker PT, which performs multi-feature point joint tracking, tracks the laparoscopic video in real time. D Specifically, it includes the following steps:
[0016] The laparoscopic video sequence of 16 consecutive frames and P2 D Input to the neural network;
[0017] Parallel dilated convolutional layers of neural networks are used to extract multi-level spatiotemporal features from 16 consecutive frames of time-series images in laparoscopic video.
[0018] Establish a multi-head attention mechanism in the time dimension and calculate P2. D The correlation weight of each feature point across consecutive frames;
[0019] Multi-level spatiotemporal features and the correlation weights of each feature point between consecutive frames are input into the spatiotemporal convolutional layer of the neural network to generate the trajectory prediction matrix and visibility confidence of each feature point.
[0020] A sliding window mechanism is used to process continuous video streams of laparoscopic video. Full temporal modeling is performed every 16 frames, retaining the spatiotemporal features of the first 8 frames as temporal context. The hidden state of the window is passed through a gated loop unit, and high-confidence trajectory points are dynamically selected to update P2. D gather;
[0021] Output the trajectory matrix of all feature points.
[0022] Optionally, in another possible implementation of the first aspect, step P2 is dynamically updated in the above step P2. D and P2 D Corresponding P3 D The steps are the same as the initialization P2 in claim 3. D and initializing P3 D The steps are the same.
[0023] Optionally, in another possible implementation of the first aspect, the above initialization of P2... D and dynamically updated P2 DThe method employs a surgical vision-guided multimodal feature sampling approach, specifically including the following steps:
[0024] A three-dimensional sampling constraint space based on anatomical saliency maps was constructed. A dynamic no-sampling region mask was generated by fusing multimodal data. The multimodal data included the high-reflectivity area and fat yellowing area detected by the reflectivity analysis model based on the HSV-LAB mixed color space, the dynamic contour of surgical instruments identified by the motion saliency detection algorithm, and the blood fog concentration distribution map calculated by the haze scattering model.
[0025] An adaptive sampling topology mesh was constructed on the surface of the initially aligned 3D kidney model. Based on curvature manifold analysis, the surface of the 3D kidney model was divided into multiple anatomical functional zones. High-density sampling anchor points were set in the renal hilum depression and tumor elevation areas, which are designated as anatomical functional zones. A gradient direction consistency detection algorithm was used to screen regions with high gradient direction consistency (gradient direction variance ≤ 0.1) in the anatomical functional zones, excluding tissue adhesion transition zones. Through bidirectional visibility testing, sampling points that meet multi-view geometric constraints were retained as candidate feature points.
[0026] Candidate feature points are uniformly resampled based on the Voronoi subdivision algorithm; a deep learning model for optical flow trajectory prediction is constructed, and potential motion blur sensitive points are predicted and then eliminated based on the prediction of the deep learning model; the spatial distribution threshold λ and texture complexity threshold τ are dynamically calculated using a pre-trained energy function; the output is a set of feature points P2 that satisfies the condition that the spatial dispersion is greater than λ and the texture entropy is less than τ. D .
[0027] Optionally, in another possible implementation of the first aspect, the feature point selection strategy FSS in step P3 above is applied to each frame of the laparoscopic video in step P2. D With P3 D Four pairs of feature points are selected, and the coordinates of these four pairs of feature points are used to calculate the camera pose parameters (T) corresponding to the current frame of the laparoscopic video in real time. t ,θ t Specifically, it includes the following steps:
[0028] Step P301: Calculate the reprojection error E of the selected feature point combination from the previous frame of the laparoscopic video in the current frame. rep If E rep If the current frame's camera pose is below the first preset threshold, the rotation change of the current frame's camera pose relative to the previous frame's camera pose is less than the first dynamic threshold, and the translation change of the current frame's camera pose relative to the previous frame's camera pose is less than the second dynamic threshold, then the camera pose parameters corresponding to the feature point combination of the previous frame are used as the camera pose parameters corresponding to the current frame; otherwise, continue with the following steps.
[0029] Step P302: Based on the 2D feature point set of the current frame, divide each feature point into four quadrant regions according to the screen space, and group and label the feature points in each quadrant region.
[0030] A total of 32 candidate feature point combinations were generated through two stages of sampling.
[0031] In the first stage, a feature point is randomly selected from each quadrant region to form a four-point combination that is evenly distributed across the four quadrants. This process is repeated 16 times to generate 16 sets of candidate feature point combinations.
[0032] In the second stage, two different quadrant regions are randomly selected, and two feature points are randomly selected from each selected quadrant region to form a four-point combination distributed across regions. This process is repeated 16 times to generate 16 sets of candidate feature point combinations.
[0033] Step P303: For each of the 32 candidate feature point combinations, calculate the comprehensive score for each candidate feature point combination.
[0034] Score = α·E rep +β·(Δθ+ΔT)
[0035] Among them, E rep Δθ is the global reprojection error corresponding to the camera pose in each candidate feature point combination, Δθ is the difference in rotation angle between the camera pose in each candidate feature point combination and the camera pose in the previous frame, ΔT is the difference in translation distance between the camera pose in each candidate feature point combination and the camera pose in the previous frame, and α and β are dynamic weighting coefficients.
[0036] Step P304: Output the candidate feature point combination with the highest comprehensive score and its corresponding pose parameters (T). opt ,θ opt ), and (T) opt ,θ opt ) as the camera pose parameters (T) corresponding to the current frame t ,θ t ).
[0037] Beneficial effects: In the technical solution of this application, the three-dimensional model of the integrated kidney structure of the detection object is first aligned with the two-dimensional 2D kidney in the initial frame of the laparoscopic video using MR, and the camera pose estimation module CPEM is initialized. The initialization of CPEM includes initializing the camera pose parameters (T... o ,θ o ), Initialize the 2D feature point set P2 D Initialize the 3D feature point set P3 D Then, based on the point tracker PT, which uses multi-feature point joint tracking, P2 in the laparoscopic video is tracked in real time. DAnd dynamically update P2 at preset time intervals. D and P2 D Corresponding P3 D Finally, based on the feature point selection strategy FSS, P2 of each frame of laparoscopic video is selected. D With P3 D Four pairs of feature points are selected, and the coordinates of these four pairs of feature points are used to calculate the camera pose parameters (T) corresponding to the current frame of the laparoscopic video in real time. t ,θ t According to (T) t ,θ t The system renders and photographs a 3D model of the integrated kidney structure, generating an image with adaptive transparency to overlay the current frame. This significantly enhances the spatial perception of kidney anatomy, reduces registration errors between the 3D integrated kidney structure model and the kidney in the laparoscopic video, and improves navigation accuracy. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a flowchart illustrating a mixed reality navigation method for laparoscopic surgery based on deep learning and dynamic point tracking, provided in an embodiment of this application.
[0040] Figure 2 This is a schematic diagram of a CTA imaging image segmentation result visualized as a 3D integrated kidney structure model according to an embodiment of this application;
[0041] Figure 3 This is a schematic diagram of the MR registration initialization result provided in an embodiment of this application;
[0042] Figure 4 This is a schematic diagram of the sampling results during the initial generation or update of 2D feature points according to an embodiment of this application;
[0043] Figure 5 This is a schematic diagram illustrating the process of making a 3D integrated kidney structure model semi-transparent and fusing it frame-by-frame with laparoscopic images in an MR system, according to an embodiment of this application. Detailed Implementation
[0044] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0045] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0046] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0047] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0048] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0049] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0050] The following is a detailed description of the mixed reality navigation method for laparoscopic surgery based on deep learning and dynamic point tracking provided in this application, with reference to the accompanying drawings.
[0051] Figure 1 The illustration shows a flowchart of a mixed reality navigation method for laparoscopic surgery based on deep learning and dynamic point tracking, provided in an embodiment of this application.
[0052] like Figure 1 As shown, this mixed reality navigation method for laparoscopic surgery based on deep learning and dynamic point tracking includes the following steps:
[0053] S1. Align the 3D model of the integrated kidney structure of the detection object with the 2D kidney in the initial frame of the laparoscopic video using Mixed Reality (MR), and initialize the camera pose estimation module (CPEM). CPEM initialization includes initializing the camera pose parameters (T... o ,θ o ), Initialize the 2D feature point set P2 D Initialize the 3D feature point set P3 D ;
[0054] Furthermore, in this embodiment, step S1 above, which aligns the 3D model of the integrated kidney structure of the detection object with the 2D kidney in the initial frame of the laparoscopic video using MR, specifically includes the following steps:
[0055] S101. Load the 3D model integrating the kidney structure into the MR system, and construct a time-series-based deep neural network F. The input of neural network F is the time-series image sequence of adjacent τ frames in the laparoscopic video [I]. t-τ ,...,I t ]∈R^{τ×H×W×3}, the output of neural network F is the dynamic anchor box B of the tip of the surgical forceps in each frame of the time sequence image. t = {x, y, w, h}, where x and y are the coordinates of the center point of the dynamic anchor frame, and w and h are the width and height of the dynamic anchor frame;
[0056] S102, Based on B t The trajectory Δp of the extraction forceps tip t = (Δx, Δy), which is mapped to the translation T of the 3D model in the camera coordinate system. t =λ·Δp t λ is the preset translation scaling factor; α is the rotation angle of the forceps body about its major axis. t After Kalman filtering, it is mapped to the Z-axis rotation θ of the 3D model around the camera coordinate system. t =κ·α t κ is the preset rotational scaling factor; d is the opening and closing distance of the surgical forceps. t Scaling factor s of the 3D model mapped by bilinear interpolation t =1+μ·d tμ is the preset scaling factor; β is the tilt angle of the surgical forceps. t Mapped to transparency parameter ρ t =σ·β t σ is the preset transparency coefficient;
[0057] S103. Model the time-series image sequence using neural network F, and process the current frame I. t Simultaneously fuse contextual features from the previous τ frames to generate a motion-continuous detection result B. t According to B t Generate three-dimensional pose control signal {T t ,θ t ,s t ,ρ t}, according to {T t ,θ t ,s t ,ρ t The 3D model integrating the kidney structure is driven to update its pose accordingly;
[0058] S104. Lock the video frame of the laparoscopic video at the current moment as the initial frame, record the pose parameters of the 3D model integrating the kidney structure at the current moment, and freeze the detection result of the neural network F at the current moment as the origin of the reference coordinate system.
[0059] Specifically, the initialization of CPEM in step S1 above includes the following steps:
[0060] Initialize camera pose parameters (T0, θ) o ): The camera position parameter T when aligning the 3D model integrating the kidney structure in the camera scene with the 2D kidney in the laparoscopic video frame. o With rotation angle parameter θ o ;
[0061] Initialize the 2D feature point set P2 D Based on the initial frame of the laparoscopic video and a semantic segmentation mask of an integrated 3D model of the kidney structure aligned with the initial frame, multiple feature points are selected in the kidney and tumor regions of the mask image. Neighborhood detection is performed on each feature point using connected component analysis and spatial uniformity detection. Feature points located at the edges of anatomical structures or in densely distributed areas within the kidney and tumor regions are removed, and feature points satisfying the neighborhood coverage condition are retained as P2. D and initialize P2 D The neighborhood coverage condition is that among the rasterized coordinate points with a distance of 20 pixels in the 8 discrete neighborhood directions around a feature point, at least 6 rasterized coordinate points have been occupied by other feature points.
[0062] Initialize the 3D feature point set P3 D : P2D The 2D coordinates are converted into ray direction vectors in the camera coordinate system. The coordinates of the 3D feature points corresponding to the 2D feature points are calculated by the intersection of the ray direction vectors and the 3D model of the integrated kidney structure. Invalid points that do not hit the anatomical structure areas of the kidney and tumor regions are filtered out from the 3D feature point set. The 3D feature point set after removing invalid points is used as P3. D and initialize P3 D .
[0063] S2. The point tracker PT, based on multi-feature point joint tracking, tracks P2 in the laparoscopic video in real time. D And dynamically update P2 at preset time intervals. D and P2 D Corresponding P3 D ;
[0064] It should be noted that, in this embodiment of the application, P2 is dynamically updated. D and P2 D Corresponding P3 D The steps are the same as the initialization P2 in the above embodiment. D and initializing P2 D The steps are the same.
[0065] Specifically, initialize P2 D and dynamically updated P2 D The method employs a surgical vision-guided multimodal feature sampling approach, including the following steps:
[0066] A three-dimensional sampling constraint space based on anatomical saliency maps was constructed. A dynamic no-sampling region mask was generated by fusing multimodal data. The multimodal data included the high-reflectivity area and fat yellowing area detected by the reflectivity analysis model based on the HSV-LAB mixed color space, the dynamic contour of surgical instruments identified by the motion saliency detection algorithm, and the blood fog concentration distribution map calculated by the haze scattering model.
[0067] An adaptive sampling topology mesh was constructed on the surface of the initially aligned 3D kidney model. Based on curvature manifold analysis, the surface of the 3D kidney model was divided into multiple anatomical functional zones. High-density sampling anchor points were set in the renal hilum depression and tumor elevation areas, which are designated as anatomical functional zones. A gradient direction consistency detection algorithm was used to screen regions with high gradient direction consistency (gradient direction variance ≤ 0.1) in the anatomical functional zones, excluding tissue adhesion transition zones. Through bidirectional visibility testing, sampling points that meet multi-view geometric constraints were retained as candidate feature points.
[0068] Candidate feature points are uniformly resampled based on the Voronoi subdivision algorithm; a deep learning model for optical flow trajectory prediction is constructed, and potential motion blur sensitive points are predicted and then eliminated based on the prediction of the deep learning model; the spatial distribution threshold λ and texture complexity threshold τ are dynamically calculated using a pre-trained energy function; the output is a set of feature points P2 that satisfies the condition that the spatial dispersion is greater than λ and the texture entropy is less than τ. D .
[0069] Furthermore, in this embodiment of the application, in step S2 above, the point tracker PT, which performs multi-feature point joint tracking, tracks P2 in the laparoscopic video in real time. D Specifically, it includes the following steps:
[0070] The laparoscopic video sequence of 16 consecutive frames and P2 D Input to the neural network;
[0071] Parallel dilated convolutional layers of neural networks are used to extract multi-level spatiotemporal features from 16 consecutive frames of time-series images in laparoscopic video.
[0072] Establish a multi-head attention mechanism in the time dimension and calculate P2. D The correlation weight of each feature point across consecutive frames;
[0073] Multi-level spatiotemporal features and the correlation weights of each feature point between consecutive frames are input into the spatiotemporal convolutional layer of the neural network to generate the trajectory prediction matrix and visibility confidence of each feature point.
[0074] A sliding window mechanism is used to process continuous video streams of laparoscopic video. Full temporal modeling is performed every 16 frames, retaining the spatiotemporal features of the first 8 frames as temporal context. The hidden state of the window is passed through a gated loop unit, and high-confidence trajectory points are dynamically selected to update P2. D gather;
[0075] Output the trajectory matrix of all feature points.
[0076] S3. Based on the feature point selection strategy, FSS is applied to P2 of each frame of laparoscopic video. D With P3 D Four pairs of feature points are selected, and the coordinates of these four pairs of feature points are used to calculate the camera pose parameters (T) corresponding to the current frame of the laparoscopic video in real time. t ,θ t According to (T) t ,θ t ) Render and photograph a 3D model of the integrated kidney structure, and generate an image with adaptive transparency to overlay the current frame.
[0077] Furthermore, in this embodiment of the application, in step S3 above, the feature point selection strategy FSS is applied to P2 of each frame of laparoscopic video. D With P3 D Four pairs of feature points are selected, and the coordinates of these four pairs of feature points are used to calculate the camera pose parameters (T) corresponding to the current frame of the laparoscopic video in real time. t ,θ t Specifically, it includes the following steps:
[0078] S301. Calculate the reprojection error E of the selected feature point combination from the previous frame of the laparoscopic video in the current frame. rep If E rep If the current frame's camera pose is below the first preset threshold, the rotation change of the current frame's camera pose relative to the previous frame's camera pose is less than the first dynamic threshold, and the translation change of the current frame's camera pose relative to the previous frame's camera pose is less than the second dynamic threshold, then the camera pose parameters corresponding to the feature point combination of the previous frame are used as the camera pose parameters corresponding to the current frame; otherwise, continue with the following steps.
[0079] S302. Based on the 2D feature point set of the current frame, divide each feature point into four quadrant regions according to the screen space, and group and label the feature points in each quadrant region.
[0080] A total of 32 candidate feature point combinations were generated through two stages of sampling.
[0081] In the first stage, a feature point is randomly selected from each quadrant region to form a four-point combination that is evenly distributed across the four quadrants. This process is repeated 16 times to generate 16 sets of candidate feature point combinations.
[0082] In the second stage, two different quadrant regions are randomly selected, and two feature points are randomly selected from each selected quadrant region to form a four-point combination distributed across regions. This process is repeated 16 times to generate 16 sets of candidate feature point combinations.
[0083] S303. For each of the 32 candidate feature point combinations, calculate the comprehensive score for each candidate feature point combination:
[0084] Score = α·E rep +β·(Δθ+ΔT)
[0085] Among them, E rep Δθ is the global reprojection error corresponding to the camera pose in each candidate feature point combination, Δθ is the difference in rotation angle between the camera pose in each candidate feature point combination and the camera pose in the previous frame, ΔT is the difference in translation distance between the camera pose in each candidate feature point combination and the camera pose in the previous frame, and α and β are dynamic weighting coefficients.
[0086] S304. Output the candidate feature point combination with the highest comprehensive score and its corresponding pose parameters (T). opt ,θ opt ), and (T) opt ,θ opt ) as the camera pose parameters (T) corresponding to the current frame t ,θ t ).
[0087] As an example, the solution of this application will be described below using an embodiment.
[0088] Step (1): Input the CTA angiography image of the target object. Use a deep learning model based on meta-grayscale adaptive DenseBiasNet to segment the kidney, tumor, artery, and vein regions. After morphological post-processing (such as hole filling and edge smoothing), a high-precision semantic mask is generated. The MarchingCubes algorithm is used to convert the mask into a 3D mesh model, rendered in white (kidney), yellow (tumor), red (artery), and blue (vein) respectively, preserving the geometric topology. See [link to relevant documentation]. Figure 2 .
[0089] The 3D model is lightweighted to reduce the number of faces to meet the requirements of real-time rendering, while anatomical features are preserved through Laplacian smoothing.
[0090] Step (2): Align the 3D kidney model with the 2D kidney in the initial frame of the laparoscopic video using mixed reality (MR) registration. See [link to MR registration]. Figure 3 Then, the camera pose estimation module CPEM is initialized. The process is divided into preoperative calibration, intraoperative initial registration, and initial camera pose (T). o ,θ o Generate, initial 2D feature points (P2) D Generate, initial 3D feature points (P3) D The process involves four sub-steps, as follows:
[0091] Step (P201): Load the 3D model integrating the kidney structure into the MR system, and construct a time-series-based deep neural network F, whose input is a time-series image sequence of adjacent τ frames in a continuous laparoscopic video stream [I]. t-τ ,...,I t ]∈R^{τ×H×W×3}, the output is the dynamic anchor box B of the tip of the surgical forceps in each frame image. t = {x, y, w, h}, where x and y are the coordinates of the center point, and w and h are the width and height;
[0092] Step (P202): Establish a dynamic mapping between the movement of the surgical forceps and the model pose: based on B... t Extract the movement trajectory Δp of the pincer tip t= (Δx, Δy), which is mapped to the translation T of the 3D model in the camera coordinate system. t =λ·Δp t The angle of rotation α of the clamp body about its major axis t After Kalman filtering, it is mapped to the model's rotation θ around the Z-axis. t =κ·α t ; Pliers opening and closing distance d t The model scaling factor s is mapped using bilinear interpolation. t =1+μ·d t ; clamp tilt angle β t Mapped to transparency parameter ρ t =σ·β t ;
[0093] Step (P203): Real-time execution in a continuous video stream: First, model the temporal image sequence using a neural network F, and process the current frame I. t The contextual features of the previous τ frames are fused in real time to generate a detection result B with motion continuity. t Subsequently, according to B t Generate three-dimensional pose control signal {T t ,θ t ,s t ,ρ t Finally, the 3D model is driven to update its pose accordingly;
[0094] Step (P204): When the model edge and the renal vascular pattern in the dynamic video reach subpixel-level alignment, trigger the foot pedal synchronous execution: lock the video frame at the current time t0 as the initial frame F0, and record the model pose parameters at this moment as {T0, θ}. o}, and freeze the detection results of neural network F. As the origin of the reference coordinate system.
[0095] Step (P205): Based on the photo of the 3D kidney model aligned with the initial frame, extract the binary mask of the entire kidney model by excluding the black background, and then extract the binary mask of the kidney body and the tumor area by excluding the mask areas corresponding to the arteries (red rendering) and veins (blue rendering).
[0096] Step (P206): Combine the RGB image data of the initial frame of the laparoscopic video to perform visual detection on the incompletely peeled fat-covered area: establish a color distribution model of adipose tissue through HSV color space conversion, perform threshold segmentation on the yellow-white and rough textured areas in the video frame, and generate a temporary mask for the unpeeled fat area.
[0097] Step (P207): Perform a Boolean subtraction operation on the kidney body and tumor mask obtained in step (P205) and the fat-unpeeled mask obtained in step (P206) to obtain the final sampling source domain, ensuring that the sampling points are only distributed on the exposed kidney surface and tumor area.
[0098] In steps (P208) and (P207), the final sampling source domains obtained in step (P207) are used to select new feature points according to the gridding rules. Through connected component analysis, spatial uniformity detection, and neighborhood detection for each grid center point, grid center points located at the edges of anatomical structures or densely distributed areas are removed, and points that meet the neighborhood coverage condition are retained as P2. D ,See Figure 4 ;
[0099] Step (P209): Place P2 D The 2D coordinates are converted into ray direction vectors in the camera coordinate system. Then, the corresponding three-dimensional coordinates are calculated through the intersection of the ray and the 3D model, and invalid points that do not hit the anatomical structure are filtered out.
[0100] Step (3): Considering the tracking drift problem caused by tissue deformation, occlusion, and lighting changes in the laparoscopic scene, the multi-feature point joint tracking point tracker (PT) is built based on the CoTracker framework. It adopts an online prediction mode to achieve continuous and stable tracking of feature points in the kidney region in the laparoscopic video stream. The system realizes the evolution of the feature point set through a "two-stage dynamic update mechanism": In the real-time tracking stage, the CoTracker model uses a processing window period of 16 frames and performs high-precision tracking of the 2D feature point trajectory in the video sequence through a spatiotemporal attention mechanism to generate feature point motion trajectory and visibility data with temporal continuity; In the dynamic reconstruction stage, the system combines the current camera pose parameters and implements a "gridized feature screening strategy" in the effective area of the kidney. It selects the largest anatomical structure area through connected component analysis, obtains candidate 2D feature points using an adaptive grid sampling method (default 16 pixel interval), and removes isolated points based on neighborhood density verification (requiring at least 6 effective points in the 8-neighborhood of each feature point) to form an optimized P2. D ,See Figure 4 The 2D set is converted into 3D spatial coordinates using a ray mapping engine: projection rays are constructed based on the current camera intrinsic matrix, rotation matrix, and position parameters. A ray-mesh collision detection algorithm is then used to precisely map the selected 2D feature points onto the surface of the kidney 3D model, generating the corresponding P3. D This allows for iterative updates of the feature point space.
[0101] Step (4): For each frame of laparoscopic video, four pairs of feature points are selected from the 2D points tracked in step (3) and the mapped 3D points using a feature point selection strategy (FSS). The coordinates of the four pairs of feature points are used to calculate the camera pose parameters (T) of the current frame in real time. t ,θ t And according to T t ,θ t A 3D anatomical model of the kidney was rendered and photographed to generate a semi-transparent image of the kidney model, which was then dynamically overlaid onto the corresponding frames of the laparoscopic video. (See attached image.) Figure 5 The specific process is as follows:
[0102] Step (P401): Calculate the reprojection error (E) of the selected feature point combination in the previous frame in the current frame. rep If E rep If the current camera pose is below the preset threshold and the rotation and translation changes (Δθ, ΔT) relative to the previous frame are both less than the dynamic threshold, then the motion continuity assumption is triggered, and the feature point combination of the previous frame is directly used, skipping the subsequent selection steps.
[0103] Step (P402): The system divides the image plane into four quadrants (Q1~Q4) and performs two-stage sampling based on the spatial distribution of feature points.
[0104] Phase I (Regional Equalization Sampling): Constructing the quadruplet S I = (q1∈Q1,q2∈Q2,q3∈Q3,q4∈Q4), 16 candidate groups are generated using the Monte Carlo method to ensure spatial coverage of anatomical structure surface features;
[0105] Phase II (Cross-Domain Dense Sampling): Randomly select two different quadrants Q a Q b Construct combination S II =(q a1 ,q a2 ∈Q a ,q b1 ,q b2 ∈Q b Repeat this process 16 times to generate cross-regional feature pairs, thereby enhancing robustness to local deformations.
[0106] Step (P403): For the candidate combinations generated in step (P402), calculate the camera pose parameters (T) corresponding to each combination using the solvePnP-AP3P algorithm. cand ,θ cand Then calculate the global E under that pose. rep The formula includes the difference in rotation angle (Δθ) and translation distance (ΔT) between the current pose and the previous frame pose, and finally, the overall score is calculated according to the formula:
[0107] Score = α·E rep +β·(Δθ+ΔT)
[0108] Where α and β are dynamic weighting coefficients, which are dynamically adjusted according to the surgical stage (initially α = 0.7, β = 0.3; β is increased when organ movement is intense, and α is increased when it is stable, α + β = 1);
[0109] Step (P404): Output the optimal feature point combination with the comprehensive score and its corresponding pose parameters (T). opt ,θ opt );
[0110] Step (P405), according to T opt ,θ opt The 3D kidney anatomical model is rendered and photographed to generate a semi-transparent kidney model image, which is then dynamically overlaid onto the corresponding frame of the laparoscopic video.
[0111] The mixed reality navigation method for laparoscopic surgery based on deep learning and dynamic point tracking provided in this application first aligns the 3D model of the integrated kidney structure of the target object with the 2D kidney in the initial frame of the laparoscopic video using MR, and then initializes the camera pose estimation module CPEM. The initialization of CPEM includes initializing the camera pose parameters (T... o ,θ o ), Initialize the 2D feature point set P2 D Initialize the 3D feature point set P3 D Then, based on the point tracker PT, which uses multi-feature point joint tracking, P2 in the laparoscopic video is tracked in real time. D And dynamically update P2 at preset time intervals. D and P2 D Corresponding P3 D Finally, based on the feature point selection strategy FSS, P2 of each frame of laparoscopic video is selected. D With P3 D Four pairs of feature points are selected, and the coordinates of these four pairs of feature points are used to calculate the camera pose parameters (T) corresponding to the current frame of the laparoscopic video in real time. t ,θ t According to (T) t ,θ t The system renders and photographs a 3D model of the integrated kidney structure, generating an image with adaptive transparency to overlay the current frame. This significantly enhances the spatial perception of kidney anatomy, reduces registration errors between the 3D integrated kidney structure model and the kidney in the laparoscopic video, and improves navigation accuracy.
[0112] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0113] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A mixed reality navigation method for laparoscopic surgery based on deep learning and dynamic point tracking, characterized in that, Includes the following steps: S1. Align the 3D model of the integrated kidney structure of the detection object with the 2D kidney in the initial frame of the laparoscopic video using Mixed Reality (MR), and initialize the camera pose estimation module (CPEM); the initialization of the CPEM includes initializing the camera pose parameters. Initialize the 2D feature point set Initialize 3D feature point set ; S2. A point tracker (PT) based on multi-feature point joint tracking tracks the laparoscopic video in real time. And dynamically update at preset time intervals. and corresponding ; S3. Feature point selection strategy FSS in each frame of laparoscopic video. and Four pairs of feature points are selected, and the coordinates of these four pairs of feature points are used to calculate the camera pose parameters corresponding to the current frame of the laparoscopic video in real time. , This represents the translation of the 3D model in the camera coordinate system. This represents the rotation of the 3D model around the Z-axis of the camera coordinate system; according to Render and photograph a 3D model integrating the kidney structure, and generate an image with adaptive transparency to overlay the current frame; In S2, the point tracker PT, which performs multi-feature point joint tracking, tracks the laparoscopic video in real time. Specifically, it includes the following steps: The time sequence of 16 consecutive frames of laparoscopic video and Input to the neural network; Parallel dilated convolutional layers of neural networks are used to extract multi-level spatiotemporal features from 16 consecutive frames of time-series images in laparoscopic video. Establish a multi-head attention mechanism in the time dimension and compute... The correlation weight of each feature point across consecutive frames; Multi-level spatiotemporal features and the correlation weights of each feature point between consecutive frames are input into the spatiotemporal convolutional layer of the neural network to generate the trajectory prediction matrix and visibility confidence of each feature point. A sliding window mechanism is employed to process continuous video streams of laparoscopic video. Full-time modeling is performed every 16 frames, retaining the spatiotemporal features of the first 8 frames as temporal context. The hidden state of the window is passed through a gated loop unit, and high-confidence trajectory points are dynamically selected to update the timeline. gather; Output the trajectory matrix of all feature points; In S3, the feature point selection strategy FSS is applied to each frame of the laparoscopic video. and Four pairs of feature points are selected, and the coordinates of the four pairs of feature points are used to calculate the camera pose parameters corresponding to the current frame of the laparoscopic video in real time. Specifically, it includes the following steps: S301. Calculate the reprojection error of the selected feature point combination from the previous frame of the laparoscopic video in the current frame. ,like If the current frame's camera pose is below the first preset threshold, the rotation change of the current frame's camera pose relative to the previous frame's camera pose is less than the first dynamic threshold, and the translation change of the current frame's camera pose relative to the previous frame's camera pose is less than the second dynamic threshold, then the camera pose parameters corresponding to the feature point combination of the previous frame are used as the camera pose parameters corresponding to the current frame; otherwise, continue with the following steps. S302. Based on the 2D feature point set of the current frame, divide each feature point into four quadrant regions according to the screen space, and group and label the feature points in each quadrant region. A total of 32 candidate feature point combinations were generated through two stages of sampling. In the first stage, a feature point is randomly selected from each quadrant region to form a four-point combination that is evenly distributed across the four quadrants. This process is repeated 16 times to generate 16 sets of candidate feature point combinations. In the second stage, two different quadrant regions are randomly selected, and two feature points are randomly selected from each selected quadrant region to form a four-point combination distributed across regions. This process is repeated 16 times to generate 16 sets of candidate feature point combinations. S303. For each of the 32 candidate feature point combinations, calculate the comprehensive score for each candidate feature point combination: ; in, Let Δθ be the global reprojection error corresponding to the camera pose in each candidate feature point combination, Δθ be the rotation angle difference between the camera pose in each candidate feature point combination and the camera pose in the previous frame, and ΔT be the translation distance difference between the camera pose in each candidate feature point combination and the camera pose in the previous frame. and These are dynamic weighting coefficients; S304. Output the candidate feature point combination with the highest comprehensive score and its corresponding pose parameters. ), and ( ) as the camera pose parameters corresponding to the current frame ( ).
2. The laparoscopic surgery mixed reality navigation method based on deep learning and dynamic point tracking according to claim 1, characterized in that, In step S1, the 3D model of the integrated kidney structure of the object being detected is aligned with the 2D kidney in the initial frame of the laparoscopic video using MR. This specifically includes the following steps: S101. Load the 3D model integrating the kidney structure into the MR system, and construct a time-series-based deep neural network F. The input of the neural network F is a time-series image sequence of adjacent τ frames in the laparoscopic video. The output of the neural network F is the dynamic anchor frame of the tip of the surgical forceps in each frame of the time-series image. ,in The coordinates of the center point of the dynamic anchor frame. The width and height of the dynamic anchor frame; S102, based on The movement trajectory of the extraction forceps tip This is mapped to the translation of the 3D model in the camera coordinate system. , The preset translation scaling factor; the rotation angle of the forceps body about its major axis. After Kalman filtering, it is mapped to the Z-axis rotation of the 3D model around the camera coordinate system. , Preset rotation ratio coefficient; the opening and closing distance of the surgical forceps. Scaling factor mapped to 3D model via bilinear interpolation , Preset scaling factor; tilt angle of the surgical forceps body. Mapped to transparency parameter , The preset transparency factor; S103. Model the time-series image sequence using neural network F and process the current frame. The contextual features of the previous τ frames are fused in real time to generate a detection result with continuous motion. ,according to Generate 3D pose control signals ,according to Drive the 3D model integrating the kidney structure to update the corresponding pose; S104. Lock the video frame of the laparoscopic video at the current moment as the initial frame, record the pose parameters of the 3D model integrating the kidney structure at the current moment, and freeze the detection result of the neural network F at the current moment as the origin of the reference coordinate system.
3. The laparoscopic surgery mixed reality navigation method based on deep learning and dynamic point tracking according to claim 2, characterized in that, The initialization of CPEM in step S1 specifically includes the following steps: Initialize camera pose parameters : Camera position parameters when aligning the 3D model integrating the kidney structure in the camera scene with the 2D kidney in the laparoscopic video frame during initialization. With rotation angle parameter ; Initialize 2D feature point set Based on the initial frame of the laparoscopic video and a semantic segmentation mask of an integrated 3D model of the kidney structure aligned with the initial frame, multiple feature points are selected in the kidney and tumor regions of the mask image. Neighborhood detection is performed on each feature point using connected component analysis and spatial uniformity detection. Feature points located at the edges of anatomical structures or in densely distributed areas within the kidney and tumor regions are removed, retaining only those that meet the neighborhood coverage condition. and initialize The neighborhood coverage condition is that among the rasterized coordinate points of the feature point in its surrounding 8 discrete neighborhood directions with a spacing of 20 pixels, at least 6 rasterized coordinate points have been occupied by other feature points. Initialize 3D feature point set :Will The 2D coordinates are converted into ray direction vectors in the camera coordinate system. The coordinates of the 3D feature points corresponding to the 2D feature points are calculated using the intersection of the ray direction vectors and the 3D model of the integrated kidney structure. Invalid points that do not hit the anatomical structure areas of the kidney and tumor regions are filtered out from the 3D feature point set. The 3D feature point set after removing invalid points is used as the final set. and initialize .
4. The laparoscopic surgery mixed reality navigation method based on deep learning and dynamic point tracking according to claim 3, characterized in that, Dynamic update in step S2 and corresponding The steps are the same as the initialization in claim 3. and initialization The steps are the same.
5. The laparoscopic surgery mixed reality navigation method based on deep learning and dynamic point tracking according to claim 4, characterized in that, The initialization and dynamic updates The method employs a surgical vision-guided multimodal feature sampling approach, specifically including the following steps: A three-dimensional sampling constraint space based on anatomical saliency maps is constructed, and a dynamic no-sampling region mask is generated by fusing multimodal data. The multimodal data includes the high-reflectivity area and fat yellowing area detected by the reflectivity analysis model based on the HSV-LAB mixed color space, the dynamic contour of surgical instruments identified by the motion saliency detection algorithm, and the blood fog concentration distribution map calculated by the haze scattering model. An adaptive sampling topology mesh was constructed on the surface of an initially aligned 3D kidney model. Based on curvature manifold analysis, the surface of the 3D kidney model was divided into multiple anatomical functional zones. High-density sampling anchor points were set in the renal hilum depression and tumor protrusion areas, which are the anatomical functional zones. A gradient direction consistency detection algorithm was used to screen regions with high gradient direction consistency (gradient direction variance ≤ 0.1) in the anatomical functional zones, excluding tissue adhesion transition zones. Through bidirectional visibility testing, sampling points that meet multi-view geometric constraints were retained as candidate feature points. Candidate feature points are uniformly resampled based on the Voronoi subdivision algorithm; a deep learning model for optical flow trajectory prediction is constructed, and potential motion blur sensitive points are predicted and then eliminated based on the prediction of the deep learning model; the spatial distribution threshold λ and texture complexity threshold τ are dynamically calculated using a pre-trained energy function; the output is a set of feature points that satisfy the condition that the spatial dispersion is greater than λ and the texture entropy is less than τ. .