Image convolution enhanced abdominal organ real-time tracking method based on binocular camera
By using a binocular camera-based graph convolution-enhanced real-time abdominal organ tracking method, combined with a 3D segmentation network and biomechanical model, the problem of real-time correction of abdominal organ deformation during surgery is solved, and efficient and accurate organ deformation modeling and dynamic tracking are achieved.
Patent Information
- Application Number
- CN202511322759.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-16
AI Technical Summary
Existing technologies make it difficult to correct the deformation of abdominal organs in real time during surgery, resulting in increased surgical risks and complexity. In addition, the computational efficiency is low and cannot meet real-time processing requirements. The accuracy of organ surface feature extraction is insufficient, affecting the accuracy and reliability of organ deformation modeling.
A binocular camera-based graph convolution-enhanced real-time abdominal organ tracking method is adopted. Preoperative CT images and intraoperative laparoscopic images are obtained through the binocular camera. Rigid registration and elastic compensation are performed by combining the 3D organ segmentation network and biomechanical model. Feature points are extracted using a graph convolutional neural network to achieve non-rigid matching and dynamic tracking.
It improves the accuracy and efficiency of soft tissue deformation modeling, enhances the accuracy and computational efficiency of boundary condition acquisition, can track the complex deformation of abdominal organs in real time, and supports high-precision deformation prediction and dynamic tracking of abdominal organs.
Smart Images

Figure CN120807588A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the cross technical field of deep learning and intraoperative navigation, in particular to a graph convolution enhanced abdominal organ real-time tracking method based on binocular cameras. BACKGROUND
[0002] In the modern medical field, image-guided surgery is a key technology that can use advanced imaging techniques to improve the accuracy and safety of surgery. Preoperative imaging techniques such as computed tomography or magnetic resonance imaging can provide detailed information about the internal structure of abdominal organs, including the spatial location of tumors and blood vessels. This information can help surgeons understand the complexity of the surgical area and predict possible complications, so it is important for the planning and execution of surgery. However, during surgery, abdominal organs and other soft tissues may undergo significant deformation due to surgical operations, organ displacement, cutting and suturing, etc. However, these preoperative images cannot reflect the changes in organ tissue in real time during surgery, increasing the risk and complexity of surgery. Therefore, they cannot be directly used for preoperative navigation. Therefore, a technology is needed that can correct this deformation in real time and update the preoperative image to ensure the accuracy of surgical navigation. Currently, there are three main methods to accurately align preoperative organ models with intraoperative anatomical structures: model-based methods, data-driven methods, and hybrid methods. Among them, the model-based method mainly relies on physical laws and biological principles to simulate organ deformation, including biomechanical models, physics-based modeling methods, and geometric alignment methods. Yang et al. proposed a finite element model that does not require pre-defined boundary conditions. This model simulates the elastic deformation of organs by solving partial differential equations related to soft tissue deformation. This method does not rely on pre-defined zero displacement and force positions, reducing the difficulty of obtaining boundary conditions during surgery. Experiments show that this method achieves a mean TRE of 3.14±1.34mm. Suwelack et al. compared the non-rigid deformation problem to an electrostatic-elastic scenario, where the preoperative liver model is treated as a charged object interacting with the opposite charge on the intraoperative liver surface. The electrostatic force promotes the alignment of the preoperative liver model with the intraoperative surface, while the elastic force is used for regularization. Experiments show that this method has good alignment results and can accurately simulate the deformation of the liver during surgery to some extent. Data-driven methods mainly utilize deep learning techniques to recover the true deformation of organs from a large amount of data, including Pfeiffer et al. using deep learning methods to predict the entire displacement field, reconstructing the deformed part of the surface directly from the intraoperative laparoscopic video stream, which uses a 3D convolutional network to recover the displacement field from the digitized organ surface during surgery, and then uses it to infer the node displacement of the entire organ. Experiments show that this method achieves a mean TRE of 5.70 mm, demonstrating its accuracy in intraoperative liver deformation prediction; Hybrid methods combine the advantages of model-based and data-driven methods, including neural network-driven biomechanical models, physically guided deep learning, and multi-stage deformation modeling methods, including Tagliabue et al. using BANet network to estimate boundary conditions from intraoperative point cloud data, and guiding the deformation of preoperative biomechanical organ models to accurately represent intraoperative organ deformation. This method combines biomechanical models and deep learning methods to estimate boundary conditions through neural networks to drive biomechanical models; However, the above three methods still have many shortcomings: High complexity of tissue deformation: Soft tissues such as the liver may undergo complex deformation during surgery, which may exceed the scope of simple rigid transformation, covering compression, stretching, shear, and other modes, and the degree of deformation may be quite significant, such as the deformation of the liver during laparoscopic surgery may exceed 10 mm. This significant and complex deformation poses great challenges to mathematical modeling and machine learning, making it difficult to build accurate and universal deformation prediction learning models; It is difficult to obtain boundary conditions: To build an accurate biomechanical model, accurate boundary conditions such as displacement or force are needed to simulate organ deformation. In actual surgery, there is a lack of effective force sensing and displacement tracking means, making it difficult to accurately measure the mechanical interaction during surgical operations in real time, resulting in missing or inaccurate model input data, which will affect the authenticity and reliability of the model simulation; Low computational efficiency, difficult to meet the real-time processing requirements in surgery: Biomechanical models usually need to solve large, sparse linear equations, and traditional numerical calculation methods are inefficient. In addition, in the surgical environment, the registration process requires real-time or near real-time response capability to ensure that the surgeon can obtain accurate navigation information in a timely manner. This requires the registration algorithm to not only be accurate but also efficient to meet the real-time processing requirements. Directly solving the finite element-based biomechanical model takes a long time, making it difficult to achieve real-time or near real-time feedback, and unable to meet the stringent requirements of computational efficiency in surgical scenarios, limiting its application and promotion in clinical practice; Insufficient accuracy of organ surface feature extraction: Existing feature extraction methods cannot accurately capture the subtle features of the organ surface, such as the corner points or edges of the liver, which makes it difficult to accurately track and match the surface during surgery. At the same time, existing feature extraction cannot accurately capture key points in the image, such as the branch points of blood vessels, which makes it difficult for the model to understand the spatial relationship between points, affecting the accuracy and reliability of organ deformation modeling. SUMMARY
[0003] The purpose of the present application is to provide a graph convolution enhanced abdominal organ real-time tracking method based on binocular camera to solve the problems of high complexity of tissue deformation, difficulty in obtaining boundary conditions, low computational efficiency, difficulty in meeting the real-time processing requirements during surgery and insufficient accuracy of organ surface feature extraction as proposed in the background art.
[0004] To this end, the present application provides a graph convolution enhanced abdominal organ real-time tracking method based on binocular camera, comprising the following steps: A graph convolution enhanced abdominal organ real-time tracking method based on binocular camera, comprising the following steps: S1: Obtain the preoperative abdominal CT scan image of the patient and the laparoscopic intraoperative image data collected by the binocular camera, and label the target abdominal organ in the preoperative CT image and the intraoperative laparoscope image, respectively; S2: Based on the preoperative CT image, a three-dimensional surface model of the abdominal organ is constructed by a 3D organ segmentation network, and a biomechanical model of the abdominal organ is established by a co-rotational finite element modeling method; S3: According to the intraoperative laparoscope image, the three-dimensional point cloud of the initial frame is extracted, which is rigidly registered with the preoperative three-dimensional surface model to estimate the rotation and translation matrix; S4: For the initial frame three-dimensional point cloud after rigid registration, the organ feature point set is extracted, and the deformation field of the feature point cloud is estimated combined with the elastic constraint of the biomechanical model to realize elastic compensation and capture the actual shape and displacement of the organ; S5: Based on the feature point set of the initial frame, the feature points in the subsequent frames are matched with the initial frame to realize dynamic tracking of the abdominal organ during the entire surgery; S6: Quantitative and qualitative evaluation is performed on the dynamic tracking result to verify the accuracy and robustness of the abdominal organ real-time tracking method.
[0005] Preferably, the S1 specifically comprises: S11: Collect CT image data, use Mimics software to label the abdominal organ contour, obtain the 3D segmentation mask corresponding to the CT image, and provide the required data set for training the subsequent 3D organ segmentation network; S12: Collect the laparoscope video collected by the intraoperative binocular camera, label the abdominal organ region through the Label-me software, obtain the time sequence segmentation mask corresponding to the video image, and provide the training required dataset for the subsequent intraoperative video organ registration network.
[0006] Preferably, the S2 specifically comprises: S21: Preprocessing the obtained preoperative CT image, the preprocessing comprising image block cropping and normalization processing on the image data with a size of 256*256*256; S22: Constructing a 3D organ segmentation network of the preoperative image based on the nnUNet framework, extracting the mask of the abdominal organ region from the preprocessed preoperative CT image, using the Dice loss function for network training, the optimizer being AdamW, the initial value of the learning rate being set to 0.001, the iteration number being set to 10000-15000 times, and using the data augmentation strategy including random scaling, Gaussian noise, contrast enhancement, Gaussian noise and random mirror flipping on the CT image data; S23: Inputting the three-dimensional mask of the abdominal organ into the Marching Cubes algorithm, first preprocessing the original data and reading into the specified array, secondly extracting a unit body and obtaining the values and coordinate positions of the 8 vertices, then comparing the function values of the 8 vertices of the unit body with the given isosurface value to obtain the state table of the unit body, then finding out the unit body edges intersecting with the isosurface according to the state table index, and using linear interpolation to calculate the intersection position, then using the central difference method to obtain the normal vector of the 8 vertices of the unit body, using linear interpolation to obtain the normal vector of each vertex of the triangular facet, and finally drawing the isosurface image according to the coordinates and normal vectors of the triangular facet vertices, and reconstructing a three-dimensional surface model composed of 5000 triangular facets for expressing the organ structure; S24: Optimizing the organ surface model, reducing the number of triangular facets to 1000 by using the quadratic error measurement algorithm, and improving the uniformity of the surface model through Laplacian smoothing and re-meshing processing; S25: Combining the optimized surface model, using the co-rotational finite element modeling method to construct a biomechanical model.
[0007] Preferably, the S25 specifically comprises: S251: Using the quasi-static integral scheme, assuming that the inertia force component is much smaller than the elastic component, regarding the soft tissue as an elastic body, and based on Newton's second law, focusing on static equilibrium and ignoring the inertia force, resulting in the overall biomechanical model where M is the mass, f represents the sum of internal and external forces corresponding to the deformed state and .
[0008] S252: Selecting a tetrahedron as the basic unit of the organ biomechanical model, for each tetrahedron element e, constructing its deformation equation as wherein is the displacement vector of the vertex, is the force acting on the element, is the stiffness matrix of the element, expressed as wherein V is the total volume, is the volume of an element, is the strain-displacement matrix, is the stress-strain matrix, represents the rotation matrix of the current state of the tetrahedron relative to the initial state; the stress-strain matrix is derived from the Young's modulus and Poisson's ratio wherein , E and ν represent the Young's modulus and Poisson's ratio, respectively.
[0009] Preferably, the S3 specifically comprises: S31: Before the start of the operation, use the binocular camera to take multiple angle calibration board shots, use the OpenCV function "cv: findChessboardCorners" to detect the checkerboard corner points in each image, then for each detected corner point, extract its three-dimensional position in the world coordinate system, to establish the mapping relationship between the two-dimensional image coordinates and the three-dimensional world coordinates; use the cv: calibrateCamera function, pass all the collected three-dimensional object points and corresponding two-dimensional image points to the function to calculate the internal parameters and distortion coefficients of the camera; then, perform distortion correction, use the cv: undistort function to correct the distortion of the original image according to the camera intrinsic parameter matrix and distortion coefficient obtained in the previous step; finally, evaluate the calibration by calculating the reprojection error to evaluate the accuracy of the calibration, and complete the general calibration of the binocular camera; S32: Based on the initial frame of the intraoperative endoscope image, a semi-global block matching algorithm is used to match the binocular left and right images. First, the horizontal Sobel operator is used to perform edge detection on the left and right images captured by the calibrated binocular camera to obtain gradient images. Then, the matching cost is calculated. For each pixel, the matching cost with the corresponding pixel under different disparities is calculated. The absolute difference is used as the cost function. The cost is composed of two parts: the BT cost (Birchfield-Tomasi cost) calculated from the preprocessed gradient image and the BT cost calculated from the original image without preprocessing. Then, the two parts are added together to minimize the energy function. The dynamic programming method is used to calculate the cumulative cost along the eight directions, and the minimum value is taken as the final cost. The disparity map is generated, and the post-processing method is used to repair the abnormal values or holes in the disparity map. After median filtering, the initial point cloud for three-dimensional reconstruction is obtained; S33: The depth information of each pixel point in the disparity map is converted into three-dimensional space coordinates to generate the three-dimensional point cloud of the abdominal organ corresponding to the initial frame; S34: A density adaptive registration algorithm based on Dirichlet process Gaussian mixture model is used to solve the registration problem using a variational Bayesian inference framework. The point cloud data and the model are expanded in a coarse-to-fine manner to improve the accuracy and efficiency of registration. The scanning weight of each point is estimated to obtain the initial rotation matrix and translation vector , thereby realizing adaptive rigid registration between the preoperative three-dimensional surface model and the intraoperative three-dimensional point cloud.
[0010] Preferably, the S4 specifically includes: S41: Based on the intraoperative endoscope image, the SiamMask target tracking algorithm is used to delineate the target region. The double-lens video data collected in the S1 step and the abdominal organ region annotation fine-tuning model weight are used. The preselected initial bounding box is used as input to quickly extract the mask of the abdominal organ region in the intraoperative image.
[0011] S42: A feature point extraction network is constructed to extract features from the abdominal organ region to obtain feature point set coordinates and description vectors. S43: The tetrahedral vertexes of the biomechanical model in S2 are used as control points and the feature point set is used for non-rigid matching to realize modeling and correction of organ elastic deformation.
[0012] Preferably, the feature point extraction network specifically includes: A graph convolutional neural network with self-attention mechanism (SA-GCNPoint) is constructed for efficiently extracting feature points of abdominal organs; the overall architecture and training method of the network refer to the SuperPoint network; a single shared encoder is included for processing and reducing the dimension of the input image, and the backend is decoupled into two decoder "heads" that perform interest point detection and interest point description, respectively; the encoder is improved and optimized, and the original VGG encoder is replaced with a graph convolutional network with self-attention mechanism (i.e. 4 groups of Transformer and 4 groups of graph convolutional network), which extracts features of abdominal organs respectively, while the configuration of the decoder remains unchanged; by fusing the self-attention mechanism with the graph convolution, the encoder can capture the complex relationships between all nodes in the organ graph at once, obtain the global context, and efficiently solve the long-distance dependency problem, thereby accurately modeling the interaction relationship between organs in a large receptive field; the network uses a self-supervised learning method to automatically learn the feature points of abdominal organs without manual annotation of data; Preferably, the non-rigid matching specifically includes: a biomechanics control point set each control point is matched according to the minimum Hamming distance and a feature point set to determine the corresponding points , and displacement constraints are established between each pair of matched points to construct the following cost function: wherein, is the stiffness coefficient of the organ model, which is of the same order of magnitude as the Young's modulus, is the number of control points, are the initial rotation matrix and translation vector obtained in S3, respectively, is the global stiffness matrix obtained based on the element stiffness matrix set; the optimal displacement vector is solved by globally minimizing the cost function, so as to realize accurate estimation and matching of the non-rigid deformation field.
[0013] Preferably, S5 specifically includes: S51: using the target tracking algorithm and the feature point extraction network described in S4 to extract feature points from each subsequent frame in the same way, denoted as set ; S52: according to the coordinates and description vectors of each feature point, calculate the Hamming distance between the current frame feature point set and the feature points of the previous frame and their corresponding description vectors, and obtain the initial point pair through nearest neighbor matching; S53: Lucas Kanade optical flow algorithm based on cv.calcOpticalFlowPyrLK function in opencv is used to estimate the displacement of each matched feature point pair, and when the displacement is greater than the mean value + 3 times the standard deviation, it is considered to be an abnormal optical flow, and is removed to obtain a feature point set; S54: the remaining unmatched feature points are matched with the organ biomechanical model according to the non-rigid matching method of claim 7, and a cost function is constructed and solved as follows , wherein ; the final cross-frame feature point displacement estimation is obtained by optimization solution, and the dynamic tracking of the abdominal organ is realized.
[0014] Preferably, the optimization solution process adopts Newton acceleration method, and the step length and momentum are dynamically adjusted to accelerate the convergence, and the steps include the following steps: selecting an initial point and setting a learning rate , in each iteration , the predicted position is calculated, the gradient of the objective function at is calculated, and the momentum is updated, and finally the updated parameters are obtained; until the convergence condition is met, the solution process of the accelerated displacement estimation is accelerated Preferably, the S6 specifically includes: On two different clinical liver open source data sets 3Dircadb (as shown in Table 1) and DePoLL (as shown in Table 2), the registration error is evaluated, the root mean square error (RMSE) measures the average Euclidean distance between the corresponding point sets after registration, is more sensitive to larger errors, and reflects the overall registration accuracy; the mean absolute error (MAE) calculates the average absolute value of the corresponding point distance, is less affected by outliers, and reflects the stability and robustness of the registration result; the maximum error (Error) takes the maximum distance in all corresponding point pairs, directly reflects the registration deviation in the worst case, and is a key indicator for evaluating the robustness of the algorithm; the effectiveness of the proposed algorithm is evaluated by calculating the root mean square error, the mean absolute error and the maximum error between the point cloud and the 3D model, and the six methods of ICP, DCP, RPM-Net, ROPNet, PREDATOR and UTOPIC are compared.
[0015] The binocular camera-based graph convolution enhanced abdominal organ real-time tracking method has the beneficial effects that: The application efficiently solves the problem of complex deformation of soft tissue by deep learning-based feature extraction and biomechanical model based on the finite element method. In the intraoperative registration stage, it is necessary to track the deformation of soft tissue in real time, accurately screen key feature points through a deep learning algorithm, and create a displacement vector field. These feature points provide boundary conditions for the biomechanical model over time, effectively simulate complex deformations including compression, stretching and shearing, and adaptively track the dynamic changes of soft tissue during surgery, significantly improving the ability to capture and model the complex deformation behavior of soft tissue. The application continuously provides time boundary conditions through sparse feature point displacement vector fields to obtain effective quality. These boundary conditions are used for biomechanical models. The accuracy and efficiency of boundary condition acquisition are improved through the cooperative operation of stereo matching and neural network extracted feature points combined with stereo reconstruction technology. The deep learning network can effectively deal with complex situations such as tissue occlusion, bleeding and light changes that may occur during surgery due to its adaptive feature extraction capability. In addition, through the combination of stereo matching and stereo reconstruction technology, high precision can be ensured, and massive image data can be efficiently processed to provide rich and real-time updated boundary condition information for the biomechanical model. The application optimizes the efficiency and accuracy of the algorithm solution. The Nesterov accelerated gradient algorithm is used for solution optimization. A new method for determining the optimal step size is proposed to further improve the efficiency of the algorithm. This method dynamically adjusts the step size in each iteration to adapt to the current optimization state, thereby speeding up the convergence speed and improving the accuracy of registration. As a momentum optimizer, the algorithm adds the momentum of previous iterations in the current iteration to form an effect of accelerating towards the minimum point, which can effectively solve the problem of low efficiency of traditional numerical solution. The application enhances the organ deformation modeling and prediction capability. Feature extraction is performed using a graph convolutional neural network, which can fully capture the structural information and spatial relationship of the abdominal organ grid, thereby achieving more accurate modeling and prediction of abdominal organ deformation. GCN can process irregular structured data, learn the feature representation of liver surface points and blood vessel branch points, etc. by using graph topology information. The self-attention mechanism is introduced into GCN, which allows the model to dynamically adjust the attention degree of neighbor nodes when processing each node, thereby capturing more detailed abdominal organ deformation features and blood vessel branch point features. Not only does it strengthen the model's ability to extract soft tissue deformation features, but also enhances the model's adaptability to different patient tissue structures, expanding the scope of application. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 A flowchart of a graph convolution enhanced abdominal organ real-time tracking method based on a binocular camera according to the application; Figure 2The framework diagram of the liver real-time tracking algorithm of the present application. DETAILED DESCRIPTION
[0017] The technical solutions of the present application will be described in detail below through specific embodiments. Please refer to Figures 1-2 The present application provides a kind of based on binocular camera's graph convolution enhancement abdominal organ real-time tracking method, comprising the following steps: S1: obtaining the preoperative abdominal CT scan image of patient and the laparoscope intraoperative image data collected by binocular camera, and respectively marking target abdominal organ in preoperative CT image and intraoperative laparoscope image; S2: based on preoperative CT image, the three-dimensional surface model of abdominal organ is constructed by 3D organ segmentation network, and the biomechanical model of the abdominal organ is established by adopting co-rotating finite element modeling method; S3: according to intraoperative laparoscope image, the three-dimensional point cloud of abdominal in initial frame is extracted, which is rigidly registered with preoperative three-dimensional surface model, and rotation and translation matrix are estimated; S4: to the three-dimensional point cloud of initial frame after rigid registration, the feature point set of organ is extracted, the deformation field of feature point cloud is estimated in combination with the elastic constraint of the biomechanical model, and elastic compensation is realized, to capture the actual morphology and displacement of organ; S5: based on the feature point set of initial frame, the feature points in subsequent frames are matched with the initial frame, and the dynamic tracking of abdominal organ in the whole process of operation is realized; S6: quantitative and qualitative evaluation is carried out on the dynamic tracking result, to verify the accuracy and robustness of the abdominal organ real-time tracking method.
[0018] S1 specifically includes: S11: CT image data is collected, abdominal organ contour is marked using Mimics software, 3D segmentation mask corresponding to CT image is obtained, and data set required for subsequent 3D organ segmentation network is provided; S22: laparoscope video collected by binocular camera during operation is collected, abdominal organ region is marked by Label-me software, time sequence segmentation mask corresponding to video image is obtained, and data set required for subsequent video organ segmentation network is provided; The data collection and marking process of this step lays a foundation for subsequent model training through multi-modal data and multi-task marking, preoperative CT scanning data of liver and gallbladder, 3D segmentation network is trained to extract three-dimensional structure of abdominal organ from CT; Laparoscope video collected by binocular camera during operation, video segmentation network is trained to segment intraoperative organ region in real time.
[0019] S2 specifically includes: S21: preprocessing the acquired preoperative CT image, the preprocessing including image data size of 256*256*256 image block cropping and normalization processing; S22: constructing a 3D organ segmentation network of the preoperative image based on the nnUNet framework, extracting a mask of the abdominal organ region from the preprocessed preoperative CT image, the network training using a Dice loss function, an optimizer being AdamW, an initial learning rate being set to 0.001, an iteration number being set to 10000-15000 times, and a data augmentation strategy including random scaling, Gaussian noise, contrast enhancement, Gaussian noise and random mirror flipping being used for the CT image data; this step trains a key model, a 3D segmentation network, for organ region identification, which uses a general medical image segmentation framework published in a natural medical journal to realize full-automatic three-dimensional organ precise modeling; S23: inputting the three-dimensional mask of the abdominal organ into the Marching Cubes algorithm, first preprocessing the original data and reading into a specified array, second extracting a unit body and obtaining the values and coordinate positions of the 8 vertices, then comparing the function values of the 8 vertices of the unit body with the given isosurface value to obtain a state table of the unit body, then finding out the unit body edges intersecting with the isosurface according to the state table index, and using linear interpolation to calculate the intersection position, then using the central difference method to obtain the normal vector of the 8 vertices of the unit body, using linear interpolation to obtain the normal vector of each vertex of the triangular facet, and finally drawing the isosurface image according to the coordinates and normal vectors of the triangular facet vertices to reconstruct a three-dimensional surface model composed of 5000 triangular facets for expressing the organ structure; S24: optimizing the organ surface model, using a quadratic error metric algorithm to reduce the number of triangular facets to 1000, and improving the uniformity of the surface model through Laplacian smoothing and re-meshing; this step of 3D surface model reconstruction uses the 3D segmentation network to extract the abdominal organ region, reconstructs the 3D surface model through the Marching Cubes algorithm, simplifies the triangular mesh using the quadratic error metric, and optimizes the model surface uniformity to provide a lightweight and high-precision geometric basis for biomechanical modeling and registration; S25: combining the optimized surface model, using a co-rotational finite element modeling method to construct a biomechanical model.
[0020] S25 specifically includes: S251: using a quasi-static integral scheme, assuming that the inertia force component is much smaller than the elastic component, regarding the soft tissue as an elastic body, and based on Newton's second law, focusing on static equilibrium and ignoring inertia force, resulting in a whole biomechanical model where M is the mass, f represents the sum of internal and external forces corresponding to the deformation state u and .
[0021] S252: Select a tetrahedron as the basic unit of the organ biomechanical model, for each tetrahedron element e, construct its deformation equation as , wherein is the displacement vector of the vertex, is the force acting on the element, is the stiffness matrix of the element, expressed as , wherein V is the total volume, is the volume of an element, is the strain-displacement matrix, is the stress-strain matrix, , the stress-strain matrix is derived from the Young's modulus and Poisson's ratio , wherein , E and ν respectively represent the Young's modulus and Poisson's ratio; this step uses the co-rotational finite element method to construct the soft tissue biomechanical model, based on the quasi-static assumption, taking the tetrahedral element as the basic element, describing the tissue mechanical properties through the stiffness matrix and the stress-strain matrix, introducing the rotation matrix to separate the rigid rotation and pure deformation, effectively dealing with large deformation problems, and ensuring the objectivity of stress calculation, the model supports the nonlinearity of the material such as the viscoelasticity of the liver tissue, the element stiffness matrix is derived through the strain energy density function, accurate modeling of complex deformation such as compression, tension and shear is realized, and physical basis is provided for intraoperative organ accurate registration and real-time deformation prediction.
[0022] S3 specifically comprises: S31: Before the operation starts, use the binocular camera to take multi-angle calibration board shots, use the OpenCV function cv: findChessboardCorners to detect the checkerboard corner points in each image, then for each detected corner point, extract its three-dimensional position in the world coordinate system, and establish the mapping relationship between the two-dimensional image coordinates and the three-dimensional world coordinates; use the cv: calibrateCamera function to pass all the collected three-dimensional object points and corresponding two-dimensional image points to the function to calculate the internal parameters and distortion coefficients of the camera; then, distortion correction is performed, using the cv: undistort function, according to the camera intrinsic parameter matrix and distortion coefficient obtained in the previous step, the original image is corrected for distortion; finally, evaluate the calibration by calculating the reprojection error to evaluate the accuracy of the calibration, complete the general calibration of the binocular camera; this step eliminates lens distortion through OpenCV camera calibration, checkerboard corner detection, internal parameter / distortion coefficient calculation; so that the binocular camera can accurately capture the precise position of the organ, and establish an accurate, stable and quantifiable three-dimensional visual reference for the entire surgical navigation system; S32: Based on the initial frame of the intraoperative endoscope image, a semi-global block matching algorithm is used to match the left and right images of the binocular camera. First, use the horizontal Sobel operator to perform edge detection on the left and right images captured by the calibrated binocular camera, obtain the gradient image, and perform matching cost calculation. For each pixel, calculate its matching cost with the corresponding pixel under different disparities. Use the absolute difference as the cost function, which is composed of two parts: the BT cost (Birchfield-Tomasi cost) calculated from the gradient image after preprocessing and the BT cost calculated from the original image without preprocessing. Then add them together to minimize the energy function. Use the dynamic programming method to calculate the cumulative cost along the 8 directions and take the minimum value as the final cost. Generate a disparity map and use post-processing methods to repair abnormal values or holes in the disparity map. After median filtering, the initial point cloud is reconstructed in three dimensions; S33: Convert the depth information of each pixel point in the disparity map to three-dimensional space coordinates to generate the three-dimensional point cloud of the abdominal organ corresponding to the initial frame; S34: Use a density adaptive registration algorithm based on Dirichlet process Gaussian mixture model to solve the registration problem using a variational Bayesian inference framework. Expand the point cloud data and model in a coarse-to-fine manner to improve the accuracy and efficiency of registration. Estimate the scanning weight of each point to obtain the initial rotation matrix and translation vector Thus, adaptive rigid registration between the preoperative three-dimensional surface model and the intraoperative three-dimensional point cloud is realized; this step generates a disparity map and reconstructs an initial point cloud using a semi-global block matching algorithm, and realizes rigid mutual matching of the preoperative 3D model and the intraoperative point cloud by using a Dirichlet process Gaussian mixture model, and a rotation matrix and a translation vector are obtained.
[0023] S4 specifically comprises: S41: based on the intraoperative endoscopic image, using the SiamMask target tracking algorithm to delineate the target region, using the binocular lens video data collected in the S1 step and the abdominal organ region label fine-tuning model weight, using the preselected initial bounding box as the input, quickly extracting the mask of the abdominal organ region in the intraoperative image.
[0024] S42: constructing a feature point extraction network to extract features from the abdominal organ region to obtain feature point set coordinates and description vectors; S43: non-rigid matching of the tetrahedral vertexes of the biomechanical model in S2 as control points with the feature point set to realize modeling and correction of the elastic deformation of the organ; this step introduces control points through elastic compensation for organ deformation, marks the fixed constraint region, minimizes the cost function containing the stiffness matrix, compensates for the preoperative-intraoperative elastic deformation difference, and realizes preliminary alignment of the anatomical structure, laying a foundation for dynamic tracking.
[0025] The specific implementation process of the above feature point extraction network is as follows: A graph convolutional neural network (SA-GCNPoint) with self-attention mechanism is constructed to efficiently extract feature points of the abdominal organ; the overall architecture and training method of this network refer to the SuperPoint network; a single shared encoder is included for processing and reducing the dimension of the input image, and the backend is decoupled into two decoders "heads" which respectively perform interest point detection and interest point description; the encoder is improved and optimized, and the original VGG encoder is replaced by a graph convolutional network with self-attention mechanism (i.e. 4 groups of Transformer and 4 groups of graph convolutional network), which respectively extracts features of the abdominal organs, while the configuration of the decoder remains unchanged; by fusing the self-attention mechanism with the graph convolution, this encoder can capture the complex relationships between all nodes in the organ graph at once, obtain the global context, and efficiently solve the long-distance dependency problem, thus accurately modeling the interaction relationship between organs in a large receptive field; this network uses a self-supervised learning method, which can automatically learn the feature points of the abdominal organ without manual annotation data; The specific calculation process of the above non-rigid matching is as follows: Each control point in the biomechanical control point set is matched with a feature point in the feature point set according to the minimum Hamming distance. Match and determine corresponding points , and establish displacement constraints between each pair of matching points, and then construct the following cost function: ;in, is the stiffness coefficient of the organ model, which is of the same order of magnitude as Young's modulus. is the number of control points, are the initial rotation matrix and translation vector obtained in S3 respectively, Based on the element stiffness matrix The global stiffness matrix obtained by the grouping; by globally minimizing the cost function, the optimal displacement vector is solved , thereby achieving accurate estimation and matching of non-rigid deformation fields.
[0026] S5 specifically includes: S51: Using the target tracking algorithm and feature point extraction network described in S4, extract the same method from each subsequent frame. feature points, recorded as a set ; S52: Calculate the feature point set of the current frame based on the coordinates and description vectors of each feature point With the feature points of the previous frame The Hamming distance between the corresponding description vectors is obtained by nearest neighbor matching For the initial point ; S53: Use the cv.calcOpticalFlowPyrLK function in OpenCV to implement the Lucas Kanade optical flow algorithm to estimate the displacement of each matching feature point pair. When the displacement is greater than the mean + 3 times the standard deviation, it is considered to be an abnormal optical flow and removed. For the feature point set; S54: The remaining Unmatched feature points are matched with the organ biomechanical model according to the non-rigid matching method described in claim 7, and the following cost function is constructed and solved. ,in The final cross-frame feature point displacement estimation is obtained by optimization solving, and the dynamic tracking of the abdominal organ is realized. In this step, the SiamMask fine-tuning model is used to segment the intraoperative abdominal organ region in real time, the LucasKanade optical flow method is used to track the displacement of the feature points, the reliable feature point set is obtained by combining the distance threshold filtering, the 3D coordinates of the feature points are reconstructed by using the triangulation method, the distance field function is calculated as the displacement constraint of the biomechanical model, the node displacement vector is dynamically updated by minimizing the cost function containing the global stiffness matrix, the Newton acceleration method is used to optimize the cost function, and the step size and momentum are adjusted dynamically to converge, so that the registration efficiency is improved, the real-time registration of the preoperative model and the intraoperative deformed organ is realized, and the dynamic tracking with sub-millimeter accuracy is supported. The dynamic change of the key position of the organ is continuously locked, and real-time, accurate and robust navigation basis is provided for doctors or surgical robots.
[0027] S6 specifically comprises: On two different clinical liver open-source data sets 3Dircadb and DePoLL, the registration error is evaluated, the root mean square error (RMSE) is used to measure the average Euclidean distance between the corresponding point sets after registration, is more sensitive to larger errors, and reflects the overall registration accuracy; the mean absolute error (MAE) is calculated to calculate the average absolute value of the corresponding point distance, which is less affected by outliers, and reflects the stability and robustness of the registration result; the maximum error (Error) takes the maximum distance in all corresponding point pairs, directly reflects the registration deviation in the worst case, and is a key indicator for evaluating the robustness of the algorithm; the effectiveness of the proposed algorithm is evaluated by calculating the root mean square error, the mean absolute error and the maximum error between the point cloud and the 3D model, and the six methods of ICP, DCP, RPM-Net, ROPNet, PREDATOR and UTOPIC are compared in the experiment; the results show that the proposed method has the optimal performance on the two data sets, the point cloud registration effect approaches the true value, and the effectiveness and robustness of the preoperative-intraoperative dynamic registration and real-time tracking of liver deformation are verified.
[0028] Table 1 Experimental results of different initial registration algorithms of the application on the 3Dircadb liver CT data set; the error indicators are in millimeters, and the lower the value, the better the performance; the method proposed in the patent achieves the lowest value in all indicators
[0029] Table 2 Experimental results of different initial registration algorithms of the application on the DePoLL laparoscopic liver surgery data set; the error indicators are in millimeters, and the lower the value, the better the performance; the method proposed in the patent achieves the lowest value in all indicators
[0030] The above merely describes preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art, according to the technical solution and inventive concept of the present application, makes equivalent replacement or change within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.
Claims
1. A binocular camera-based graph convolution-enhanced real-time abdominal organ tracking method, comprising the following steps: S1: Obtain the patient's preoperative abdominal CT scan image and laparoscopic intraoperative image data collected by a binocular camera, and mark the target abdominal organs in the preoperative CT image and intraoperative laparoscopic image respectively; S2: Based on preoperative CT images, a 3D surface model of the abdominal organs is constructed using a 3D organ segmentation network, and a biomechanical model of the abdominal organs is established using a co-rotational finite element modeling method; S3: Based on the intraoperative laparoscopic image, the abdominal 3D point cloud of the initial frame is extracted, and it is rigidly registered with the preoperative 3D surface model to estimate the rotation and translation matrix; S4: extracting the organ feature point set from the initial frame 3D point cloud after rigid registration, estimating the deformation field of the feature point cloud in combination with the elastic constraint of the biomechanical model, and realizing elastic compensation to capture the actual morphology and displacement of the organ; S5: Based on the feature point set of the initial frame, cross-frame matching is performed on the feature points in subsequent frames with the initial frame to achieve dynamic tracking of abdominal organs throughout the entire surgical process; S6: Performing quantitative and qualitative evaluation on the dynamic tracking results to verify the accuracy and robustness of the real-time abdominal organ tracking method.
2. According to the method for real-time tracking of abdominal organs using graph convolution enhancement based on binocular cameras in claim 1, S2 specifically comprises: S21: Preprocessing the acquired preoperative CT image, wherein the preprocessing includes cropping the image data into image blocks with a size of 256×256×256 and normalizing the image data; S22: Build a 3D organ segmentation network for preoperative images based on the nnUNet framework to extract the mask of the abdominal organ region from the preprocessed preoperative CT images; S23: inputting the three-dimensional mask of the abdominal organ into the Marching Cubes algorithm to reconstruct a three-dimensional surface model consisting of 5000 triangular facets to express the organ structure; S24: Optimizing the organ surface model, reducing the number of triangles to 1000 using a quadratic error metric algorithm, and improving the uniformity of the surface model through Laplace smoothing and re-meshing. S25: Combined with the optimized surface model, the co-rotational finite element modeling method was used to construct the biomechanical model.
3. The method for real-time tracking of abdominal organs using graph convolution enhancement based on a binocular camera according to claim 2, wherein S25 specifically comprises: S251: A quasi-static integration scheme is used, based on Newton's second law, focusing on static equilibrium and ignoring inertial forces, resulting in an overall biomechanical model , where M is the mass, f represents the state corresponding to the deformation u and The sum of internal and external forces; S252: Tetrahedron is selected as the basic unit of the organ biomechanics model. For each tetrahedron element e, its deformation equation is constructed as follows: ,in is the displacement vector of the vertex, is the force acting on the element, The stiffness matrix of the element is expressed as , where V is the total volume, is the volume of an element, is the strain-displacement matrix, is the stress-strain matrix, The rotation matrix representing the current state of the tetrahedron relative to the initial state; The stress-strain matrix It is derived from Young's modulus and Poisson's ratio; S253: Perform dynamic solution through time-stepping linear method to construct the unit and overall biomechanical model.
4. The method for real-time tracking of abdominal organs using graph convolution-enhanced binocular cameras according to claim 1, wherein S3 specifically comprises: S31: Before the operation begins, use the binocular camera to shoot a multi-angle calibration plate, establish a mapping relationship between two-dimensional image coordinates and three-dimensional world coordinates through corner detection, calculate the camera's internal parameters and distortion coefficients, and complete the universal calibration of the binocular camera; S32: Based on the initial frame of the intraoperative laparoscopic image, a semi-global block matching algorithm is used to match the binocular left and right images to generate a disparity map; S33: converting the depth information of each pixel point in the disparity map into three-dimensional space coordinates to generate a three-dimensional point cloud of the abdominal organs corresponding to the initial frame; S34: Using a density adaptive registration algorithm based on a Dirichlet process Gaussian mixture model, the three-dimensional point cloud is aligned with the preoperative three-dimensional surface model described in S2 in terms of overall position, direction, and scale, and a rotation matrix is calculated. and translation vectors , achieving rigid registration.
5. The method for real-time tracking of abdominal organs using graph convolution-enhanced binocular cameras according to claim 1, wherein S4 specifically comprises: S41: Based on the intraoperative laparoscopic images, the SiamMask target tracking algorithm is used to obtain the mask of the abdominal organ area in the intraoperative images; S42: Construct a feature point extraction network to extract features from the abdominal organ region and obtain the coordinates and description vectors of the feature point set; S43: Use the vertices of the tetrahedron of the biomechanical model described in S2 as control points With the feature point set Perform non-rigid matching to achieve modeling and correction of organ elastic deformation.
6. According to the binocular camera-based graph convolution-enhanced real-time abdominal organ tracking method of claim 5, the feature point extraction network specifically includes an encoder and two decoders that perform feature point detection and feature point description, respectively; the encoder is composed of a graph convolutional network with a self-attention mechanism, and the decoder is composed of a convolutional neural network.
7. According to the binocular camera-based graph convolution enhanced real-time abdominal organ tracking method of claim 5, the non-rigid matching specifically comprises: The biomechanical control point set Each control point in Based on the minimum Hamming distance and feature point set Match and determine corresponding points , and establish displacement constraints between each pair of matching points, and then construct the following cost function: , ;in, is the stiffness coefficient of the organ model, which is of the same order of magnitude as Young's modulus. is the number of control points, are the initial rotation matrix and translation vector obtained in S3 respectively, Based on the element stiffness matrix The global stiffness matrix obtained by the grouping; by globally minimizing the cost function, the optimal displacement vector is solved , thereby achieving accurate estimation and matching of non-rigid deformation fields.
8. The method for real-time tracking of abdominal organs using graph convolution enhancement based on binocular cameras according to claim 1, wherein S5 specifically comprises: S51: Using the target tracking algorithm and feature point extraction network described in S4, extract the same method from each subsequent frame. feature points, recorded as a set ; S52: Calculate the feature point set of the current frame based on the coordinates and description vectors of each feature point With the feature points of the previous frame The Hamming distance between the corresponding description vectors is obtained by nearest neighbor matching For the initial point ; S53: Use the optical flow method to estimate the displacement of each matching feature point pair, and remove the feature points whose displacement is greater than the mean plus three times the standard deviation, to obtain For the feature point set; S54: The remaining Unmatched feature points are matched with the organ biomechanical model according to the non-rigid matching method described in claim 7, and the following cost function is constructed and solved. ,in ; The final cross-frame feature point displacement estimation is obtained through optimization solution to achieve dynamic tracking of abdominal organs.
9. The method for real-time abdominal organ tracking based on binocular cameras enhanced by graph convolution according to claim 8, wherein the optimization process adopts Newton acceleration method and accelerates convergence by dynamically adjusting step size and momentum, specifically comprising the following steps: Select the starting point And set the learning rate , in each iteration Execute in sequence: Calculate the predicted position , calculate the objective function in The gradient at , and update the momentum , and finally update the parameters according to ; until the convergence conditions are met, the displacement estimation solution process is accelerated.
10. The method for real-time tracking of abdominal organs using graph convolution enhancement based on binocular cameras according to claim 1, wherein S6 specifically comprises: The registration error was evaluated on two different clinical liver open source datasets, 3Dircadb and DePoLL. The effectiveness of the proposed algorithm was evaluated by calculating the root mean square error, mean absolute error, and maximum error between the point cloud and the biomechanical model. The six methods, including ICP, DCP, RPM-Net, ROPNet, PREDATOR, and UTOPIC, were experimentally compared.
Citation Information
Patent Citations
Tumor placeholder brain function partition nerve influence image segmentation method based on deep learning model
CN120526142A
Locally rigid vessel based registration for laparoscopic liver surgery
GB201506842D0
Tomography-Based and MRI-Based Imaging Systems
US20110142316A1