Intraoperative residual organ volume estimation system for surgical planning assistance

By registering preoperative and intraoperative tissue mesh models in minimally invasive surgery, and combining physician annotations with binocular endoscopic depth estimation, the problem of inaccurate measurement of residual organ volume during surgery was solved, achieving precise volume measurement and surgical planning assistance.

CN116993805BActive Publication Date: 2025-11-14HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310419428.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-14
Publication Date
2025-11-14
Estimated Expiration
2043-04-14

AI Technical Summary

Technical Problem

Current technology cannot accurately measure the volume of remaining organs during minimally invasive surgery, resulting in inaccurate measurement results and failing to provide effective intraoperative guidance for doctors.

Method used

The preoperative tissue mesh model is registered with the intraoperative tissue mesh model through the registration module. Combined with the doctor's annotation, the volume of the remaining organs is obtained. The intraoperative tissue mesh model is obtained by an online self-supervised learning depth estimation method based on binocular endoscopy. Multi-level features and attention mechanism are used for accurate registration and volume calculation.

Benefits of technology

It enables precise intraoperative measurement of remaining organ volume, reduces interference from complex tissue deformation and invisible areas on volume prediction, and provides highly referential and actively selective volume measurement information to assist doctors in adjusting surgical plans and improve surgical efficiency and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116993805B_ABST
    Figure CN116993805B_ABST
Patent Text Reader

Abstract

This invention provides an intraoperative residual organ volume estimation system for surgical planning assistance, relating to the field of minimally invasive surgery. By matching the preoperative and intraoperative tissue mesh models, this invention transforms the measurement of residual organ volume into a volume measurement of the corresponding region in the preoperative tissue mesh model, avoiding interference from complex tissue deformation and invisible areas on volume prediction. Furthermore, it considers intraoperative interaction with the surgeon, receiving simple annotations from the surgeon regarding regions of interest to obtain accurate volume measurement information, achieving active selectivity and high reference value during surgery. In addition, the introduced online self-supervised learning depth estimation method based on binocular endoscopy utilizes a binocular depth estimation network with rapid overlearning capabilities, continuously adapting to new scenarios using self-supervised information, thereby ensuring the accuracy of the intraoperative tissue mesh model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of minimally invasive surgery, and more specifically to an intraoperative residual organ volume estimation system for surgical planning assistance. Background Technology

[0002] Compared to traditional open surgery, minimally invasive surgery (such as endoscopic surgery) has advantages such as smaller incisions, less bleeding, and faster recovery, and is gradually being widely adopted. There are many evaluation indicators for minimally invasive surgery, including operation time, intraoperative blood loss, postoperative complication rate, and recovery time. However, for organ removal surgery, the volume of remaining organ is one of the important evaluation indicators.

[0003] Depending on the type of organ and the specific circumstances, there are currently several methods for measuring the remaining volume of an organ. The following are some commonly used technical solutions for measuring the remaining volume of an organ:

[0004] (1) Preoperative or postoperative CT and MRI scans: By performing three-dimensional CT or MRI scans, the volume of organs such as the liver and lungs can be calculated, thereby assessing the volume and proportion of remaining organs before and after surgery. However, this method cannot accurately distinguish between organs and surrounding tissues, that is, the boundaries of the measured organs may be unclear, resulting in inaccurate measurement results.

[0005] (2) Intraoperative ultrasound examination: Ultrasound examination is a non-invasive examination method that can assess the size and shape of organs such as the liver, gallbladder, and pancreas, thereby indirectly calculating the remaining volume. However, firstly, this method requires doctors to have a certain level of skill and experience, and the operational level and experience of different doctors may affect the accuracy of the measurement results; secondly, the measurement of organs is affected by factors such as depth, angle, distance, and scanning plane, which can easily lead to errors; finally, there are limitations in the measurement of some organs, such as the measurement of deep organs such as the heart and lungs, which are obstructed by the ribs, making it difficult to obtain accurate volume data.

[0006] (3) Direct postoperative measurement method: This method involves directly removing the organ during surgery or dissection and immersing it in a liquid, then measuring the amount of liquid displacement to estimate the organ's volume. However, this method can only be used for organs removed postoperatively and cannot be applied to the volume measurement of living organs, nor can it provide doctors with intraoperative guidance or prompts.

[0007] Therefore, it is necessary to provide a more accurate intraoperative residual organ volume estimation system. Summary of the Invention

[0008] (a) Technical problems to be solved

[0009] To address the shortcomings of existing technologies, this invention provides an intraoperative residual organ volume estimation system for surgical planning assistance, which solves the technical problem of inaccurate intraoperative measurement results.

[0010] (II) Technical Solution

[0011] To achieve the above objectives, the present invention provides the following technical solution:

[0012] An intraoperative residual organ volume estimation system for surgical planning assistance includes:

[0013] The registration module is used to register the preoperative tissue mesh model and the intraoperative tissue mesh model to obtain an overall tissue mesh model that displays the internal tissue information of the preoperative tissue mesh model in the intraoperative tissue mesh model.

[0014] Specifically, the intraoperative tissue mesh model is obtained based on the depth value of the specified binocular endoscopic image frame;

[0015] The first acquisition module is used to receive the region to be excised marked by the doctor on the region of interest of the specified binocular endoscope image frame, and to acquire the corresponding vertex set of all pixels in the region to be excised in the overall tissue mesh model.

[0016] The second acquisition module is used to acquire and visualize the corresponding vertex set on the preoperative tissue mesh model based on the corresponding vertex set in the overall tissue mesh model.

[0017] The solution module is used to receive the cutting direction and manually selected candidate regions marked by the doctor on the preoperative tissue mesh model, and combine them with the corresponding vertex set on the preoperative tissue mesh model to obtain the local volume of the organ corresponding to the region to be removed, and finally obtain the remaining organ volume or the percentage of the remaining organ volume.

[0018] Preferably, the registration module includes:

[0019] The first modeling unit is used to obtain a preoperative tissue mesh model with tissue semantic information;

[0020] The second modeling unit is used to obtain an intraoperative tissue mesh model based on the depth value of a specified binocular endoscopic image frame.

[0021] The feature extraction unit is used to obtain corresponding multi-level features based on the preoperative tissue mesh model and the intraoperative tissue mesh model, respectively.

[0022] The overlap prediction unit is used to obtain the overlapping area of ​​the preoperative tissue mesh model and the intraoperative tissue mesh model based on the multi-level features, and to obtain the pose transformation relationship of the vertices of the preoperative tissue mesh model within the overlapping area.

[0023] The global fusion unit is used to obtain the registered vertex coordinates of the preoperative tissue mesh model based on the coordinates and pose transformation relationship of the vertices in the overlapping region and the coordinates of the vertices in the non-overlapping region.

[0024] The information display unit is used to obtain the overall tissue grid model that displays the internal tissue information of the preoperative tissue grid model in the intraoperative tissue grid model based on the coordinates of all vertices after registration of the preoperative tissue grid model.

[0025] Preferably, the feature extraction unit uses Chebyshev spectral map convolution to extract multi-level features of the preoperative tissue mesh model and the intraoperative tissue mesh model:

[0026]

[0027]

[0028] Among them, the preoperative tissue mesh model M is defined. pre =(V pre E pre V pre E represents the three-dimensional coordinates of the vertices of the preoperative tissue mesh model. pre The edges between vertices in the preoperative tissue mesh model; the intraoperative tissue mesh model M in =(V in E in V in E represents the three-dimensional coordinates of the vertices of the preoperative tissue mesh model. in Represents the edges between vertices of the intraoperative tissue mesh model;

[0029] and Let the downsampling scale features of the (n+1)th and nth layers of the preoperative tissue model be represented respectively, and initialized. For V pre ; and Let the features of the (n+1)th and nth layers of the intraoperative tissue model be represented respectively, and initialized. For V in ;

[0030] The b-th order Chebyshev polynomials calculated from their respective vertices and their B-ring neighborhoods. They are respectively from edge E in E pre Calculate the scaled Laplacian matrix. These are the learning parameters of the neural network;

[0031] And / or the overlap prediction unit is specifically used for:

[0032] An attention mechanism is used to obtain the overlapping region of the preoperative tissue mesh model and the intraoperative tissue mesh model, including:

[0033]

[0034]

[0035] Among them, O pre Represents the preoperative tissue mesh model M pre Mask of overlapping regions; O in Intraoperative tissue mesh model M in The mask for the overlapping region; cross and self represent the self-attention and cross-attention operations, respectively; and These represent the m-th downsampling scale features of the vertices of the preoperative tissue mesh model and the intraoperative tissue mesh model, respectively.

[0036] According to mask O pre and O in Get the vertices that are within the overlapping region. and its characteristics The preoperative tissue mesh model M was calculated using a multilayer perceptron (MLP). pre Vertex in Corresponding point:

[0037]

[0038] in, It is the intraoperative tissue mesh model M in The vertices in the model correspond to the preoperative tissue mesh model M. pre Vertex in This indicates the calculation of cosine similarity. This indicates that the position encoding operation is performed on the vertices of the intraoperative tissue mesh model that are within the overlapping area;

[0039] Vertex construction using KNN (Knowledge Neighbor Network) For the local neighborhood, the rotation matrix is ​​solved using Singular Value Decomposition (SVD), as shown in the following formula:

[0040]

[0041] in, Represents vertices The rotation matrix; This indicates that the KNN algorithm is used to construct the vertex... A local neighborhood; It is the vertex of the preoperative tissue mesh model. neighborhood points, It corresponds to the neighborhood point The vertices of the intraoperative tissue mesh model;

[0042] Using rotation matrix Change the point cloud coordinates to obtain Using MLP to predict vertices The displacement vector is given by the following formula:

[0043]

[0044] in, The displacement vectors of the vertices in the overlapping region of the preoperative tissue mesh model are compared with the rotation matrix. This constitutes the pose transformation relationship;

[0045] And / or the global fusion unit is specifically used for:

[0046] The rotation matrix and displacement vector of all vertices of the preoperative tissue mesh model were obtained by MLP regression:

[0047]

[0048] Among them, R pre ,t pre These represent the rotation matrix and displacement vector of all vertices in the preoperative tissue mesh model, respectively; Indicates based on vertices within the overlapping region All vertices v of the preoperative tissue mesh model pre The weights for distance calculation;

[0049]

[0050] in, This represents the coordinates of all vertices after registration of the preoperative tissue mesh model.

[0051] Preferably, during the training phase of the intraoperative residual organ volume estimation system, a training set is generated based on real data:

[0052] Based on the specified feature point pairs between the binocular endoscopic image frames and the preoperative tissue mesh model, a non-rigid algorithm is used to register the preoperative and intraoperative tissue mesh models using these feature points. For any given feature point:

[0053]

[0054] Wherein, Non_rigid_ICP represents the non-rigid registration algorithm ICP. This represents the a-th feature point in the preoperative tissue mesh model used for non-rigid registration. correspond Feature points of the intraoperative tissue mesh model, T G T is the global transition matrix of the preoperative tissue mesh model. l,a It belongs to feature point v pre,a The local deformation transfer matrix;

[0055] The local deformation transfer matrix T of all vertices in the preoperative tissue mesh model was obtained by using quaternion interpolation. l Vertices v in the preoperative tissue mesh model are obtained by transforming the relationships. pre Registered coordinate labels

[0056] Preferably, during the training phase of the intraoperative residual organ volume estimation system, the following supervised loss function is constructed:

[0057]

[0058] Among them, Loss s This represents the supervised loss function during the training phase.

[0059] β s γ s These represent the coefficients of the supervised loss term;

[0060] N1 represents the preoperative combined mesh model M. pre The number of vertices;

[0061] This represents the L2 ground truth loss based on a manually labeled dataset. Represents the coordinates of all vertices after registration of the preoperative tissue mesh model;

[0062] I c +II c +III c I represents the Cauchygreen invariant, used to constrain the degree of tissue deformation in vivo. c The arc distance between two points on the constrained surface remains constant, II c The surface area of ​​the constrained tissue remains unchanged, III c The volume of the constrained tissue remains constant.

[0063] Preferably, the registration module further includes:

[0064] The precision fine-tuning unit is used to introduce an unsupervised loss fine-tuning network to assist the global fusion unit in obtaining the registered vertex coordinates of the preoperative tissue mesh model.

[0065] And / or the unsupervised loss fine-tuning network described above, during application, constructs the following unsupervised loss function:

[0066]

[0067] Among them, Loss u Represents the unsupervised loss function;

[0068] β u ,γ u They represent the coefficients of the unsupervised loss term, and These are the vertex coordinates after registration of the preoperative tissue mesh model during unsupervised training. This represents the vertices of the preoperative tissue mesh model after distance registration in the intraoperative tissue mesh model. The closest point, Represents vertices and European distance, This represents the distance v from the vertex of the intraoperative tissue mesh model in the preoperative tissue mesh model after registration. in,b The closest point, Represents vertex v in,b and vertex Euclidean distance;

[0069] N1 represents the preoperative tissue mesh model M. pre The number of vertices, N2 represents the intraoperative tissue mesh model M. in The number of vertices;

[0070] Describes the Cauchygreen invariant during unsupervised training. The arc distance between two points on the constrained surface remains constant. The surface area of ​​the constrained tissue remains unchanged. The volume of the constrained tissue remains constant.

[0071] Preferably, all pixels within the area to be excised are defined in the intraoperative tissue mesh model M. in The vertex set on is P s All pixels within the area to be removed are in the overall tissue mesh model M. trans The corresponding vertex set in is P trans Where any vertex is p trans The preoperative tissue mesh model M pre The corresponding vertex set on is P ct Where any vertex is p ct ;

[0072] The first acquisition module is specifically used to obtain P sThe nearest neighbor algorithm is used to obtain the nearest neighbor in M. trans P on trans Any vertex p in trans ;

[0073] And / or the second acquisition module is used to obtain R pre t pre and p trans Obtain and visualize M pre P on ct Any vertex p in ct ;

[0074] p ct =R pre p trans +t pre .

[0075] Preferably, the solution module is specifically used for:

[0076] Traversing P ct vertex p in ct Obtain the candidate region within the manually selected area and the cutting direction v. pro Find the nearest vertices, calculate the average of these vertices, and obtain the corresponding vertex p′. ct The coordinates of the points form the corresponding vertex set P′. ct ;

[0077] According to vertex set P ct and P′ ct Pairing spatial point p ct and p′ ct Multiple irregular pentahedrons are generated within the region S to be excised.

[0078] Calculate the volume of each irregular pentahedron and sum them to obtain the local volume V of the organ corresponding to the region to be removed. cut Combined with the overall volume V of the organ before surgery all Finally, the remaining organ volume V was obtained. remain =V all -V cut or percentage of remaining organ volume

[0079] The steps for calculating the volume of any pentahedron A1A2A3B1B2B3 are as follows:

[0080] Draw auxiliary lines B1′A2 and A2B3′ along the direction parallel to B1B2 and B2B3 to divide the irregular pentahedron into a triangular prism B1B2B3B1′A2B3′ and a square pyramid A1A3B3′B1′A2;

[0081] Calculate the area of ​​plane B1B2B3 and the perpendicular distance h between plane A2 and plane B1B2B3. prism Obtain the volume of the triangular prism B1B2B3B1′A2B3′:

[0082] V prism =S ΔB1B2B3 ×h prism

[0083] Calculate the area of ​​plane A1A3B3′B1′ and the perpendicular distance h between A2 and plane A1A3B3B1. pyramid Obtain the volume of the square pyramid A1A3B3′B1′A2:

[0084]

[0085] Then the volume V of each pentahedron final for:

[0086] V final =V prism +V pyramid .

[0087] Preferably, the second modeling unit uses an online self-supervised learning depth estimation method based on binocular endoscope to obtain the depth value of the specified binocular endoscope image frame; the binocular depth estimation network used by the online self-supervised learning depth estimation method has the ability to quickly overlearn and can continuously adapt to new scenes using self-supervised information;

[0088] In real-time reconstruction mode, the second modeling unit is specifically used to overfit continuous video frames to obtain the depth value of a specified binocular endoscopic image frame, including:

[0089] Extraction subunits are used to acquire binocular endoscope images, and the encoder network of the current binocular depth estimation network is used to extract multi-scale features of the current frame image;

[0090] The fusion subunit is used to fuse multi-scale features using the decoder network of the current binocular depth estimation network to obtain the disparity of each pixel in the current frame image;

[0091] The conversion subunit is used to convert parallax into depth based on camera intrinsic and extrinsic parameters and output it as the result of the current frame image.

[0092] The first estimation subunit is used to update the parameters of the current stereo depth estimation network using self-supervised loss without introducing external ground truth, for depth estimation of the next frame image.

[0093] Preferably, in the precise measurement mode, the second modeling unit is specifically used to overfit key image video frames, including:

[0094] The second estimation subunit, without introducing external truth values, updates the parameters of the aforementioned binocular depth estimation network in real-time reconstruction mode based on the binocular depth estimation network obtained from the previous frame of the specified binocular endoscopic image frame using the self-supervised loss corresponding to the specified binocular endoscopic image frame until convergence, and uses the converged binocular depth estimation network to accurately estimate the depth of the specified binocular endoscopic image frame, thereby obtaining the depth value of the specified binocular endoscopic image frame.

[0095] (III) Beneficial Effects

[0096] This invention provides an intraoperative residual organ volume estimation system for surgical planning assistance. Compared with existing technologies, it has the following advantages:

[0097] This invention transforms the measurement of remaining organ volume into the volume measurement of the corresponding region in the preoperative tissue mesh model by matching the preoperative tissue mesh model with the intraoperative tissue mesh model, thus avoiding interference from complex tissue deformation and invisible areas on volume prediction. Furthermore, it considers the interaction between the surgeon and the patient during the operation, receiving simple annotations from the surgeon for the region of interest to obtain accurate volume measurement information, achieving active selectivity and high reference value during the operation. Attached Figure Description

[0098] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0099] Figure 1 This is a structural block diagram of an intraoperative residual organ volume estimation system for surgical planning assistance provided by an embodiment of the present invention;

[0100] Figure 2 A schematic diagram illustrating the organ volume corresponding to the area to be removed, provided in an embodiment of the present invention;

[0101] Figure 3 A schematic diagram illustrating the volume calculation of an irregular pentahedron provided in an embodiment of the present invention;

[0102] Figure 4 This is a schematic diagram illustrating the technical framework of an online self-supervised learning depth estimation method based on binocular endoscopy, provided in an embodiment of the present invention. Detailed Implementation

[0103] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0104] This application provides an intraoperative residual organ volume estimation system for surgical planning assistance, which solves the technical problem of inaccurate intraoperative measurement results.

[0105] The technical solution in this application is to solve the above-mentioned technical problems, and the general idea is as follows:

[0106] The embodiments of this invention are primarily applied to, but not limited to, surgical endoscopic scenarios such as laparoscopic surgery. As mentioned in the background section, existing solutions either can only be verified postoperatively and cannot measure the volume of removed organs in real time during surgery to provide doctors with prompts and guidance; or they rely on expensive intraoperative equipment and highly skilled surgeons, making them unsuitable for large-scale deployment. Therefore, the embodiments of this invention propose an intraoperative residual organ volume estimation system for surgical planning assistance, based on a fast, non-invasive, and low-cost algorithm.

[0107] In laparoscopic surgery, surgeons can only see the surface of tissues, while information such as the location of blood vessels and lesion areas within the tissue relies on their experience and judgment. Preoperative CT / MRI reconstruction models contain information about blood vessels and lesion areas within the tissue. Non-rigid registration and fusion algorithms can register the preoperative tissue mesh model to the intraoperative tissue mesh model and present internal tissue information to the surgeon using conventional display techniques, assisting in clinical decision-making, reducing surgical risks, and improving surgical efficiency. Without relying on additional equipment, the matching relationship between the preoperative and intraoperative tissue mesh models transforms the measurement of remaining organ volume into a volume measurement of the corresponding area in the preoperative tissue mesh model, avoiding interference from complex tissue deformation and invisible areas on volume prediction.

[0108] Furthermore, through simple interaction with the doctor during the operation, the volume of the organ to be cut can be measured in a timely manner, and important information such as the scope of organ resection and the functional areas to be preserved can be determined. This helps the doctor to better adjust the surgical plan, assess surgical risks and predict surgical results, while also helping to ensure the safety and effectiveness of the operation.

[0109] Furthermore, an intraoperative tissue mesh model can be obtained based on the depth values ​​of specified binocular endoscopic image frames. Specifically, an online self-supervised learning depth estimation method based on binocular endoscopy can be used to obtain the depth values ​​of the specified binocular endoscopic image frames. The binocular depth estimation network used in this online self-supervised learning depth estimation method has the ability to quickly overlearn and can continuously adapt to new scenarios using self-supervised information. Moreover, the online self-supervised learning depth estimation method also provides two modes: a real-time reconstruction mode and a precise measurement mode, for determining the depth values ​​of the specified binocular endoscopic image frames.

[0110] The dual-mode switching depth estimation can provide real-time point clouds of intraoperative anatomical structures to help doctors intuitively understand the intraoperative three-dimensional structure. It can also achieve high-precision reconstruction of the binocular endoscopic image frames specified by the doctor based on single-frame overfitting, providing a basis for subsequent processing and balancing speed and accuracy in application.

[0111] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0112] Example:

[0113] like Figure 1 As shown, embodiments of the present invention provide

[0114] An intraoperative residual organ volume estimation system for surgical planning assistance includes:

[0115] The registration module is used to register the preoperative tissue mesh model and the intraoperative tissue mesh model to obtain an overall tissue mesh model that displays the internal tissue information of the preoperative tissue mesh model in the intraoperative tissue mesh model.

[0116] Specifically, the intraoperative tissue mesh model is obtained based on the depth value of the specified binocular endoscopic image frame;

[0117] The first acquisition module is used to receive the region to be excised marked by the doctor on the region of interest of the specified binocular endoscope image frame, and to acquire the corresponding vertex set of all pixels in the region to be excised in the overall tissue mesh model.

[0118] The second acquisition module is used to acquire and visualize the corresponding vertex set on the preoperative tissue mesh model based on the corresponding vertex set in the overall tissue mesh model.

[0119] The solution module is used to receive the cutting direction and manually selected candidate regions marked by the doctor on the preoperative tissue mesh model, and combine them with the corresponding vertex set on the preoperative tissue mesh model to obtain the local volume of the organ corresponding to the region to be removed, and finally obtain the remaining organ volume or the percentage of the remaining organ volume.

[0120] This invention does not rely on additional equipment. By matching the preoperative tissue mesh model with the intraoperative tissue mesh model, the measurement of the remaining organ volume is transformed into the volume measurement of the corresponding area of ​​the preoperative tissue mesh model, avoiding interference from complex tissue deformation and invisible areas on volume prediction. Furthermore, it considers the interaction between the surgeon and the patient during the operation, receiving simple annotations from the surgeon for the area of ​​interest to obtain accurate volume measurement information, achieving active selectivity and high reference value during the operation.

[0121] The following section will detail each component module of the above technical solution:

[0122] The registration module is used to register the preoperative tissue mesh model and the intraoperative tissue mesh model to obtain an overall tissue mesh model that displays the internal tissue information of the preoperative tissue mesh model in the intraoperative tissue mesh model.

[0123] The registration module includes a first modeling unit, a second modeling unit, a feature extraction unit, an overlap prediction unit, a global fusion unit, and a precision fine-tuning unit. Specifically:

[0124] The first modeling unit is used to obtain a preoperative tissue mesh model with tissue semantic information.

[0125] For example, this unit uses software such as 3D Slicer to reconstruct CT / MRI tissues to obtain a three-dimensional mesh model. Then, deep learning algorithms such as DeepLab or manual segmentation are used to divide tissues such as blood vessels and liver, ultimately forming a preoperative tissue mesh model M with tissue semantic information. pre =(V pre E pre ), where V pre E represents the three-dimensional coordinates of the model's vertices. pre This represents the edges between vertices.

[0126] The second modeling unit is used to obtain an intraoperative tissue mesh model based on the depth values ​​of a specified binocular endoscopic image frame.

[0127] For example, this unit uses an online self-supervised learning depth estimation method based on binocular endoscopy (see below for details) to estimate the depth value D of a pixel; and calculates the spatial coordinates of the pixel in the camera coordinate system using a pinhole camera model, the formula of which is:

[0128]

[0129]

[0130] z = D

[0131] Where D is the depth estimate of the pixel; x, y, and z represent the x, y, and z coordinates in the camera coordinate system, respectively.

[0132] c x ,c y ,f x ,f y The intrinsic parameter matrix between the left or right endoscope and the camera in a binocular endoscope. The corresponding parameters in the image are used to convert the image into a point cloud V. in ={v in,a |a=1,2,…N1}, where v in,a This represents the spatial coordinates of the a-th pixel.

[0133] Finally, Delaunay triangulation was used to generate the point cloud V. in Adjacent edge E in Ultimately, an intraoperative tissue mesh model M is formed. in =(V in E in ).

[0134] The feature extraction unit is used to obtain corresponding multi-level features based on the preoperative tissue mesh model and the intraoperative tissue mesh model, respectively.

[0135] Specifically, the feature extraction unit uses Chebyshev spectral map convolution to extract multi-level features from the preoperative and intraoperative tissue mesh models:

[0136]

[0137]

[0138] Among them, the preoperative tissue mesh model M is defined. pre =(V pre E pre V pre E represents the three-dimensional coordinates of the vertices of the preoperative tissue mesh model. pre The edges between vertices in the preoperative tissue mesh model; the intraoperative tissue mesh model M in =(V in E in V in E represents the three-dimensional coordinates of the vertices of the preoperative tissue mesh model. in Represents the edges between vertices of the intraoperative tissue mesh model;

[0139] and Let the downsampling scale features of the (n+1)th and nth layers of the preoperative tissue model be represented respectively, and initialized. For V pre; and Let the features of the (n+1)th and nth layers of the intraoperative tissue model be represented respectively, and initialized. For V in ;

[0140] The b-th order Chebyshev polynomials calculated from their respective vertices and their B-ring neighborhoods. They are respectively from edge E in E pre Calculate the scaled Laplacian matrix. These are the learning parameters of the neural network.

[0141] The overlap prediction unit is used to obtain the overlapping area of ​​the preoperative tissue mesh model and the intraoperative tissue mesh model based on the multi-level features, and to obtain the pose transformation relationship of the vertices of the preoperative tissue mesh model within the overlapping area.

[0142] Specifically, the overlap prediction unit is used for:

[0143] An attention mechanism is used to obtain the overlapping region of the preoperative tissue mesh model and the intraoperative tissue mesh model, including:

[0144]

[0145]

[0146] Among them, O pre Represents the preoperative tissue mesh model M pre Mask of overlapping regions; O in Intraoperative tissue mesh model M in The mask for the overlapping region; cross and self represent the self-attention and cross-attention operations, respectively; and These represent the m-th downsampling scale features of the vertices of the preoperative tissue mesh model and the intraoperative tissue mesh model, respectively.

[0147] According to mask O pre and O in Get the vertices that are within the overlapping region. and its characteristics The preoperative tissue mesh model M was calculated using a multilayer perceptron (MLP). pre Vertex in Corresponding point:

[0148]

[0149] in, It is the intraoperative tissue mesh model M in The vertices in the model correspond to the preoperative tissue mesh model M. pre Vertex in This indicates the calculation of cosine similarity. This indicates that the position encoding operation is performed on the vertices of the intraoperative tissue mesh model that are within the overlapping area;

[0150] Vertex construction using KNN (Knowledge Neighbor Network) For the local neighborhood, the rotation matrix is ​​solved using Singular Value Decomposition (SVD), as shown in the following formula:

[0151]

[0152] in, Represents vertices The rotation matrix; This indicates that the KNN algorithm is used to construct the vertex... A local neighborhood; It is the vertex of the preoperative tissue mesh model. neighborhood points, It corresponds to the neighborhood point The vertices of the intraoperative tissue mesh model;

[0153] Using rotation matrix Change the point cloud coordinates to obtain Using MLP to predict vertices The displacement vector is given by the following formula:

[0154]

[0155] in, The displacement vectors of the vertices in the overlapping region of the preoperative tissue mesh model.

[0156] The global fusion unit is used to obtain the coordinates of all vertices of the preoperative tissue mesh model after registration, based on the coordinates and pose transformation relationship of the vertices in the overlapping region of the preoperative tissue mesh model and the coordinates of the vertices in the non-overlapping region.

[0157] Specifically, the global fusion unit is used for:

[0158] The rotation matrix and displacement vector of all vertices of the preoperative tissue mesh model were obtained by MLP regression:

[0159]

[0160] Among them, R pre ,t pre These represent the rotation matrix and displacement vector of all vertices in the preoperative tissue mesh model, respectively; Indicates based on vertices within the overlapping region All vertices v of the preoperative tissue mesh model pre The weights for distance calculation (where all vertices include vertices in overlapping regions and vertices in non-overlapping regions);

[0161]

[0162] in, This represents the coordinates of all vertices after registration of the preoperative tissue mesh model.

[0163] Accordingly, it can be clearly stated that the multi-mode fusion network proposed in this embodiment of the invention is based on grid data. It predicts the overlapping region and its displacement field through overlapping prediction units, and combines Cauchy Green invariant constraints on the non-rigid deformation of the preoperative tissue grid model, making the multi-mode fusion model more reasonable and reducing errors in multi-mode fusion.

[0164] The information display unit is used to obtain the overall tissue grid model that displays the internal tissue information of the preoperative tissue grid model in the intraoperative tissue grid model based on the coordinates of all vertices after registration of the preoperative tissue grid model.

[0165] For example, in this unit, VR glasses can be used to display the two registered 3D models in a unified coordinate system, or the registered preoperative tissue mesh model can be superimposed on the endoscopic image based on the basic principles of camera imaging. Both of these optional display methods can present internal tissue information to doctors, assist doctors in making clinical decisions, reduce surgical risks, and improve surgical efficiency.

[0166] The precision fine-tuning unit is used to introduce an unsupervised loss fine-tuning network to assist the global fusion unit in obtaining the registered vertex coordinates of the preoperative tissue mesh model.

[0167] The reason for introducing the precision fine-tuning unit is that, in this embodiment of the invention, when registering a specified binocular endoscopic image frame, the reconstructed intraoperative tissue mesh model may differ from the dataset due to differences in endoscopic lighting and individual patient characteristics. These differences may lead to a decrease in registration accuracy. Using an unsupervised loss fine-tuning network can improve the registration accuracy.

[0168] Therefore, in the application of the unsupervised loss fine-tuning network, the following unsupervised loss function needs to be constructed:

[0169]

[0170] Among them, Loss u Represents the unsupervised loss function;

[0171] β u ,γ u They represent the coefficients of the unsupervised loss term, and These are the vertex coordinates after registration of the preoperative tissue mesh model during unsupervised training. This represents the vertices of the preoperative tissue mesh model after distance registration in the intraoperative tissue mesh model. The closest point, Represents vertices and European distance, This represents the distance v from the vertex of the intraoperative tissue mesh model in the preoperative tissue mesh model after registration. in, The closest point, Represents vertex v in, and vertex Euclidean distance;

[0172] N1 represents the preoperative tissue mesh model M. pre The number of vertices, N2 represents the intraoperative tissue mesh model M. in The number of vertices;

[0173] Describes the Cauchygreen invariant during unsupervised training. The arc distance between two points on the constrained surface remains constant. The surface area of ​​the constrained tissue remains unchanged. The volume of the constrained tissue remains constant.

[0174] This invention constructs an unsupervised fine-tuning mechanism with bidirectional nearest neighbor as the loss function to achieve accurate fusion of the preoperative combined mesh model and the intraoperative tissue mesh model under a specified binocular endoscopic image frame.

[0175] It should be noted that, compared with the virtual registration datasets constructed by biomechanical models in the prior art, the embodiments of the present invention use real endoscopic images and medical test data to construct a dataset that is tailored to the characteristics of the flexible dynamic environment in vivo. The network trained on this dataset has higher registration accuracy.

[0176] Specifically, during the training phase of the registration module, a training set is generated based on real data, including:

[0177] Based on the specified feature point pairs between the binocular endoscopic image frames and the preoperative tissue mesh model, a non-rigid algorithm is used to register the preoperative and intraoperative tissue mesh models using these feature points. For any given feature point:

[0178]

[0179] Wherein, Non_rigid_ICP represents the non-rigid registration algorithm ICP. This represents the a-th feature point in the preoperative tissue mesh model used for non-rigid registration. correspond Feature points of the intraoperative tissue mesh model, T G T is the global transition matrix of the preoperative tissue mesh model. l,a It belongs to feature point v pre,a The local deformation transfer matrix;

[0180] The local deformation transfer matrix T of all vertices in the preoperative tissue mesh model was obtained by using quaternion interpolation. l Vertices v in the preoperative tissue mesh model are obtained by transforming the relationships. pre Registered coordinate labels

[0181] Accordingly, during the training phase of the registration module, the following supervised loss function needs to be constructed:

[0182]

[0183] Among them, Loss s This represents the supervised loss function during the training phase.

[0184] β s γ s These represent the coefficients of the supervised loss term;

[0185] N1 represents the preoperative combined mesh model M. pre The number of vertices;

[0186] This represents the L2 ground truth loss based on a manually labeled dataset. Represents the coordinates of all vertices after registration of the preoperative tissue mesh model;

[0187] I c +II c +III c I represents the Cauchygreen invariant, used to constrain the degree of tissue deformation in vivo. c The arc distance between two points on the constrained surface remains constant, II c The surface area of ​​the constrained tissue remains unchanged, III c The volume of the constrained tissue remains constant.

[0188] The first acquisition module is used to receive the region to be excised marked by the doctor on the region of interest of the specified binocular endoscopic image frame, and to obtain the corresponding vertex set of all pixels in the region to be excised in the overall tissue mesh model.

[0189] Define all pixels within the area to be excised in the intraoperative tissue mesh model M. in The vertex set on is P s All pixels within the area to be removed are in the overall tissue mesh model M. trans The corresponding vertex set in is P trans Where any vertex is p trans The preoperative tissue mesh model M pre The corresponding vertex set on is P ct Where any vertex is p ct ;

[0190] The first acquisition module is specifically used to obtain P s The nearest neighbor algorithm is used to obtain the nearest neighbor in M. trans P on trans Any vertex p in trans .

[0191] It is easy to understand that the second modeling unit of the aforementioned registration module obtains the vertex set P by calculating the spatial coordinates of the pixels in the camera coordinate system based on the depth values ​​of all pixels in the area to be removed and by using the pinhole camera model. s .

[0192] The second acquisition module is used to acquire and visualize the corresponding vertex set on the preoperative tissue mesh model based on the corresponding vertex set in the overall tissue mesh model.

[0193] The second acquisition module is used to obtain R based on the data acquired by the registration module. pre t pre , and p obtained by the first acquisition module trans Obtain and visualize M pre P on ct Any vertex p in ct ;

[0194] p ct =R pre p trans +t pre .

[0195] The solution module receives the cutting direction and manually selected candidate regions marked by the doctor on the preoperative tissue mesh model, combines them with the corresponding vertex set on the preoperative tissue mesh model, obtains the local volume of the organ corresponding to the region to be removed, and finally obtains the remaining organ volume or the percentage of the remaining organ volume.

[0196] Specifically, the solution module is used for:

[0197] like Figure 2 As shown, traverse P ct vertex p inct Along the cutting direction v pro Obtain the candidate region within the manually selected area and the cutting direction v. pro Find the nearest vertices, calculate the average of these vertices, and obtain the corresponding vertex p′. ct The coordinates of the points form the corresponding vertex set P′. ct ;

[0198] According to vertex set P ct and P′ ct Pairing spatial point p ct and p′ ct Multiple irregular pentahedrons are generated within the region S to be excised.

[0199] Calculate the volume of each irregular pentahedron and sum them to obtain the local volume V of the organ corresponding to the region to be removed. cut Combined with the overall volume V of the organ before surgery all Finally, the remaining organ volume V was obtained. remain =V all -V cut or percentage of remaining organ volume

[0200] Among them, such as Figure 3 As shown, the steps for calculating the volume of any pentahedron A1A2A3B1B2B3 are as follows:

[0201] Draw auxiliary lines B1′A2 and A2B3′ along the direction parallel to B1B2 and B2B3 to divide the irregular pentahedron into a triangular prism B1B2B3B1′A2B3′ and a square pyramid A1A3B3′B1′A2;

[0202] Calculate the area of ​​plane B1B2B3 and the perpendicular distance h between plane A2 and plane B1B2B3. prism Obtain the volume of the triangular prism B1B2B3B1′A2B3′:

[0203]

[0204] Calculate the area of ​​plane A1A3B3′B1′ and the perpendicular distance h between A2 and plane A1A3B3B1. pyramid Obtain the volume of the square pyramid A1A3B3′B1′A2:

[0205]

[0206] Then the volume V of each pentahedron final for:

[0207] V final =V prism +Vpyramid .

[0208] In addition to the factors mentioned above that may affect the fusion accuracy, how the second modeling unit obtains the depth value of the specified binocular endoscopic image frame is also a key factor, as it directly affects the accuracy of the intraoperative tissue mesh model.

[0209] As mentioned above, the second modeling unit uses an online self-supervised learning depth estimation method based on binocular endoscope to obtain the depth value of the specified binocular endoscope image frame; the binocular depth estimation network used by the online self-supervised learning depth estimation method has the ability to quickly overlearn and can continuously adapt to new scenarios using self-supervised information;

[0210] In real-time reconstruction mode, the second modeling unit is specifically used to overfit continuous video frames to obtain the depth value of a specified binocular endoscopic image frame, including:

[0211] Extraction subunits are used to acquire binocular endoscope images, and the encoder network of the current binocular depth estimation network is used to extract multi-scale features of the current frame image;

[0212] The fusion subunit is used to fuse multi-scale features using the decoder network of the current binocular depth estimation network to obtain the disparity of each pixel in the current frame image;

[0213] The conversion subunit is used to convert parallax into depth based on camera intrinsic and extrinsic parameters and output it as the result of the current frame image.

[0214] The first estimation subunit is used to update the parameters of the current stereo depth estimation network using self-supervised loss without introducing external ground truth, for depth estimation of the next frame image.

[0215] This depth estimation scheme utilizes the similarity of consecutive frames to extend the overfitting idea from a pair of binocular images to overfitting over time series. By continuously updating the model parameters through online learning, it can obtain high-precision tissue depth in various binocular endoscopic surgical environments.

[0216] The pre-training stage of the stereo depth estimation network abandons the traditional training mode and adopts the idea of ​​meta-learning, which allows the network to learn the depth of one image to predict the depth of another image, thereby calculating the loss and updating the network. This can effectively promote the network's generalization to new scenes and improve its robustness to low-texture complex lighting, while significantly reducing the time required for subsequent overfitting.

[0217] like Figure 4 As shown in section b, the initial model parameters corresponding to the stereo depth estimation network are obtained through meta-learning training, specifically including:

[0218] S100, Randomly select an even number of pairs of stereo images {e1,e2,…,e 2K} and equally divided into support sets and query set and Images are randomly paired to form K tasks

[0219] S200, Inner Circulation Training: Based on The loss is calculated from the support set image to perform a parameter update;

[0220]

[0221] in, This represents the network parameters after the inner loop update; Let α represent the derivative, where α is the learning rate of the inner loop. For the support set image of the k-th task, It is based on the initial parameters φ of the model m The calculated loss; f represents the stereo depth estimation network;

[0222] S300, External Loop Training: Based on The query set image is used to calculate the meta-learning loss using the updated model, and the initial parameters φ of the model are directly updated. m For φ m+1 ;

[0223]

[0224] Where β is the learning rate of the outer loop; This is the query set image for the k-th task. This is the learning loss of the meta-learning.

[0225] The following is a detailed description of each sub-unit included in the second modeling unit:

[0226] For extracting sub-units, such as Figure 4 As shown in section a, it acquires binocular endoscopic images and uses the encoder network of the current binocular depth estimation network to extract multi-scale features of the current frame image.

[0227] For example, the encoder of the stereo depth estimation network in this subunit uses a ResNet18 network to extract feature maps at five scales for the current frame image (left and right eyes) respectively.

[0228] For fused subunits, such as Figure 4As shown in section a, it employs the decoder network of the current binocular depth estimation network to fuse multi-scale features and obtain the disparity of each pixel in the current frame image; specifically, it includes:

[0229] The decoder network described above processes the coarse-scale feature map through convolutional blocks and upsampling, concatenates it with the fine-scale feature map, and then performs feature fusion through convolutional blocks again. The convolutional blocks are constructed by combining reflection padding, convolutional layers, and nonlinear activation subunits (ELUs).

[0230] Calculate the disparity directly based on the output with the highest network resolution:

[0231] d = k·(sigmoid(conv(Y))-TH)

[0232] Where d represents the disparity estimate of a pixel; k is the preset maximum disparity range; Y is the highest resolution output; TH represents a parameter related to the type of binocular endoscope, which is 0.5 when the endoscope image has negative disparity and 0 when all endoscope images have positive disparity; conv is a convolutional layer; and sigmoid performs range normalization.

[0233] For the transformation subunit, it converts disparity into depth based on camera intrinsic and extrinsic parameters and outputs it as the result of the current frame image.

[0234] In this sub-unit, converting parallax to depth means:

[0235]

[0236] Among them, c x1 , These are the intrinsic parameter matrices of the left and right eye endoscopes and cameras in a binocular endoscope. The corresponding parameter in; if f x Take the corresponding internal parameters of the left eye camera When f is the left-eye pixel, then d takes the disparity estimate of the left-eye pixel, and D is the depth estimate of the left-eye pixel; if f x Take the corresponding internal parameters of the right eye camera Then d is the disparity estimate of the right eye pixel, and D is the depth estimate of the right eye pixel; b is the baseline length, i.e. the extrinsic parameter of the binocular camera.

[0237] For the first estimation unit, such as Figure 4 As shown in section b, it uses self-supervised loss to update the parameters of the current stereo depth estimation network without introducing external ground truth, for depth estimation of the next frame image.

[0238] It is easy to understand that the "external truth value" mentioned in the embodiments of the present invention is the label (or "supervision information"), which is a well-known expression in the art.

[0239] In this sub-unit, such as Figure 4 As shown in part b, the self-supervised loss is expressed as:

[0240]

[0241] Among them, L self The value represents the self-supervised loss; α1, α2, α3, and α4 are all hyperparameters, l corresponds to the left figure, and r corresponds to the right figure.

[0242] Since both eyes are observing the same scene, the values ​​of corresponding pixels on the left and right depth maps should be equal when transformed to the same coordinate system. Therefore, we introduce... and

[0243] (1) The geometric consistency loss is represented by the left figure:

[0244]

[0245] Wherein, P1 represents the first set of valid pixels (i.e., the valid pixels of the right eye); The effective pixel p represents the left-eye depth obtained from the right-eye depth map after camera pose transformation, and D represents the left-eye depth. l ′(p) represents the effective pixel point p using the predicted right-side disparity Dis. R The left eye depth is obtained by sampling on the left eye depth map.

[0246] (2) The geometric consistency loss is shown in the right figure:

[0247]

[0248] Wherein, P2 represents the second set of valid pixels (i.e., the valid pixels of the left eye); The effective pixel p represents the right-eye depth obtained from the left-eye depth map after camera pose transformation, and D represents the right-eye depth. r ′(p) represents the effective pixel point p using the predicted left image disparity Dis. L The right eye depth is obtained by sampling on the right eye depth map.

[0249] By incorporating geometric consistency constraints into the training loss, the network's general applicability to hardware is ensured, enabling it to autonomously adapt to unconventional binocular images such as surgical endoscopes.

[0250] Assuming constant brightness and spatial smoothness during endoscopic surgery, reprojection between left and right eye images can achieve reconstruction of another objective. However, this introduces structural similarity loss. The brightness, contrast, and structure of the two images are normalized and compared, and then... and

[0251] (3) The left image shows the luminous loss:

[0252]

[0253] Among them, I L (p) represents the left figure, I L ′(p) represents the parallax between the right image and the predicted left image Dis. L (p) Reconstructed image from the left eye endoscope, λ i and λ s To balance the parameters, SSIM LL′ (p) represents I L (p) and I L Image structural similarity of ′(p);

[0254] (4) The image on the right shows the luminous loss:

[0255]

[0256] Among them, I R (p) represents the right figure, I′ R (p) indicates the use of the disparity between the left image and the predicted right image. R (p) Generated reconstructed image from the right eye endoscope, SSIM RR′ (p) represents I R (p) and I′ R Image structural similarity (p).

[0257] In low-texture and monochromatic organizational regions, smoothing priors are used to aid inference, and depth regularization is applied, introducing... and

[0258] (5) The smoothing loss is shown in the left figure:

[0259]

[0260] in, This represents the normalized left eye depth map. and This represents the first derivative along the horizontal and vertical directions of the image;

[0261] (6) The smoothing loss is shown in the right figure:

[0262]

[0263] in, This represents the normalized depth map of the right eye. and This represents the first derivative along the horizontal and vertical directions of the image.

[0264] Specifically, the process of obtaining the first set of valid pixels P1 and the second set of valid pixels P2 is as follows:

[0265] Define the left eye disparity predicted by the current binocular depth estimation network as: Right eye parallax is The formulaic expression for the cross-validation mask for the left and right eyes is as follows:

[0266]

[0267]

[0268] in, These are used to determine whether the pixel at position (i,j) in the left and right eye images is within the stereo matching range; i takes the value of any integer between [1,W]; j takes the value of any integer between [1,H]; W represents the image width, and H represents the image height;

[0269] Let c take the value L or R, when If the value is within the range of stereo matching under the current calculation method, then the pixel at position (i,j) is within the range of stereo matching; otherwise, it is not within the range of stereo matching.

[0270] By using a pinhole camera model, binocular pose transformation, and predicted depth for projection, an effective region mask based on 3D points is obtained. Take 0 or 1, when If the value is within the range of stereo matching under the current calculation method, then the pixel at position (i,j) is within the range of stereo matching; otherwise, it is not within the range of stereo matching.

[0271] Obtain the final valid region mask

[0272]

[0273] If pixel p satisfies When c is R, the first set of valid pixels P1 is obtained; when c is L, the second set of valid pixels P2 is obtained.

[0274] In the corrected stereo image, additional regions caused by viewpoint shift cannot find matching pixels. However, this embodiment of the invention considers that low texture and uneven illumination of in vivo tissues can make local features less obvious, and pixels in these invalid regions often find similar pixels in neighboring regions. Therefore, as mentioned above, this embodiment of the invention proposes a cross-validation-based binocular effective region recognition algorithm, which eliminates the misleading effect of self-supervised loss of invalid region pixels on network learning and improves the accuracy of depth estimation.

[0275] In addition, to avoid insufficient robustness of depth estimation in pure texture or low-light scenes, a method is also introduced.

[0276] (7) Indicates the loss of sparse optical flow:

[0277]

[0278] Among them, Dis L (p) represents the predicted left eye disparity map, OF L (p) represents the sparse disparity map of the left eye. r (p) represents the predicted right eye disparity map, OF R (p) represents the right eye sparse disparity map; P3 represents the left eye sparse disparity map OF. L The third set of valid pixels in (p); P4 represents the right eye sparse disparity map OF. R The fourth set of valid pixels in (p); γ1 and γ2 are balancing parameters, both of which are non-negative and not both of which are 0 at the same time.

[0279] Specifically, the process of obtaining the third set of valid pixels P3 and the fourth set of valid pixels P4 is as follows:

[0280] Using the LK (Lucas-Kanade) optical flow solution algorithm, sparse optical flow (Δx, Δy) is calculated every n pixels in the row and column directions, where Δx represents the horizontal offset of the pixel and Δy represents the vertical offset of the pixel.

[0281] When solving for the optical flow from the left figure to the right figure, only when If Δx > thred1, the disparity at that pixel location is retained as Δx, where KT and thred1 are the corresponding preset thresholds. If the above conditions are not met or the disparity at the sparse optical flow location is not calculated, it is set to 0 to obtain the final sparse disparity map OF. L (p), OF L Pixels whose (p)≠0 constitute the third set of valid pixels P3;

[0282] When solving for the optical flow from the right figure to the left figure, only when And Δx < thred2, the disparity at this pixel position is retained as Δx, where thred2 is the corresponding preset threshold. For those that do not meet the above conditions or for which the sparse optical flow position disparity is not calculated, it is set to 0 to obtain the final sparse disparity map OF. R (p), OF R Pixels where (p)≠0 form the fourth set of valid pixels P4.

[0283] As mentioned above, in the embodiments of the present invention, traditional Lucas-Kanade optical flow is introduced to deduce the sparse disparity between binocular images, giving the network a reasonable learning direction, improving the fast learning ability and reducing the probability of falling into local optimum.

[0284] It should be particularly emphasized that, in addition to the real-time reconstruction mode, the online self-supervised learning depth estimation method adopted by the second modeling unit in the embodiments of the present invention also sets a precise measurement mode. As Figure 4 shown in part b, in the precise measurement mode, the second modeling unit is specifically used for overfitting the key image video frames, including:

[0285] The second estimation sub-unit, without introducing external ground truth, according to the binocular depth estimation network obtained from the previous frame image of the specified binocular endoscope image frame in the real-time reconstruction mode, uses the self-supervised loss corresponding to the specified binocular endoscope image frame to update the parameters of the aforementioned binocular depth estimation network until convergence, and uses the converged binocular depth estimation network for precise depth estimation of the specified binocular endoscope image frame to obtain the depth value of the specified binocular endoscope image frame.

[0286] It should be noted that the technical details such as the depth estimation network, self-supervised loss function, effective region mask calculation, and meta-learning pre-training method in the precise measurement mode are all consistent with the technical details extended in the real-time reconstruction mode, and will not be elaborated here.

[0287] In summary, compared with the prior art, the present invention has the following beneficial effects:

[0288] 1. The embodiments of the present invention do not rely on additional devices. By the matching relationship between the preoperative tissue mesh model and the intraoperative tissue mesh model, the measurement of the remaining organ volume is transformed into the volume measurement of the corresponding area of the preoperative tissue mesh model, avoiding the interference of tissue complex deformation and invisible areas on volume prediction.

[0289] 2. The embodiments of the present invention generate training data by artificial annotation and interpolation based on real data, train the multi-modal registration and fusion network in a supervised manner, and finally further improve the registration accuracy through unsupervised fine-tuning.

[0290] 3. The embodiments of the present invention take into account the interaction between the surgeon and the patient during the operation, receive simple annotations from the surgeon for the area of ​​interest, obtain accurate volume measurement information, and achieve active selectivity and high reference value during the operation.

[0291] 4. This invention provides an online self-supervised learning depth estimation method based on binocular endoscopy, the beneficial effects of which include at least:

[0292] 4.1 The switching depth estimation can provide real-time point cloud of intraoperative anatomical structures to help doctors intuitively understand the intraoperative three-dimensional structure. It can also achieve high-precision reconstruction of key frames selected by doctors based on single-frame overfitting, providing a basis for subsequent measurements, so that speed and accuracy can be balanced in application.

[0293] 4.2 By utilizing the similarity of consecutive frames, the overfitting concept on a pair of binocular images is extended to overfitting on time series. Through online learning, the model parameters are continuously updated, enabling high-precision tissue depth measurements to be obtained in various binocular endoscopic surgical environments.

[0294] 4.4 The pre-training stage of the network model abandons the traditional training mode and adopts the idea of ​​meta-learning, which allows the network to learn the depth of one image to predict the depth of another image, thereby calculating the loss and updating the network. This can effectively promote the generalization of the network to new scenes and improve its robustness to low-texture complex lighting, while significantly reducing the time required for subsequent overfitting.

[0295] 4.4 By incorporating geometric consistency constraints into the training loss, the network's general applicability to hardware is ensured, enabling autonomous adaptation to unconventional binocular images such as surgical endoscopes.

[0296] 4.5. Depth estimation of each frame of stereo image is treated as an independent task, and high-precision models suitable for the current frame are obtained through real-time overfitting; and new scenes can be learned quickly through online learning to obtain high-precision depth estimation results.

[0297] 4.6 The cross-validation binocular effective region recognition algorithm eliminates the misleading effect of the self-supervised loss of invalid region pixels on network learning and improves the accuracy of depth estimation.

[0298] 4.7. Introducing the traditional Lucas-Kanade optical flow to derive the sparse parallax between binocular images provides the network with a reasonable learning direction, improves its rapid learning ability, and reduces the probability of getting trapped in local optima.

[0299] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0300] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A system for estimating the volume of residual organs during surgery to aid surgical planning, characterized in that, include: The registration module is used to register the preoperative tissue mesh model and the intraoperative tissue mesh model to obtain an overall tissue mesh model that displays the internal tissue information of the preoperative tissue mesh model in the intraoperative tissue mesh model. Specifically, the intraoperative tissue mesh model is obtained based on the depth value of the specified binocular endoscopic image frame; The first acquisition module is used to receive the region to be excised marked by the doctor on the region of interest of the specified binocular endoscope image frame, and to acquire the corresponding vertex set of all pixels in the region to be excised in the overall tissue mesh model. The second acquisition module is used to acquire and visualize the corresponding vertex set on the preoperative tissue mesh model based on the corresponding vertex set in the overall tissue mesh model. The solution module is used to receive the cutting direction and manually selected candidate regions marked by the doctor on the preoperative tissue mesh model, and combine them with the corresponding vertex set on the preoperative tissue mesh model to obtain the local volume of the organ corresponding to the region to be removed, and finally obtain the remaining organ volume or the percentage of the remaining organ volume. The solution module is specifically used for: Traversing P ct vertex p in cy Along the cutting direction v pro Obtain the candidate region within the manually selected area and the cutting direction v. pro Find the nearest vertices, calculate the average of these vertices, and obtain the corresponding vertex p′. ct The coordinates of the points form the corresponding vertex set P′. ct ; According to vertex set P ct and P′ ct Pairing spatial point p ct and p′ ct Multiple irregular pentahedrons are generated within the region S to be excised. Calculate the volume of each irregular pentahedron and sum them to obtain the local volume V of the organ corresponding to the region to be removed. cut Combined with the overall volume V of the organ before surgery all Finally, the remaining organ volume V was obtained. remain =V all -V cut or percentage of remaining organ volume The steps for calculating the volume of any pentahedron A1A2A3B1B2B3 are as follows: Draw auxiliary lines B′1A2 and A2B′3 along the direction parallel to B1B2 and B2B3 to divide the irregular pentahedron into a triangular prism B1B2B3B′1A2B′3 and a square pyramid A1A3B′3B′1A2; Calculate the area of ​​plane B1B2B3 and the perpendicular distance h between plane A2 and plane B1B2B3. prism Obtain the volume of the triangular prism B1B2B3B′1A2B′3: Calculate the area of ​​plane A1A3B′3B′1 and the perpendicular distance h between plane B2 and plane B1B3B3B1. pyramid Obtain the volume of the square pyramid A1A3B′3B′1A2: Then the volume V of each pentahedron final for: V final =V prism +V pyramid 。 2. The intraoperative residual organ volume estimation system as described in claim 1, characterized in that, The registration module includes: The first modeling unit is used to obtain a preoperative tissue mesh model with tissue semantic information; The second modeling unit is used to obtain an intraoperative tissue mesh model based on the depth value of a specified binocular endoscopic image frame. The feature extraction unit is used to obtain corresponding multi-level features based on the preoperative tissue mesh model and the intraoperative tissue mesh model, respectively. The overlap prediction unit is used to obtain the overlapping area of ​​the preoperative tissue mesh model and the intraoperative tissue mesh model based on the multi-level features, and to obtain the pose transformation relationship of the vertices of the preoperative tissue mesh model within the overlapping area. The global fusion unit is used to obtain the registered vertex coordinates of the preoperative tissue mesh model based on the coordinates and pose transformation relationship of the vertices in the overlapping region and the coordinates of the vertices in the non-overlapping region. The information display unit is used to obtain the overall tissue grid model that displays the internal tissue information of the preoperative tissue grid model in the intraoperative tissue grid model based on the coordinates of all vertices after registration of the preoperative tissue grid model.

3. The intraoperative residual organ volume estimation system as described in claim 2, characterized in that, The feature extraction unit uses Chebyshev spectral map convolution to extract multi-level features from the preoperative and intraoperative tissue mesh models: Among them, the preoperative tissue mesh model M is defined. pre =(V pre E pre V pre E represents the three-dimensional coordinates of the vertices of the preoperative tissue mesh model. pre The edges between vertices in the preoperative tissue mesh model; the intraoperative tissue mesh model M in =(V in E in V in E represents the three-dimensional coordinates of the vertices of the preoperative tissue mesh model. in Represents the edges between vertices of the intraoperative tissue mesh model; and Let the downsampling scale features of the (n+1)th and nth layers of the preoperative tissue model be represented respectively, and initialized. For V ore ; and Let the features of the (n+1)th and nth layers of the intraoperative tissue model be represented respectively, and initialized. For V in ; The b-th order Chebyshev polynomials calculated from their respective vertices and their B-ring neighborhoods. They are respectively from edge E in E pre Calculate the scaled Laplacian matrix. These are the learning parameters of the neural network; And / or the overlap prediction unit is specifically used for: An attention mechanism is used to obtain the overlapping region of the preoperative tissue mesh model and the intraoperative tissue mesh model, including: Among them, O pre Represents the preoperative tissue mesh model M pre Mask of overlapping regions; O in Intraoperative tissue mesh model M in The mask for the overlapping region; cross and self represent the self-attention and cross-attention operations, respectively; and These represent the m-th downsampling scale features of the vertices of the preoperative tissue mesh model and the intraoperative tissue mesh model, respectively. According to mask O pre and O in Get the vertices that are within the overlapping region. and its characteristics The preoperative tissue mesh model M was calculated using a multilayer perceptron (MLP). pre Vertex in Corresponding point: in, It is the intraoperative tissue mesh model M in The vertices in the model correspond to the preoperative tissue mesh model M. pre Vertex in This indicates the calculation of cosine similarity. This indicates that the position encoding operation is performed on the vertices of the intraoperative tissue mesh model that are within the overlapping area; Vertex construction using KNN (Knowledge Neighbors) For the local neighborhood, the rotation matrix is ​​solved using Singular Value Decomposition (SVD), as shown in the following formula: in, Represents vertices The rotation matrix; This indicates that the KNN algorithm is used to construct the vertex... A local neighborhood; It is the vertex of the preoperative tissue mesh model. neighborhood points, It corresponds to the neighborhood point The vertices of the intraoperative tissue mesh model; Using rotation matrix Change the point cloud coordinates to obtain Using MLP to predict vertices The displacement vector is given by the following formula: in, The displacement vectors of the vertices in the overlapping region of the preoperative tissue mesh model are compared with the rotation matrix. This constitutes the pose transformation relationship; And / or the global fusion unit is specifically used for: The rotation matrix and displacement vector of all vertices of the preoperative tissue mesh model were obtained by MLP regression: Among them, R pre ,t pre These represent the rotation matrix and displacement vector of all vertices in the preoperative tissue mesh model, respectively; Indicates based on vertices within the overlapping region All vertices v of the preoperative tissue mesh model pre The weights for distance calculation; in, This represents the coordinates of all vertices after registration of the preoperative tissue mesh model.

4. The intraoperative residual organ volume estimation system as described in claim 1, characterized in that, During the training phase of the intraoperative residual organ volume estimation system, a training set is generated based on real data: Based on the specified feature point pairs between the binocular endoscopic image frames and the preoperative tissue mesh model, a non-rigid algorithm is used to register the preoperative and intraoperative tissue mesh models using these feature points. For any given feature point: Wherein, Non_rigid_ICP represents the non-rigid registration algorithm ICP. This represents the a-th feature point in the preoperative tissue mesh model used for non-rigid registration. correspond Feature points of the intraoperative tissue mesh model, T G T represents the global transition matrix of the preoperative tissue mesh model. l,a It belongs to feature point v pre,a The local deformation transfer matrix; The local deformation transfer matrix T of all vertices in the preoperative tissue mesh model was obtained by using quaternion interpolation. l Vertices v in the preoperative tissue mesh model are obtained by transforming the relationships. pre Registered coordinate labels 5. The intraoperative residual organ volume estimation system as described in claim 4, characterized in that, During the training phase of the intraoperative residual organ volume estimation system, the following supervised loss function is constructed: Among them, Loss s This represents the supervised loss function during the training phase. β s γ s These represent the coefficients of the supervised loss term; N1 represents the preoperative combined mesh model M. pre The number of vertices; This represents the L2 ground truth loss based on a manually labeled dataset. Represents the coordinates of all vertices after registration of the preoperative tissue mesh model; I c +II c +III c I represents the Cauchygreen invariant, used to constrain the degree of tissue deformation in vivo. c The arc distance between two points on the constrained surface remains constant, II c The surface area of ​​the constrained tissue remains unchanged, III c The volume of the constrained tissue remains constant.

6. The intraoperative residual organ volume estimation system as described in claim 2, characterized in that, The registration module also includes: The precision fine-tuning unit is used to introduce an unsupervised loss fine-tuning network to assist the global fusion unit in obtaining the registered vertex coordinates of the preoperative tissue mesh model. And / or the unsupervised loss fine-tuning network described above, during application, constructs the following unsupervised loss function: Among them, Loss u Represents the unsupervised loss function; β u ,γ u They represent the coefficients of the unsupervised loss term, and These are the vertex coordinates after registration of the preoperative tissue mesh model during unsupervised training. This represents the vertices of the preoperative tissue mesh model after distance registration in the intraoperative tissue mesh model. The closest point, Represents vertices and European distance, This represents the distance v from the vertex of the intraoperative tissue mesh model in the preoperative tissue mesh model after registration. in,b The closest point, Represents vertex v in,b and vertex Euclidean distance; N1 represents the preoperative tissue mesh model M. pre The number of vertices, N2 represents the intraoperative tissue mesh model M. in The number of vertices; Describes the Cauchygreen invariant during unsupervised training. The arc distance between two points on the constrained surface remains constant. The surface area of ​​the constrained tissue remains unchanged. The volume of the constrained tissue remains constant.

7. The intraoperative residual organ volume estimation system as described in claim 3, characterized in that, Define all pixels within the area to be excised in the intraoperative tissue mesh model M. in The vertex set on is P s All pixels within the area to be removed are in the overall tissue mesh model M. trans The corresponding vertex set in is P trans Where any vertex is p trans The preoperative tissue mesh model M pre The corresponding vertex set on is P ct Where any vertex is p ct ; The first acquisition module is specifically used to obtain P s The nearest neighbor algorithm is used to obtain the nearest neighbor in M. trans P on trans Any vertex p in trans ; And / or the second acquisition module is used to obtain R pre t pre and p trans Obtain and visualize M pre P on ct Any vertex p in ct ; p ct =R pre p trans +t pre 。 8. The intraoperative residual organ volume estimation system as described in claim 2, characterized in that, The second modeling unit uses an online self-supervised learning depth estimation method based on binocular endoscope to obtain the depth value of the specified binocular endoscope image frame; the binocular depth estimation network used by the online self-supervised learning depth estimation method has the ability to quickly overlearn and can continuously adapt to new scenes using self-supervised information; In real-time reconstruction mode, the second modeling unit is specifically used to overfit continuous video frames to obtain the depth value of a specified binocular endoscopic image frame, including: Extraction subunits are used to acquire binocular endoscope images, and the encoder network of the current binocular depth estimation network is used to extract multi-scale features of the current frame image; The fusion subunit is used to fuse multi-scale features using the decoder network of the current binocular depth estimation network to obtain the disparity of each pixel in the current frame image; The conversion subunit is used to convert parallax into depth based on camera intrinsic and extrinsic parameters and output it as the result of the current frame image. The first estimation subunit is used to update the parameters of the current stereo depth estimation network using self-supervised loss without introducing external ground truth, for depth estimation of the next frame image.

9. The intraoperative residual organ volume estimation system as described in claim 8, characterized in that, In precise measurement mode, the second modeling unit is specifically used to overfit key image video frames, including: The second estimation subunit, without introducing external truth values, uses the binocular depth estimation network obtained in real-time reconstruction mode based on the previous frame of the specified binocular endoscopic image frame. It then updates the parameters of the aforementioned binocular depth estimation network using the self-supervised loss corresponding to the specified binocular endoscopic image frame until convergence. The converged binocular depth estimation network is then used to accurately estimate the depth of the specified binocular endoscopic image frame, thereby obtaining the depth value of the specified binocular endoscopic image frame.

Citation Information

Patent Citations

  • Online self-supervised learning depth estimation method based on binocular endoscope

    CN115359104A

  • Surgical systems for generating three dimensional constructs of anatomical organs and coupling identified anatomical structures thereto

    US20210196385A1