Operation target pose tracking method and device, electronic equipment and storage medium
By optimizing and calculating the error of the initial posture of the surgical target in the current frame and using the posture optimization objective function, the real-time and accuracy problems of surgical target posture tracking in the existing technology are solved, and efficient posture tracking is achieved, which is suitable for complex surgical environments.
Patent Information
- Application Number
- CN202510636301.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-10-17
Smart Images

Figure CN120807575A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and in particular to a surgical target pose tracking method and device, electronic equipment and a storage medium. BACKGROUND
[0002] With the rapid development of intelligent surgery and robot-assisted surgery systems, higher requirements are put forward for the spatial positioning and six degrees of freedom (6DoF) pose tracking capability of key instruments and anatomical structures in the surgical process. In particular, in the fields of neurosurgery, orthopedics, otolaryngology and other fields involving fine operations, how to obtain and track the spatial pose of the surgical target such as the intraoperative instrument or anatomical site in real time so that accurate registration and slice navigation can be achieved with preoperative CT / MRI medical images has become a key technical bottleneck for improving the safety, accuracy and automation level of surgery.
[0003] Currently, the main methods for obtaining and tracking the pose of the surgical target in real time include an optical positioning system based on an infrared reflective ball (such as NDI Polaris, Vicon, etc.) and a sparse point cloud reconstruction, simultaneous localization and mapping (SLAM) method based on vision.
[0004] However, although the optical positioning system based on the infrared reflective ball has high static accuracy, the sparse point cloud reconstruction, SLAM method based on vision and other methods can replace high-cost optical systems, but both have the problem of being difficult to support continuous pose updates in the intraoperative dynamic changes and intense interactions of multiple instruments. SUMMARY
[0005] The present application provides a surgical target pose tracking method, device, electronic equipment and storage medium to solve the defect that it is difficult to support continuous pose updates in the intraoperative dynamic changes and intense interactions of multiple instruments in the prior art, and to realize a high-precision and real-time pose tracking scheme.
[0006] The present application provides a surgical target pose tracking method, comprising: optimizing a current frame initial pose of a surgical target to be tracked at a current time to obtain a current frame estimated pose; determining a current frame virtual image of the surgical target to be tracked based on the current frame estimated pose; calculating a first error value based on the current frame virtual image and a current frame observation image; In a case that the first error value is greater than a preset error threshold, a plurality of first pose candidates of the to-be-tracked surgical target are obtained based on the current frame observation image; a second error value of the first pose candidate is calculated based on a preset pose optimization objective function, and the first pose candidate is continuously optimized with the second error value being minimum as an optimization objective, to obtain a plurality of second pose candidates; and a target estimated pose is determined from the second pose candidates; In a case that the first error value is less than or equal to the preset error threshold, the target estimated pose is determined based on the current frame estimated pose. The pose of the to-be-tracked surgical target is tracked based on the target estimated pose, and a current frame initial pose at a next time is determined based on the target estimated pose.
[0007] According to the surgical target pose tracking method provided by the application, the current frame observation image includes a current frame observation color image and a current frame observation depth image; the second error value of the first pose candidate is determined based on a first re-projection error of a color channel and a second re-projection error of a depth channel; the first re-projection error is determined based on the current frame observation color image and a first virtual color image; the second re-projection error is determined based on the current frame observation depth image and a first virtual depth image; and the first virtual color image and the first virtual depth image are both determined based on the first pose candidate.
[0008] According to the surgical target pose tracking method provided by the application, the expression of the pose optimization objective function is as follows: ; Wherein, is the second error value of the first pose candidate; is a rotation parameter of the first pose candidate; is an offset parameter of the first pose candidate; is the current frame observation color image; is a first virtual color image determined based on the first pose candidate; is the current frame observation depth image; is a first virtual depth image determined based on the first pose candidate; is a depth error weighting factor; is an effective pixel set of a mask region of the to-be-tracked surgical target; is a mask region pixel in the effective pixel set .
[0009] The application provides a surgical target pose tracking method. The current frame observation color image is input into a pre-trained mask region extraction model to obtain a mask region of the surgical target to be tracked output by the mask region extraction model; target depth pixels are extracted from a current frame observation depth image based on the mask region; a local point cloud region is generated based on the target depth pixels; a plurality of initial pose candidates are generated based on the current frame observation color image, the current frame observation depth image, a three-dimensional model of the surgical target to be tracked and a geometric center point of the local point cloud region; and the plurality of first pose candidates are determined based on the plurality of initial pose candidates.
[0010] The application provides a surgical target pose tracking method. The initial pose candidate is superimposed with a preset amplitude of translation disturbance to generate the first pose candidate; and the number of the first pose candidates is greater than the number of the initial pose candidates.
[0011] The application provides a surgical target pose tracking method. A second virtual depth image, a second virtual color image and a virtual mask region of each second pose candidate are determined; a first score is determined based on the mean absolute error between the second virtual depth image and the current frame observation depth image; a second score is determined based on the degree of overlap between the virtual mask region and the mask region; a third score is determined based on the edge map error term between the second virtual color image and the current frame observation color image; a matching quality score of each second pose candidate is determined based on the first score, the second score and the third score; and the target estimated pose is determined based on the second pose candidate with the highest matching quality score.
[0012] The application provides a surgical target pose tracking method. The first error value is calculated based on the mean square error between the mask region pixels of the current frame virtual depth image and the mask region pixels of the current frame observation depth image.
[0013] The application provides a surgical target pose tracking method. The third error value of the initial posture of the current frame is calculated based on the posture optimization objective function, and the initial posture of the current frame is continuously optimized with the minimum third error value as the optimization goal to obtain the estimated posture of the current frame.
[0014] According to a surgical target posture tracking method provided by the present invention, the method of determining a current frame virtual image of the surgical target to be tracked based on the current frame posture estimation includes: Based on the estimated posture of the current frame, the initial three-dimensional points of the three-dimensional model of the surgical target to be tracked are transformed to obtain transformed three-dimensional points; based on the internal parameter matrix of the shooting device of the current frame observation image, the transformed three-dimensional points are projected to obtain plane pixel coordinates; based on the plane pixel coordinates, the depth value of the three-dimensional model and the texture information of the three-dimensional model, the current frame virtual color image and the current frame virtual depth image are synthesized.
[0015] The present invention also provides a surgical target posture tracking device, comprising: A posture estimation module is used to optimize the initial posture of the current frame of the surgical target to be tracked at the current moment to obtain the estimated posture of the current frame; A virtual image generation module, configured to determine a current frame virtual image of the surgical target to be tracked based on the estimated posture of the current frame; an error calculation module, configured to calculate a first error value based on the current frame virtual image and the current frame observation image; a first posture determination module configured to, when the first error value is greater than a preset error threshold, obtain a plurality of first posture candidates of the surgical target to be tracked based on the current frame observation image; calculate a second error value of the first posture candidate based on a preset posture optimization objective function, continuously optimize the first posture candidate with the second error value minimized as an optimization objective, to obtain a plurality of second posture candidates; and determine an estimated target posture from the second posture candidates; A second posture determination module is configured to determine the target estimated posture based on the current frame estimated posture when the first error value is less than or equal to the preset error threshold; The posture tracking module is used to track the posture of the surgical target to be tracked based on the estimated posture of the target, and determine the initial posture of the current frame at the next moment based on the estimated posture of the target.
[0016] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, it implements any of the above-described surgical target posture tracking methods.
[0017] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method for tracking the surgical target posture as described above is implemented.
[0018] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned surgical target posture tracking methods.
[0019] The surgical target posture tracking method, device, electronic device and storage medium provided by the present invention, by designing an error trigger mechanism in the surgical target posture tracking at continuous moments, defaults to using the estimated posture used for posture tracking at the previous moment as the initial posture at the current moment, and defaults to prioritizing only the current frame initial posture at the current moment to quickly optimize the current frame estimated posture, and judges whether the tracking fails based on the error between the current frame virtual image corresponding to the current frame estimated posture and the actually collected current frame observation image. It can quickly respond to complex clinical environments such as drastic dynamic changes of surgical targets and frequent interactions of multiple surgical instruments, adjust posture tracking in real time, and ensure the effectiveness of detection and tracking; and directly adjust the posture tracking when the tracking is successful based on whether the tracking fails. The estimated pose of the target is determined by using the estimated pose of the current frame. The pose initialization and pose optimization process is triggered only when tracking fails. The pose initialization operation is performed according to the current frame observation image collected in real time, and the pose candidates are continuously optimized with the goal of minimizing the error calculated by the preset pose optimization objective function. The target estimated pose is then confirmed from the optimized pose candidates. This can not only restore the final stability, but also eliminate the need to generate a large number of candidate poses at each moment, significantly reducing the computational overhead and computational delay, and improving the overall processing efficiency. Therefore, without the need for physical markers and without relying on expensive external positioning equipment, it achieves strong adaptability, high precision, strong robustness, and strong real-time overall pose tracking for arbitrary rigid bodies in complex surgical scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 It is a flow chart of the surgical target posture tracking method provided by the present invention.
[0022] Figure 2 It is a schematic diagram of the process of generating the first candidate posture provided by the present invention.
[0023] Figure 3is a structural schematic view of a surgical target pose tracking device provided by the present application.
[0024] Figure 4 is a structural schematic view of an electronic device provided by the present application. DETAILED DESCRIPTION
[0025] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present application.
[0026] It should be noted that, in the description of the present application, the term “comprising” or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The specific meaning of the above-mentioned term in the present application can be understood by those skilled in the art according to specific circumstances.
[0027] The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by “first”, “second” and the like are generally a class, and are not limited to the number of objects, for example, the first object can be one or more.
[0028] The present application provides a surgical target pose tracking method, device, electronic device and storage medium. Figures 1-4 The present application provides a surgical target pose tracking method, device, electronic device and storage medium.
[0029] At present, the main methods for real-time acquisition and tracking of the pose of the surgical target in the operation include optical positioning systems based on infrared reflecting balls (such as NDI Polaris, Vicon, etc.) and sparse point cloud reconstruction, SLAM methods based on vision.
[0030] Optical positioning systems based on infrared reflective spheres have high static accuracy, but face the following bottlenecks: (1) The hardware of optical positioning systems is expensive, and the installation and maintenance are complex, which has a high dependence on the operating room space and layout; (2) It is easily affected by occlusion and strong light interference, resulting in unstable pose tracking; (3) It generally needs to customize and modify surgical instruments or add external markers, which increases the preoperative preparation burden and affects the universality of surgical instruments; (4) It is difficult to adapt to the complex clinical environment with severe intraoperative dynamic changes and frequent multi-instrument interactions.
[0031] Introducing visual-based sparse point cloud reconstruction and SLAM methods into medical scenarios can replace high-cost optical systems, but when dealing with actual problems such as intraoperative strong occlusion, high dynamic changes, and multi-target interference, the tracking method is not adjusted for the different poses of each frame of surgical target, which has many limitations such as high redundant calculation, poor real-time performance, insufficient tracking accuracy, high calculation delay, and poor registration robustness, especially in supporting continuous intraoperative pose updating and efficient preoperative image slice display.
[0032] Therefore, it is necessary to provide a surgical target pose tracking method, device, electronic equipment and storage medium to solve at least one of the technical problems.
[0033] Figure 1 is a flowchart of the surgical target pose tracking method provided by the present application, as Figure 1 shown, the surgical target pose tracking method includes but is not limited to steps 110 to 150.
[0034] It should be noted that the execution subject of the surgical target pose tracking method provided by the present application is the corresponding surgical target pose tracking device, which can be a server, a computer device, such as a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a wearable device, an Ultra-Mobile Personal Computer (UMPC), a netbook or a Personal Digital Assistant (PDA), etc.
[0035] Step 110: optimizing the current frame initial pose of the to-be-tracked surgical target at the current time to obtain a current frame estimated pose.
[0036] Among them, the to-be-tracked surgical target includes but is not limited to at least one of various surgical instruments, various anatomical positions and other intraoperative targets.
[0037] The current frame initial pose and the current frame estimated pose both include six degrees of freedom (6DoF) pose information of the to-be-tracked surgical target.
[0038] Optionally, the initial pose of the current frame of the to-be-tracked surgical target at the current time is the target estimated pose of the to-be-tracked surgical target at the last time; or the initial pose of the current frame of the to-be-tracked surgical target at the current time is determined based on the target estimated pose of the to-be-tracked surgical target at the last time.
[0039] Specifically, at the current time, any one of a filtering method, a nonlinear least square method, a geometric optimization method, or the like is used to optimize the initial pose of the current frame of the to-be-tracked surgical target, to obtain the current frame estimated pose of the to-be-tracked surgical target.
[0040] Step 120: determining the current frame virtual image of the to-be-tracked surgical target based on the current frame estimated pose.
[0041] The current frame virtual image is a two-dimensional image obtained by mapping the three-dimensional current frame estimated pose of the to-be-tracked surgical target to a plane.
[0042] Specifically, at the current time, the three-dimensional current frame estimated pose is mapped to an image plane by coordinate system transformation, geometric projection, and visualization technology, and the obtained two-dimensional image is taken as the current frame virtual image. By comparing the current frame virtual image with the current frame observation image actually collected by the shooting device, it can be determined whether the optimized current frame estimated pose is suitable for being applied to pose tracking.
[0043] Optionally, the current frame virtual image includes a current frame virtual color image and a current frame virtual depth image.
[0044] As an optional embodiment, the determining the current frame virtual image of the to-be-tracked surgical target based on the current frame estimated pose includes: performing pose transformation on the initial three-dimensional points of the three-dimensional model of the to-be-tracked surgical target based on the current frame estimated pose, to obtain transformed three-dimensional points; projecting the transformed three-dimensional points based on an internal parameter matrix of the shooting device of the current frame observation image, to obtain plane pixel coordinates; and synthesizing a current frame virtual color image and a current frame virtual depth image based on the plane pixel coordinates, a depth value of the three-dimensional model, and texture information of the three-dimensional model.
[0045] Specifically, before the virtual image is generated, the three-dimensional model of the to-be-tracked surgical target and the internal parameter matrix of the shooting device used to shoot the observation image of the to-be-tracked surgical target are given in advance First, the initial three-dimensional points of the three-dimensional model are transformed by pose transformation based on the current frame estimated pose, to obtain transformed three-dimensional points of the current frame estimated pose in the camera coordinate system .
[0046] The internal parameter matrix , the transformed three-dimensional point is projected to obtain a planar pixel coordinate on the image plane. .
[0047] The current frame virtual color image and the current frame virtual depth image of the to-be-tracked surgical target corresponding to the current frame estimated pose are synthesized by a Volume Rendering method of the three-dimensional model according to the planar pixel coordinate, the depth value of the three-dimensional model, and the texture information of the three-dimensional model.
[0048] Optionally, the expression of the pose transformation of the initial three-dimensional point of the three-dimensional model of the to-be-tracked surgical target based on the current frame estimated pose to obtain the transformed three-dimensional point is as follows: ; wherein, is the transformed three-dimensional point; is a rotation parameter of the current frame estimated pose; is the initial three-dimensional point; is an offset parameter of the current frame estimated pose.
[0049] Optionally, the expression of the projection of the transformed three-dimensional point based on the internal parameter matrix of the photographing device of the current frame observation image to obtain the planar pixel coordinate is as follows: ; wherein, is the planar pixel coordinate; is the internal parameter matrix; is the transformed three-dimensional point.
[0050] Optionally, the three-dimensional model of the to-be-tracked surgical target is acquired based on a CAD design file or a CT image reconstruction result in a preoperative surgical preparation stage and is saved in a standard format such as STL or PLY.
[0051] In the case where the three-dimensional model is acquired based on the CT image reconstruction result, in the three-dimensional model acquisition process, the CT data is segmented and reconstructed by a reconstruction algorithm, and it is ensured that the model accuracy of the three-dimensional model is better than a preset accuracy threshold (such as 1 mm).
[0052] Step 130: based on the current frame virtual image and the current frame observation image, a first error value is calculated.
[0053] The first error value is an error value calculated according to a predetermined error index item (such as mean absolute error, mean square error, edge map intensity error, etc.) by using the current frame virtual image and the current frame observation image at any current time. The first error value is used to determine whether the current frame estimated pose obtained by optimization is suitable for application to the pose tracking at the current time.
[0054] Specifically, the first error value is calculated according to a predetermined error index item by using all or part of the current frame virtual image and the current frame observation image at the current time.
[0055] By calculating the first error value at each current time to determine whether the current frame estimated pose is suitable for application to the pose tracking at the current time, the complex clinical environment of the surgical target dynamic change being severe and the multiple surgical instruments interacting frequently can be quickly and real-timely responded, and thus the pose tracking is adjusted.
[0056] Optionally, the calculation of the first error value based on the current frame virtual image and the current frame observation image comprises: calculating the first error value based on a current frame virtual color image and a current frame observation color image; or, calculating the first error value based on a current frame virtual depth image and a current frame observation depth image; or, calculating a first color error value based on a current frame virtual color image and a current frame observation color image; calculating a first depth error value based on a current frame virtual depth image and a current frame observation depth image; and calculating the first error value based on a weighted sum of the first color error value and the first depth error value.
[0057] As an optional embodiment, the calculation of the first error value based on the current frame virtual image and the current frame observation image comprises: calculating the first error value based on a mean square error between mask region pixels of a current frame virtual depth image and mask region pixels of a current frame observation depth image.
[0058] Optionally, an expression for calculating the first error value based on the mean square error between the mask region pixels of the current frame virtual depth image and the mask region pixels of the current frame observation depth image is as follows: ; wherein, is the first error value; is a current frame observation depth image; is a current frame virtual depth image; is an effective pixel set of a mask region of a surgical target to be tracked; is the effective pixel set mask region pixels in the current frame virtual depth image.
[0059] Step 141: in the case that the first error value is greater than the preset error threshold, obtaining a plurality of first attitude candidates of the surgical target to be tracked based on the current frame observation image; calculating a second error value of the first attitude candidate based on a preset attitude optimization objective function, and continuously optimizing the first attitude candidate with the second error value being minimum as an optimization objective to obtain a plurality of second attitude candidates; determining a target estimated attitude from the second attitude candidates.
[0060] The preset error threshold is used to determine whether the current frame estimated attitude obtained by optimization is suitable for being applied to the pose tracking at the current time.
[0061] It should be noted that the preset error threshold can be determined comprehensively according to factors such as surgical precision requirements, shooting device performance, and pre-determined error index items, and the present application does not limit this.
[0062] For example, in the case that the first error value is calculated based on the mean square error between the mask region pixels of the current frame virtual depth image and the mask region pixels of the current frame observation depth image, the preset error threshold is set to 0.1. .
[0063] Specifically, in the case that the first error value is greater than the preset error threshold, it indicates that the estimated attitude obtained by optimization at the current time has a large error with the actual observation attitude of the surgical target to be tracked at the current time, and cannot be used for high-precision surgical target pose tracking, so it is determined that the tracking fails, and the pose initialization operation needs to be performed to generate new attitude candidates and select an estimated attitude that can be used for pose tracking.
[0064] The second error value is an error value calculated according to a preset attitude optimization objective function, and is used to continuously optimize the first attitude candidate to obtain the second attitude candidate.
[0065] First, the current frame observation image shot by the shooting device at the current time is processed to roughly generate a plurality of first attitude candidates of the surgical target to be tracked.
[0066] Secondly, the second error value of each first attitude candidate is calculated by using the preset attitude optimization objective function. For each first attitude candidate, the second error value of the first attitude candidate is continuously optimized by using a gradient optimization method or the like with the second error value of the first attitude candidate being minimum as an optimization objective, to obtain the second attitude candidate corresponding to the first attitude candidate. The step of optimizing the first attitude candidate to obtain the second attitude candidate is repeated until the second attitude candidates corresponding to all the first attitude candidates are obtained, Finally, a target estimated pose is selected from the second pose candidate, so as to use the target estimated pose for pose tracking of the surgical target to be tracked at the current time.
[0067] Optionally, the number of optimization iteration steps for continuously optimizing the first pose candidate is less than a preset number (such as 5, 10, 15, etc.). By setting the number of optimization iteration steps for continuously optimizing the first pose candidate to be less than the preset number, the real-time performance of intraoperative surgical target tracking can be ensured.
[0068] Step 142: In the case where the first error value is less than or equal to the preset error threshold, determining the target estimated pose based on the current frame estimated pose.
[0069] Specifically, in the case where the first error value is less than or equal to the preset error threshold, it indicates that the estimated pose obtained by optimization at the current time has a small error with the actual observed pose of the surgical target to be tracked at the current time, and can meet the high-precision pose tracking of the surgical target at the current time, without the need to regenerate a pose candidate and determine an estimated pose. At this time, the current frame estimated pose obtained by optimizing the current frame initial pose at the current time can be directly used as the target estimated pose of the surgical target to be tracked.
[0070] Step 150: Based on the target estimated pose, performing pose tracking on the surgical target to be tracked, and determining a current frame initial pose at the next time based on the target estimated pose.
[0071] Specifically, after obtaining the target estimated pose, on the one hand, the target estimated pose is used for pose tracking of the surgical target to be tracked, including spatial registration of the target estimated pose and the preoperative CT or MRI image, unification to a medical image coordinate system to realize coordinate unification, and real-time superposition of the target estimated pose on an intraoperative visualization terminal (such as a surgical screen, an augmented reality device, etc.) to dynamically display the navigation tracking result, so as to assist the doctor in precise identification and path planning, and significantly improve the intelligent operation and navigation assistance capability of minimally invasive surgery; on the other hand, the target estimated pose is used as the current frame initial pose at the next time.
[0072] The surgical target pose tracking method provided by the application can quickly respond to complex clinical environments such as rapid dynamic changes of the surgical target and frequent interactions of multiple surgical instruments, and can adjust the pose tracking in real time to ensure the effectiveness of the detection tracking. In addition, according to whether the tracking fails, the target estimated pose is directly determined by using the current frame estimated pose when the tracking succeeds, and the pose initialization and attitude optimization process is triggered only when the tracking fails. The pose initialization operation is performed according to the real-time acquisition of the current frame observation image, and the target estimated pose is confirmed from the optimized attitude candidate. The final stability can be restored, and a large number of candidate poses do not need to be generated at each time, which significantly reduces the computational overhead and computational delay, improves the overall processing efficiency, and realizes the overall pose tracking with strong adaptability, high precision, strong robustness and strong real-time performance for any rigid body in a complex surgical scene without physical markers and expensive external positioning devices.
[0073] Based on the above embodiment, as an optional embodiment, the current frame observation image includes a current frame observation color image and a current frame observation depth image; the second error value of the first attitude candidate is determined based on a first re-projection error of a color channel and a second re-projection error of a depth channel; the first re-projection error is determined based on the current frame observation color image and a first virtual color image; the second re-projection error is determined based on the current frame observation depth image and a first virtual depth image; and the first virtual color image and the first virtual depth image are both determined based on the first attitude candidate.
[0074] Specifically, in the optimization process of the first attitude candidate, for each first attitude candidate, the first virtual color image and the first virtual depth image corresponding to the first attitude candidate are obtained through coordinate system transformation, geometric projection and visualization technology. Further, the second error value of the first attitude candidate is calculated according to the attitude optimization objective function, and the calculation process includes calculating the first re-projection error of the color channel according to the current frame observation color image and the first virtual color image, calculating the second re-projection error of the depth channel according to the current frame observation depth image and the first virtual depth image, and further fusing the first re-projection error of the color channel and the second re-projection error of the depth channel to calculate the second error value of the first attitude candidate.
[0075] The surgical target posture tracking method provided by the present invention continuously optimizes the first posture candidate and obtains the second posture candidate by continuously comparing the first posture candidate with the actual collected current frame observation color image and current frame observation depth image during the optimization process of the first posture candidate, fusing the reprojection errors of the color channel and the depth channel, avoiding the strong dependence of traditional optimization methods on feature points or external markers, and can be applied to surgical target tracking and navigation in complex surgical scenarios.
[0076] As an optional embodiment, the first virtual color image and the first virtual depth image are determined based on the following steps: Based on the first posture candidate, the initial three-dimensional points of the three-dimensional model of the surgical target to be tracked are transformed in posture to obtain transformed three-dimensional points; based on the internal parameter matrix of the shooting device of the current frame observation image, the transformed three-dimensional points are projected to obtain plane pixel coordinates; based on the plane pixel coordinates, the depth value of the three-dimensional model and the texture information of the three-dimensional model, the first virtual color image and the first virtual depth image are synthesized.
[0077] Based on the above embodiment, as an optional embodiment, the expression of the posture optimization objective function is as follows: ; in, The second error value of the first pose candidate; is the rotation parameter of the first pose candidate; is the offset parameter of the first pose candidate; Observing a color image for the current frame; A first virtual color image determined based on the first pose candidate; Observing a depth image for the current frame; A first virtual depth image determined based on the first pose candidate; is the depth error weighting factor; is a set of valid pixels in the mask area of the surgical target to be tracked; The effective pixel set The pixels in the mask area.
[0078] Among them, the effective pixel set The mask area pixels in , i.e., the valid pixels in the mask area refer to the pixels in the mask area of the current frame observation image except for invalid pixels such as noise pixels and depth error pixels.
[0079] Specifically, in the refinement optimization process of each pose candidate, according to the expression of the aforementioned pose optimization objective function, starting from the effective pixels of the virtual image corresponding to the first pose candidate and the mask region of the current frame observation image, and taking the pixel-level rendering error in the mask region as the basis, the second error value of each first pose candidate is calculated, and the optimization target is to minimize the second error value of the first pose candidate, and the first pose candidate is continuously optimized by using gradient optimization and the like to obtain the second pose candidate corresponding to the first pose candidate.
[0080] The surgical target pose tracking method provided by the application can continuously compare each mask region pixel of the candidate virtual depth image and the candidate virtual color image corresponding to the first pose candidate in the optimization process with each mask region pixel of the actually collected current frame observation color image and current frame observation depth image, continuously optimize the first pose candidate by fusing the re-projection errors of the color channel and the depth channel, and then obtain the second pose candidate. The photometric consistency and geometric consistency of the image are fully utilized, the strong dependence on feature points or external markers in the traditional optimization method is avoided, the accuracy and robustness of the candidate pose refinement can be effectively improved, and the surgical target tracking and navigation in a complex surgical scene are particularly suitable.
[0081] Based on the above embodiment, as an optional embodiment, the optimization of the current frame initial pose of the surgical target to be tracked at the current time is performed to obtain a current frame estimated pose, which includes: The third error value of the current frame initial pose is calculated based on the pose optimization objective function, and the current frame initial pose is continuously optimized with the minimum third error value as the optimization target to obtain the current frame estimated pose.
[0082] Specifically, in the optimization process of the current frame initial pose at the current time, according to the expression of the aforementioned pose optimization objective function, starting from the effective pixels of the current frame virtual image corresponding to the current frame initial pose and the mask region of the current frame observation image, the third error value of the current frame initial pose is calculated, and the optimization target is to minimize the third error value of the current frame initial pose, and the current frame initial pose is continuously optimized by using gradient optimization and the like to obtain the current frame estimated pose.
[0083] Optionally, the number of optimization iteration steps for continuously optimizing the current frame initial pose is less than a preset number of times (such as 5 times, 10 times, 15 times, etc.). By setting the number of optimization iteration steps to be less than the preset number of times, the pose optimization efficiency and the surgical target tracking accuracy can be maximally considered.
[0084] The surgical target pose tracking method provided by the application uses the pre-designed pose optimization target function as the loss function, and performs pose optimization by fitting the pose optimization target function in the optimization process of the first pose candidate and the initial pose of the current frame, fully utilizes the photometric consistency and geometric consistency of the image, avoids the strong dependence of the traditional optimization method on the feature points or external markers, and ensures the consistency of the pose optimization in the surgical target pose tracking, which can effectively improve the accuracy and robustness of the candidate pose refinement.
[0085] Based on the above embodiment, as an optional embodiment, the obtaining of the plurality of first pose candidates of the surgical target to be tracked based on the current frame observation image comprises: inputting the current frame observation color image into the pre-trained mask region extraction model to obtain a mask region of the surgical target to be tracked output by the mask region extraction model; extracting target depth pixels from the current frame observation depth image based on the mask region; generating a local point cloud region based on the target depth pixels; generating a plurality of initial pose candidates based on the current frame observation color image, the current frame observation depth image, the three-dimensional model of the surgical target to be tracked, and the geometric center point of the local point cloud region; determining a plurality of first pose candidates based on a plurality of initial pose candidates.
[0086] Optionally, the current frame observation color image and the current frame observation depth image are both collected in real time by an RGBD camera in synchronization, the image resolution of the current frame observation color image is not less than 640px*480px, and the depth accuracy of the current frame observation depth image is better than ±1mm.
[0087] Specifically, in the preoperative preparation stage, the mask region extraction model needs to be constructed using an instance segmentation network in advance, and a plurality of training samples are collected to complete the pre-training of the mask region extraction model. Each training sample includes an observation color image sample of the surgical target to be tracked and a corresponding mask label.
[0088] Figure 2 is a flowchart of generating a first candidate pose provided by the application, as shown in Figure 2 The current frame observation color image and the current frame observation depth image are collected in synchronization by a shooting device, and frame synchronization processing is performed to ensure the frame synchronization of the current frame observation color image and the current frame observation depth image. The current frame observation color image at the current time is input into the pre-trained mask region extraction model to obtain a mask region of the surgical target to be tracked output by the mask region extraction model, i.e. a mask image.
[0089] The target depth pixels overlapping with the mask region are extracted from the current frame observation depth image, and the target depth pixels are converted to the camera coordinate system by using the calibration parameters of the shooting device, projected to the three-dimensional space to generate a local point cloud region of the dense scene, and a geometric center point of the local point cloud region is determined as an input for subsequent pose estimation.
[0090] Further, the current frame observation color image, the current frame observation depth image, the three-dimensional model of the to-be-tracked surgical target and the geometric center point of the local point cloud region are used to generate a plurality of rough initial pose candidates, and finally the initial pose candidates are directly taken as the first candidate poses to obtain a plurality of first candidate poses.
[0091] Optionally, the generating a plurality of initial pose candidates based on the current frame observation color image, the current frame observation depth image, the three-dimensional model of the to-be-tracked surgical target and the geometric center point of the local point cloud region comprises: constructing an input vector based on the current frame observation color image, the current frame observation depth image, the three-dimensional model of the to-be-tracked surgical target and the geometric center point of the local point cloud region; inputting the input vector into a pre-trained pose candidate generation model to obtain a plurality of initial pose candidates output by the pose candidate generation model; wherein the pose candidate generation model is pre-trained by using a plurality of vector samples and a plurality of pose candidate sample labels corresponding to the plurality of vector samples; each vector sample is determined based on a frame-synchronized observation color image sample and an observation depth image sample, a three-dimensional model sample of a to-be-tracked surgical target sample and a geometric center point sample of a local point cloud region sample.
[0092] Optionally, the pose candidate generation model is constructed based on a shallow lightweight neural network.
[0093] Optionally, the mask region extraction model supports intraoperative multi-type to-be-tracked surgical target identification and can adapt to visual differences under different surgical procedures.
[0094] Optionally, the number of the first pose candidates is less than a preset number threshold, so as to avoid excessive calculation overhead.
[0095] The surgical target pose tracking method provided by the application can quickly extract the mask region of the surgical target to be tracked from the current frame observation color image by using the pre-trained mask region extraction model at each time in the surgical target pose tracking at continuous time, and generate a local point cloud region according to the mask region and the current frame observation depth image, so as to simultaneously use the current frame observation color image, the current frame observation depth image, the three-dimensional model and the geometric center point of the local point cloud region to quickly generate the first pose candidate, and a better balance between the calculation efficiency and the tracking accuracy can be achieved, and the high redundant calculation overhead under the traditional unit'spherical sampling + disturbance' rule can be avoided.
[0096] Based on the above embodiment, as an optional embodiment, the first pose candidate is determined based on a plurality of the initial pose candidates, and the method comprises: The initial pose candidate is superimposed with a translation disturbance of a preset amplitude to generate the first pose candidate. The number of the first pose candidates is more than the number of the initial pose candidates.
[0097] Specifically, in combination with Figure 2 As shown in the figure, after generating a plurality of initial pose candidates, a small-amplitude translation disturbance of a preset amplitude (such as ±5mm) is superimposed on each initial pose candidate to generate a plurality of first pose candidates, each of which includes a rotation parameter and an offset parameter.
[0098] For example, the number of the first pose candidates is 50.
[0099] The surgical target pose tracking method provided by the application can generate a plurality of first pose candidates by generating a small-amplitude translation disturbance on the initial pose candidate generated by using the current frame observation color image, the current frame observation depth image, the three-dimensional model and the geometric center point of the local point cloud region, so as to ensure that the pose space has sufficient coverage, obtain a change range covering multiple viewing angles and close to the true value, and achieve a good balance between the calculation efficiency and the tracking accuracy while avoiding high calculation overhead.
[0100] Based on the above embodiment, as an optional embodiment, the target estimated pose is determined from the second pose candidate, and the method comprises: A second virtual depth image, a second virtual color image and a virtual mask region of each of the second pose candidates are determined. A first score is determined based on the average absolute error between the second virtual depth image and the current frame observation depth image. A second score is determined based on the overlapping degree between the virtual mask region and the mask region. determining a third score based on an edge map error term between the second virtual color image and the current frame observed color image; determining a match quality score for each second pose candidate based on the first score, the second score, and the third score; The estimated target pose is determined based on the second pose candidate with the highest matching quality score.
[0101] Specifically, when determining the estimated target pose from all second pose candidates, the second virtual depth image, the second virtual color image, and the virtual mask area corresponding to each second pose candidate are first determined.
[0102] A matching quality score is calculated for each second pose candidate. The first score is calculated by calculating the mean absolute error (MAE) of all masked pixels between the second virtual depth image and the observed depth image of the current frame. The second score is obtained by using the Intersection over Union (IoU) metric to assess the degree of overlap between the virtual masked area corresponding to the second pose candidate and the masked area corresponding to the observed depth image of the current frame. This is used to evaluate edge alignment and improve edge structure consistency. Furthermore, an edge map error term is introduced to compare edge strength differences and texture similarity based on the edge map error term between the second virtual color image and the observed color image of the current frame, yielding a third score.
[0103] The matching quality score of each second pose candidate is determined by using a weighted sum of the first score, the second score, and the third score of each second pose candidate.
[0104] Optionally, the weight of the first score , the weight of the second score and the weight of the third score The sum is 1.
[0105] According to the matching quality score of each second pose candidate, all the second pose candidates obtained after refinement and optimization are sorted, and the second pose candidate with the highest matching quality score is used as the final output target estimated pose, so as to use the target estimated pose to track the surgical target pose.
[0106] Optionally, the calculation formula for determining the first score based on the mean absolute error between the second virtual depth image and the current frame observed depth image is as follows: ; in, For the a first score of a second pose candidate; Observe the depth image for the current frame; a second virtual depth image for the first second pose candidate; a second virtual depth image for the first second pose candidate; a set of valid pixels of a mask region of the surgical target to be tracked; a mask region pixel in the set of valid pixels a mask region pixel in the set of valid pixels
[0107] Optionally, the second score is determined based on an overlap between the virtual mask region and the mask region, and a calculation formula of the second score is as follows: ; wherein, a second score of the first second pose candidate; a virtual mask region of the first second pose candidate; a virtual mask region of the first second pose candidate; a mask region.
[0108] Optionally, the third score is determined based on an edge map error term between the second virtual color image and the current frame observation color image, and a calculation formula of the third score is as follows: ; wherein, a third score of the first second pose candidate; a current frame observation color image; a second virtual color image of the first second pose candidate; a current frame observation color image; a second virtual color image of the first second pose candidate; a set of valid pixels of a mask region of the surgical target to be tracked; a mask region pixel in the set of valid pixels a mask region pixel in the set of valid pixels
[0109] The surgical target pose tracking method provided by the application realizes pose candidate screening without neural network training, and has good generalization ability and interpretability, and is especially suitable for complex surgical environments such as occlusion, reflection and weak texture.
[0110] As an optional embodiment, the second virtual depth image, the second virtual color image and the virtual mask region of each second pose candidate are determined by: Based on the second pose candidate, the initial three-dimensional points of the three-dimensional model of the to-be-tracked surgical target are pose-transformed to obtain transformed three-dimensional points; based on an internal parameter matrix of a shooting device of the current frame observation image, the transformed three-dimensional points are projected to obtain plane pixel coordinates; based on the plane pixel coordinates, a depth value of the three-dimensional model and texture information of the three-dimensional model, the second virtual depth image and a second virtual color image are synthesized; and the second pose candidate is input into a pre-trained mask region extraction model to obtain the virtual mask region of the second pose candidate output by the mask region extraction model.
[0111] In combination Figure 2 As shown in the whole, the surgical target pose tracking method provided by the application forms a complete "estimation-judgment-optimization-feedback" closed loop structure in the process design, first acquires a three-dimensional model reconstructed by CAD or CT in the preoperative preparation stage, uses an RGBD camera to collect a current frame observation color image and a current frame observation depth image at each moment in the operation, triggers an initialization mechanism at a time when simple tracking fails to avoid frequent re-computation, extracts a model extraction mask region from the current frame observation color image using a mask region, generates a local point cloud region in combination with the current frame observation depth image, and uses a pose candidate generation model constructed based on a shallow lightweight neural network to generate a plurality of initial pose candidates with higher computational efficiency, obtains a plurality of first pose candidates after a small translation disturbance, further optimizes the plurality of first pose candidates based on rendering errors to obtain a plurality of second pose candidates, and determines a target estimated pose output from the second pose candidate to realize tracking of the surgical target. The target pose selection scoring method designed by fusing three-channel visual indicators can increase the robustness of complex environments, and is particularly suitable for intraoperative scenes such as occlusion, reflection and weak texture structure. By supporting continuous frame real-time tracking, seamless registration with medical images (CT / MRI) can be realized, and the visualization and operation accuracy are enhanced; the real-time 6DoF pose tracking of any rigid body target can be supported, and the adaptability is stronger; physical markers are not required, and the method is suitable for various surgical instruments or anatomical structures, thereby reducing the system deployment cost and use complexity; high robustness and high precision spatial pose tracking can be realized in a complex and dynamic surgical environment, and the real-time, adaptability and robustness of tracking are greatly improved, and the method can be widely applied to the fields of surgical navigation, medical image registration and augmented reality guidance.
[0112] Figure 3 is a structural schematic diagram of a surgical target pose tracking device provided by the application, as Figure 3 shown, the surgical target pose tracking device includes but is not limited to a pose estimation module 310, a virtual image generation module 320, an error calculation module 330, a first pose determination module 341, a second pose determination module 342 and a pose tracking module 350.
[0113] The pose estimation module 310 is configured to optimize the initial pose of the current frame of the to-be-tracked surgical target at the current time, to obtain a current frame estimated pose.
[0114] The virtual image generation module 320 is configured to determine a current frame virtual image of the to-be-tracked surgical target based on the current frame estimated pose.
[0115] The error calculation module 330 is configured to calculate a first error value based on the current frame virtual image and a current frame observation image.
[0116] The first pose determination module 341 is configured to, in a case where the first error value is greater than a preset error threshold, acquire a plurality of first pose candidates of the to-be-tracked surgical target based on the current frame observation image, calculate a second error value of the first pose candidates based on a preset pose optimization objective function, continuously optimize the first pose candidates with the second error value being minimum as an optimization objective, to obtain a plurality of second pose candidates, and determine a target estimated pose from the second pose candidates.
[0117] The second pose determination module 342 is configured to, in a case where the first error value is less than or equal to the preset error threshold, determine the target estimated pose based on the current frame estimated pose.
[0118] The pose tracking module 350 is configured to perform pose tracking on the to-be-tracked surgical target based on the target estimated pose, and determine the initial pose of the current frame at the next time based on the target estimated pose.
[0119] It should be noted that the surgical target pose tracking device provided by the present application can execute the surgical target pose tracking method described in any of the above embodiments when it is actually running, and the present embodiment will not be described here.
[0120] The surgical target pose tracking device provided by the application can quickly respond to complex clinical environments such as rapid dynamic changes of a surgical target, frequent interaction of multiple surgical instruments and the like, and can adjust the pose tracking in real time to ensure the detection tracking effectiveness. In addition, according to the different tracking failures, the target estimated pose is directly determined by using the current frame estimated pose when the tracking is successful, and only when the tracking fails, the pose initialization and attitude optimization process are triggered, the pose initialization operation is performed according to the real-time acquisition of the current frame observation image, the attitude candidate is continuously optimized by taking the minimum error calculated by the preset attitude optimization target function as the target, and the target estimated pose is confirmed from the optimized attitude candidate. The final stability can be restored, and a large number of candidate poses do not need to be generated at each time, the calculation overhead and the calculation delay are significantly reduced, and the overall processing efficiency is improved, so that the overall pose tracking with strong adaptability, high precision, strong robustness and strong real-time performance for any rigid body in a complex surgical scene is realized without physical markers and expensive external positioning devices.
[0121] Figure 4 is a structural schematic diagram of an electronic device provided by the application, like Figure 4As shown, the electronic device can include a processor 410, a communications interface 420, a memory 430, and a communications bus 440, wherein the processor 410, the communications interface 420, and the memory 430 complete communications with each other through the communications bus 440. The processor 410 can invoke a logical instruction in the memory 430 to execute the surgical target pose tracking method provided by any of the above embodiments, which includes but is not limited to the following steps: optimizing a current frame initial pose of a surgical target to be tracked at a current time to obtain a current frame estimated pose; determining a current frame virtual image of the surgical target to be tracked based on the current frame estimated pose; calculating a first error value based on the current frame virtual image and a current frame observation image; in a case where the first error value is greater than a preset error threshold, acquiring a plurality of first pose candidates of the surgical target to be tracked based on the current frame observation image; calculating a second error value of the first pose candidate based on a pre-set pose optimization objective function, and continuously optimizing the first pose candidate to obtain a plurality of second pose candidates with the second error value being the minimum as the optimization objective; determining a target estimated pose from the second pose candidate; in a case where the first error value is less than or equal to the preset error threshold, determining the target estimated pose based on the current frame estimated pose; tracking the pose of the surgical target to be tracked based on the target estimated pose, and determining a current frame initial pose at a next time based on the target estimated pose.
[0122] In addition, the logical instruction in the memory 430 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0123] In another aspect, the present application also provides a computer program product comprising a computer program, which can be stored on a non-transitory computer readable storage medium, and when the computer program is executed by a processor, the computer can execute the surgical target pose tracking method provided by any of the above embodiments, which comprises but is not limited to the following steps: optimizing a current frame initial pose of a surgical target to be tracked at a current time to obtain a current frame estimated pose; determining a current frame virtual image of the surgical target to be tracked based on the current frame estimated pose; calculating a first error value based on the current frame virtual image and a current frame observation image; in the case that the first error value is greater than a preset error threshold, obtaining a plurality of first pose candidates of the surgical target to be tracked based on the current frame observation image; calculating a second error value of the first pose candidate based on a pre-set pose optimization objective function, and continuously optimizing the first pose candidate to obtain a plurality of second pose candidates with the minimum second error value as the optimization objective; determining a target estimated pose from the second pose candidates; in the case that the first error value is less than or equal to the preset error threshold, determining the target estimated pose based on the current frame estimated pose; and tracking the pose of the surgical target to be tracked based on the target estimated pose, and determining a current frame initial pose at a next time based on the target estimated pose.
[0124] In yet another aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the surgical target pose tracking method provided by any of the above embodiments is implemented, which comprises but is not limited to the following steps: optimizing a current frame initial pose of a surgical target to be tracked at a current time to obtain a current frame estimated pose; determining a current frame virtual image of the surgical target to be tracked based on the current frame estimated pose; calculating a first error value based on the current frame virtual image and a current frame observation image; in the case that the first error value is greater than a preset error threshold, obtaining a plurality of first pose candidates of the surgical target to be tracked based on the current frame observation image; calculating a second error value of the first pose candidate based on a pre-set pose optimization objective function, and continuously optimizing the first pose candidate to obtain a plurality of second pose candidates with the minimum second error value as the optimization objective; determining a target estimated pose from the second pose candidates; in the case that the first error value is less than or equal to the preset error threshold, determining the target estimated pose based on the current frame estimated pose; and tracking the pose of the surgical target to be tracked based on the target estimated pose, and determining a current frame initial pose at a next time based on the target estimated pose.
[0125] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0126] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0127] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A surgical target posture tracking method, characterized in that: include: Optimizing the initial pose of the current frame of the surgical target to be tracked at the current moment to obtain the estimated pose of the current frame; Determining a current frame virtual image of the surgical target to be tracked based on the estimated posture of the current frame; Calculating a first error value based on the current frame virtual image and the current frame observation image; When the first error value is greater than a preset error threshold, a plurality of first posture candidates of the surgical target to be tracked are obtained based on the current frame observation image; a second error value of the first posture candidate is calculated based on a preset posture optimization objective function, and the first posture candidate is continuously optimized with minimizing the second error value as the optimization objective, to obtain a plurality of second posture candidates; Determining an estimated target pose from the second pose candidates; When the first error value is less than or equal to the preset error threshold, determining the target estimated pose based on the current frame estimated pose; Based on the estimated target posture, the posture of the surgical target to be tracked is tracked, and based on the estimated target posture, the initial posture of the current frame at the next moment is determined.
2. The surgical target posture tracking method according to claim 1, characterized in that: The current frame observation image includes a current frame observation color image and a current frame observation depth image; The second error value of the first posture candidate is determined based on the first reprojection error of the color channel and the second reprojection error of the depth channel; the first reprojection error is determined based on the current frame observed color image and the first virtual color image; the second reprojection error is determined based on the current frame observed depth image and the first virtual depth image; the first virtual color image and the first virtual depth image are both determined based on the first posture candidate.
3. The surgical target posture tracking method according to claim 2, characterized in that: The expression of the posture optimization objective function is as follows: ; in, The second error value of the first pose candidate; is the rotation parameter of the first pose candidate; is the offset parameter of the first pose candidate; Observing a color image for the current frame; A first virtual color image determined based on the first pose candidate; Observing a depth image for the current frame; A first virtual depth image determined based on the first pose candidate; is the depth error weighting factor; is a set of valid pixels in the mask area of the surgical target to be tracked; The effective pixel set The pixels in the mask area.
4. The surgical target posture tracking method according to claim 1, characterized in that: The step of acquiring a plurality of first posture candidates of the surgical target to be tracked based on the current frame observation image includes: Inputting the current frame observation color image into a pre-trained mask region extraction model to obtain the mask region of the surgical target to be tracked output by the mask region extraction model; Extracting target depth pixels from the current frame observation depth image based on the mask area; Generating a local point cloud region based on the target depth pixels; generating a plurality of initial pose candidates based on the current frame observation color image, the current frame observation depth image, the three-dimensional model of the surgical target to be tracked, and the geometric center point of the local point cloud area; Based on the plurality of initial pose candidates, a plurality of first pose candidates are determined.
5. The surgical target posture tracking method according to claim 4, characterized in that: The determining of a plurality of first posture candidates based on the plurality of initial posture candidates comprises: superimposing a translation disturbance of a preset magnitude on the initial posture candidate to generate the first posture candidate; The number of the first pose candidates is greater than the number of the initial pose candidates.
6. The surgical target posture tracking method according to claim 4, characterized in that: Determining the estimated target pose from the second pose candidates includes: determining a second virtual depth image, a second virtual color image, and a virtual mask area for each of the second pose candidates; determining a first score based on a mean absolute error between the second virtual depth image and the current frame observed depth image; determining a second score based on a degree of overlap between the virtual mask area and the mask area; determining a third score based on an edge map error term between the second virtual color image and the current frame observed color image; determining a match quality score for each second pose candidate based on the first score, the second score, and the third score; The estimated target pose is determined based on the second pose candidate with the highest matching quality score.
7. The surgical target posture tracking method according to claim 1, characterized in that: The calculating a first error value based on the current frame virtual image and the current frame observation image includes: The first error value is calculated based on a mean square error between pixels in the mask area of the virtual depth image of the current frame and pixels in the mask area of the observed depth image of the current frame.
8. The surgical target posture tracking method according to claim 3, characterized in that: The step of optimizing the initial pose of the current frame of the surgical target to be tracked at the current moment to obtain the estimated pose of the current frame includes: The third error value of the initial posture of the current frame is calculated based on the posture optimization objective function, and the initial posture of the current frame is continuously optimized with the minimum third error value as the optimization goal to obtain the estimated posture of the current frame.
9. The surgical target posture tracking method according to claim 1, characterized in that: The step of estimating the posture of the current frame and determining the current frame virtual image of the surgical target to be tracked includes: Based on the estimated posture of the current frame, performing posture transformation on the initial three-dimensional points of the three-dimensional model of the surgical target to be tracked to obtain transformed three-dimensional points; Projecting the transformed three-dimensional point based on the internal parameter matrix of the shooting device of the current frame observation image to obtain plane pixel coordinates; A current frame virtual color image and a current frame virtual depth image are synthesized based on the plane pixel coordinates, the depth value of the three-dimensional model and the texture information of the three-dimensional model.
10. A surgical target posture tracking device, characterized in that: include: A posture estimation module is used to optimize the initial posture of the current frame of the surgical target to be tracked at the current moment to obtain the estimated posture of the current frame; A virtual image generation module, configured to determine a current frame virtual image of the surgical target to be tracked based on the estimated posture of the current frame; an error calculation module, configured to calculate a first error value based on the current frame virtual image and the current frame observation image; a first posture determination module configured to obtain, based on the current frame observation image, a plurality of first posture candidates of the surgical target to be tracked when the first error value is greater than a preset error threshold; calculate a second error value of the first posture candidate based on a preset posture optimization objective function, and continuously optimize the first posture candidate with minimizing the second error value as an optimization objective, to obtain a plurality of second posture candidates; Determining an estimated target pose from the second pose candidates; A second posture determination module is configured to determine the target estimated posture based on the current frame estimated posture when the first error value is less than or equal to the preset error threshold; The posture tracking module is used to track the posture of the surgical target to be tracked based on the estimated posture of the target, and determine the initial posture of the current frame at the next moment based on the estimated posture of the target.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the surgical target posture tracking method according to any one of claims 1 to 9 is implemented.
12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the surgical target posture tracking method according to any one of claims 1 to 9 is implemented.
13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the surgical target posture tracking method according to any one of claims 1 to 9 is implemented.