Comprehensive method for preoperative focus identification and intraoperative guidance
By fusing OCT and microscopic images with deep learning and combining them with servo control, the problems of positioning accuracy and real-time performance in preoperative lesion identification and intraoperative guidance have been solved, achieving high-precision lesion identification and instrument guidance, and improving the safety and operational precision of the surgery.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SMART VISION MEDICAL ROBOT (HARBIN) CO LTD
- Filing Date
- 2025-12-11
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies suffer from insufficient accuracy in lesion localization, inadequate image segmentation, poor real-time registration during surgery, and limited instrument guidance accuracy in preoperative lesion identification and intraoperative guidance. In particular, micron-level errors have a significant impact on surgical outcomes in microsurgical or minimally invasive procedures.
By fusing optical coherence tomography (OCT) technology with microscopic images, combined with deep learning and servo control, high-precision alignment between intraoperative images and preoperative models is achieved through multimodal image fusion, 3D fine modeling, and lesion identification, and errors during the operation are compensated in real time.
It achieves high-precision lesion identification, high model stability, and high-reliability instrument guidance, improving the accuracy and safety of surgery and meeting the real-time navigation needs of microsurgery and minimally invasive surgery.
Smart Images

Figure CN121937367A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a comprehensive method for preoperative lesion identification and intraoperative guidance, belonging to the technical field of preoperative lesion identification and intraoperative guidance. Background Technology With the development of medical imaging equipment and intelligent algorithms, preoperative lesion identification and intraoperative guidance technologies have gradually become important supporting links for precision medicine and intelligent surgery. Through comprehensive analysis of preoperative image data, doctors can understand the spatial morphology, depth distribution, and relationship with surrounding tissues of lesions before surgery, thereby formulating a more reasonable surgical path. However, in clinical applications, traditional image processing and navigation methods still have significant shortcomings and cannot meet the real-time and stability requirements of high-precision surgery.
[0002] Existing lesion identification methods primarily rely on single-modality imaging, such as structural imaging data like MRI or CT. While these images provide information on tissue layers, their ability to identify lesion boundaries, fine structures, and tissue texture is limited. Although optical coherence tomography (OCT) and surgical microscopy offer high resolution, their limited field of view and penetration depth make it difficult to fully present the global spatial information of lesions. The limitations of single-modality imaging lead to insufficient lesion localization accuracy, especially in cases of complex organ surface morphology or low tissue contrast, where lesion areas are prone to misidentification and blurred boundaries.
[0003] Furthermore, traditional image registration and model reconstruction methods are mostly based on geometric features or grayscale similarity. However, differences in resolution, viewing angle deviations, and imaging noise exist between different modalities of images, making it difficult for conventional algorithms to guarantee precise consistency in spatial correspondence. This results in positional deviations between the preoperative reconstructed model and the actual intraoperative image, affecting the accurate localization of the lesion and the precision of surgical instrument guidance. Especially in microsurgical or minimally invasive surgical scenarios, even micrometer-level errors can significantly impact surgical outcomes.
[0004] During the surgical procedure, surgeons need to determine the position of instruments and the boundaries of lesions in real time based on microscope or endoscopic images. However, traditional navigation systems often cannot achieve dynamic registration between intraoperative images and preoperative models, leading to unstable real-time guidance. When surgical procedures cause tissue deformation, displacement, or changes in perspective, the system struggles to update model information in a timely manner, resulting in the accumulation of navigation errors. Furthermore, existing instrument guidance often employs rigid methods based on optical or mechanical positioning, lacking the ability to adaptively compensate for tissue changes and effectively cope with the effects of nonlinear deformations in the surgical environment.
[0005] In summary, existing technologies still face significant bottlenecks in the coordination of preoperative lesion identification and intraoperative guidance, primarily manifested in insufficient lesion localization accuracy, imprecise image segmentation, poor real-time intraoperative registration, and limited instrument guidance precision. How to combine multimodal image fusion, deep learning segmentation, spatial registration, and intelligent control technologies to construct a comprehensive system capable of operating throughout the entire preoperative and intraoperative process is a key technical problem that urgently needs to be solved in the field of intelligent surgery. Summary of the Invention
[0006] This invention addresses the aforementioned shortcomings by proposing a comprehensive solution that integrates multimodal imaging to achieve detailed lesion modeling and high-precision intraoperative guidance, providing technical support for intelligent and precise surgery.
[0007] This invention aims to propose a comprehensive method for preoperative lesion identification and intraoperative guidance. Through multimodal image fusion, three-dimensional fine modeling, lesion identification, and intraoperative fine registration, it achieves high-precision alignment between intraoperative images and preoperative models, and provides instrument movement guidance, thereby improving the accuracy and safety of surgery.
[0008] To address the above problems, this invention provides a comprehensive method for preoperative lesion identification and intraoperative guidance, comprising the following steps: Step 1: Acquire structural data of deep tissues in the target area using optical coherence tomography (OCT) technology; Step 2: Deeply fuse the spatially registered microscope images with OCT 3D data to construct a 3D fusion model and a 3D coordinate system that simultaneously contains rich texture information and fine structural information; Step 3: Apply a pre-trained deep segmentation network to perform automated segmentation and spatial localization of the lesion region in the 3D fusion model; the final output of this step is the precise contour of the lesion and its detailed information in the 3D coordinate system. Step 4: Perform a comprehensive 3D reconstruction of the patient's preoperative imaging data to generate a detailed preoperative planning model; on this model, we will clearly mark the lesion area and plan the ideal surgical path; Step 5: Real-time acquisition of microscope video stream and OCT B-scan images during the surgery; Step Six: Construct an advanced neural network model with a dual-stream multimodal deep network structure; this model, trained on a large amount of data, can efficiently process multimodal images acquired during surgery and identify the location of lesions in real time and accurately. Step 7: Accurately map the three-dimensional deformation field output by the dual-stream multimodal deep network and the real-time lesion location onto the three-dimensional coordinate system of the three-dimensional fusion model established before surgery; Ultimately, the movement of the instruments is dynamically adjusted through a servo control system, thereby enabling real-time compensation for errors that occur during the surgical procedure.
[0009] Furthermore, step one also includes: constructing a high-resolution three-dimensional tissue volume model using a tomographic reconstruction algorithm.
[0010] Furthermore, step one also includes: using a joint optimization algorithm based on feature points and surface curvature to accurately register the microscope image; the joint optimization algorithm accurately maps the microscope image onto the surface of the three-dimensional tissue volume model through feature matching and photometric correction, thereby ensuring the spatial consistency of the multimodal image.
[0011] Furthermore, in step two, during the deep fusion process, cross-modal features are extracted using a deep learning network. This aims to enhance the contrast between tissue boundaries and lesion areas, ultimately achieving a three-dimensional reconstruction that can maintain high detail.
[0012] Furthermore, the preoperative imaging data includes OCT volumetric data and structural images of the retina.
[0013] Furthermore, step five also includes preprocessing the real-time acquired data. This preprocessing includes frame synchronization, image denoising, data standardization, and time calibration. After these processes, a time-series data stream that can be directly used by deep learning models will be generated.
[0014] Furthermore, the steps for constructing an advanced neural network model of a dual-stream multimodal deep network include: acquiring multi-frame OCT B-Scan slice sequence data from the patient's fundus. I OCT ( x, z A layer extraction algorithm using a 3D U-Net network was employed to segment the retinal layer of OCT data, obtaining depth location functions of the main tissue interfaces within the retina, which were then converted into real-space 3D coordinates: P xyz ( idx, kdy, jdz Obtain the 3D point cloud data of the corresponding layer. P ILM , P IPL A three-dimensional distance scalar field (Volume) to the point cloud surface is constructed using a fast nearest neighbor search algorithm. x, y, zUsing the Marching Cubes algorithm, linear interpolation is performed at the boundary locations to solve for the intersection points, forming triangular facets, which are then combined to generate a three-dimensional mesh model of multi-layered retinal tissue. High-frequency surface motion (optical flow / texture) from microscopic video is fused with low-frequency but scaled depth information from intraoperative OCT to reconstruct a dense 3D deformation field. The preoperative retinal model is then remapped intraoperatively to the current anatomical morphology, thereby reflecting and compensating for retinal deformation in real time.
[0015] Furthermore, step seven also includes: by calculating the registration matrix in real time and continuously comparing it with the position of the surgical instrument tip, the system can accurately identify any deviation.
[0016] The beneficial effects of this invention are: This invention employs a combined acquisition method of OCT and microscopic images, integrating tomographic reconstruction algorithms with photometric correction registration strategies to achieve spatial consistency between different modalities. Through a joint optimization method based on feature points and surface curvature, it effectively aligns deep tissue and surface texture information, significantly improving the registration accuracy and spatial integrity of multimodal images, thus providing a reliable data foundation for subsequent lesion identification.
[0017] By employing a deep learning-driven multimodal fusion network to collaboratively model structural and textural information, the system enhances the contrast and morphological details of lesion boundaries, achieving high-fidelity 3D reconstruction. This provides a precise model foundation for subsequent identification and navigation. A multi-scale segmentation network automatically identifies lesion regions, and an attention mechanism enables refined segmentation and spatial localization, reducing reliance on manual intervention and improving segmentation accuracy and stability. During the intraoperative phase, multimodal images are acquired in real-time and processed synchronously and standardized, with dynamic updates to lesion locations achieved using a depth model. The system automatically calculates errors and performs servo compensation based on registration results and instrument positions, ensuring high precision and continuity during the guidance process.
[0018] In summary, this invention achieves high precision in lesion identification, high stability in model registration, and high reliability in instrument guidance by integrating multimodal imaging, depth modeling, and real-time control, providing efficient and safe technical support for microsurgical and minimally invasive surgery. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the overall technical process of the present invention.
[0020] Figure 2 This is a schematic diagram of the structure for modeling the fusion of OCT and microscope images in this invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] In this invention, unless otherwise explicitly specified and limited, the terms "connected," "linked," and "fixed" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0023] In this invention, the terms "first" and "second" are used only to distinguish similar components / parts in different positions or with different characteristics, and have no other limiting meaning; "upper" refers to the direction in which each component is away from the ground, and "lower" refers to the direction in which each component is away from the ground.
[0024] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can include direct contact between the first and second features, or contact between the first and second features through another feature between them. Furthermore, "above," "over," and "on top" of the second feature includes the first feature directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature includes the first feature directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature.
[0025] This invention provides a comprehensive method for preoperative lesion identification and intraoperative guidance, comprising the following steps: Step 1: Acquire structural data of deep tissues in the target area using optical coherence tomography (OCT) technology; Step 2: Deeply fuse the spatially registered microscope images with OCT 3D data to construct a 3D fusion model and a 3D coordinate system that simultaneously contains rich texture information and fine structural information; Step 3: Apply a pre-trained deep segmentation network to perform automated segmentation and spatial localization of the lesion region in the 3D fusion model; the final output of this step is the precise contour of the lesion and its detailed information in the 3D coordinate system. Step 4: Perform a comprehensive 3D reconstruction of the patient's preoperative imaging data to generate a detailed preoperative planning model; on this model, we will clearly mark the lesion area and plan the ideal surgical path; Step 5: Real-time acquisition of microscope video stream and OCT B-scan images during the surgery; Step Six: Construct an advanced neural network model of a dual-stream multimodal deep network; this model, trained on a large amount of data, can efficiently process multimodal images acquired during surgery and identify the location of lesions in real time and accurately. Step 7: Accurately map the three-dimensional deformation field and real-time lesion location output by the dual-stream multimodal deep network into the three-dimensional coordinate system of the three-dimensional fusion model established before surgery; Ultimately, the movement of the instruments is dynamically adjusted through a servo control system, thereby enabling real-time compensation for errors that occur during the surgical procedure.
[0026] Furthermore, step one also includes: constructing a high-resolution three-dimensional tissue volume model using a tomographic reconstruction algorithm.
[0027] Furthermore, step one also includes: using a joint optimization algorithm based on feature points and surface curvature to accurately register the microscope image; the joint optimization algorithm accurately maps the microscope image onto the surface of the three-dimensional tissue volume model through feature matching and photometric correction, thereby ensuring the spatial consistency of the multimodal image.
[0028] Furthermore, in step two, during the deep fusion process, cross-modal features are extracted using a deep learning network. This aims to enhance the contrast between tissue boundaries and lesion areas, ultimately achieving a three-dimensional reconstruction that can maintain high detail.
[0029] Furthermore, the preoperative imaging data includes OCT volumetric data and structural images of the retina.
[0030] Furthermore, step five also includes preprocessing the real-time acquired data. This preprocessing includes frame synchronization, image denoising, data standardization, and time calibration. After these processes, a time-series data stream that can be directly used by deep learning models will be generated.
[0031] Furthermore, the steps for constructing an advanced neural network model of a dual-stream multimodal deep network include: acquiring multi-frame OCT B-Scan slice sequence data from the patient's fundus. I OCT ( x, z A layer extraction algorithm using a 3D U-Net network was employed to segment the retinal layer of OCT data, obtaining depth location functions of the main tissue interfaces within the retina, which were then converted into real-space 3D coordinates: Pxyz ( idx, kdy, jdz Obtain the 3D point cloud data of the corresponding layer. P ILM , P IPL A three-dimensional distance scalar field (Volume) to the point cloud surface is constructed using a fast nearest neighbor search algorithm. x, y, z Using the Marching Cubes algorithm, linear interpolation is performed at the boundary locations to solve for the intersection points, forming triangular facets, which are then combined to generate a three-dimensional mesh model of multi-layered retinal tissue. High-frequency surface motion (optical flow / texture) from microscopic video is fused with low-frequency but scaled depth information from intraoperative OCT to reconstruct a dense 3D deformation field. The preoperative retinal model is then remapped intraoperatively to the current anatomical morphology, thereby reflecting and compensating for retinal deformation in real time.
[0032] Furthermore, step seven also includes: by calculating the registration matrix in real time and continuously comparing it with the position of the surgical instrument tip, the system can accurately identify any deviation.
[0033] Example 1 The overall process of this invention consists of core stages such as OCT image modeling, microscopic image registration, multimodal fusion, lesion identification, preoperative planning, deep learning recognition module, spatial registration module, and instrument guidance.
[0034] In the preoperative stage, deep tissue structure data of the target area of the patient were acquired using a Fourier Domain Optical Coherence Tomography (OCT) system. The scanning wavelength was centered at 850 nm, with an axial resolution of approximately 5 μm and a lateral resolution of approximately 10 μm. The original volume data size was 512×512×512 voxels, and the sampling interval was 0.01 mm. A three-dimensional tissue model was constructed using tomographic interpolation and scattering compensation algorithms, and noise suppression and voxel homogenization were performed to obtain a high-fidelity three-dimensional structural morphology.
[0035] In the microscopic image registration stage, the system first acquires a microscope image corresponding to the OCT imaging area and performs multimodal geometric alignment based on this. First, key texture feature points are extracted from the microscope image, and an initial matching relationship is established by combining surface curvature information. Then, a local iterative optimization algorithm based on Iterative Closest Point (ICP) is used to progressively refine the spatial registration result. After geometric alignment is completed, illumination consistency constraints and a color mapping function are introduced to perform photometric correction and color fusion on the microscope image, thereby achieving high-precision consistency between the OCT image and the microscope image at both the spatial structure and illumination feature levels.
[0036] In the multimodal fusion stage, deep structural features from OCT images and surface texture features from microscopic images are extracted using convolutional neural networks, and cross-modal fusion is performed at the feature level. The network structure adopts a dual-stream encoder-decoder architecture, which contains two parallel processing paths that extract features from data of different modalities. The structure channel is responsible for volumetric morphology reconstruction, while the texture channel is responsible for detail enhancement. The two are fused at the decoding layer to output a high-fidelity 3D model with consistent texture and structure.
[0037] In the lesion segmentation and localization stage, based on the 3D fusion model, an improved U-Net network is used for lesion mask prediction. The model weights are jointly optimized using Dice loss and cross-entropy loss, with the loss function being: ,in The Dice loss is used for contour consistency constraints, and BCE is used for pixel-level classification. The output includes the 3D volume of the lesion, its center coordinates, and boundary surface parameters.
[0038] In the preoperative planning stage, a geometric analysis of the 3D lesion model (preoperative planning model) is performed to conduct in-depth geometric morphological analysis of the reconstructed 3D lesion model. First, the least-squares ellipsoid fitting algorithm is used to accurately obtain the parameterized equations of the lesion surface. Based on this parameterized model, the system can accurately define the reachable space of surgical instruments in 3D space. To determine the optimal surgical path, we introduce an optimization strategy based on inverse kinematics. This strategy is centered on a dual objective function: "shortest surgical path length" and "maximum perpendicularity of instrument incident angle". To solve this multi-objective optimization problem, a genetic algorithm is used, with the following key parameters set: population size of 50, maximum number of iterations of 100, crossover probability of 0.8, and mutation probability of 0.05. Through iterative optimization using this algorithm, the system can automatically calculate the optimal surgical manipulation point and instrument incident angle. The final planning results can be imported into an intraoperative navigation system or a virtual surgical simulation platform for surgical path verification and risk assessment.
[0039] During the intraoperative phase, the microscope and OCT system simultaneously acquire image streams, which are then processed through frame synchronization, denoising, illumination equalization, and standardization before being input into the deep learning recognition module. The global encoder uses ResNet-50 to extract spatial features from the microscope images, while the local encoder uses U-Net to extract layered structural features from the OCT images. The registration module calculates a cross-modal feature correlation matrix through an attention mechanism to enhance the spatial correspondence between the two images and outputs the transformation parameters of the microscope images relative to the preoperative 3D model.
[0040] The decoder generates a lesion region segmentation mask and its spatial pose prediction results. The instrument guidance module calculates displacement compensation amounts Δx, Δy, and Δz based on the deviation between the prediction results and the preoperative planned path, and transmits this information to the control system in real time. The servo control module then adjusts the end effector trajectory of the robotic arm accordingly to ensure that the instrument accurately follows the target path, achieving dynamic error compensation and real-time guidance.
[0041] In the experimental verification, 20 cases of fundus retinal drug injection imaging data were selected for testing and compared with the traditional template matching method. The results showed that the average registration error of the method of the present invention was 0.087 mm (46.2% lower than the control group), the instrument guidance error was less than 0.12 mm, and the model inference speed could reach 32 FPS, which can meet the timeliness requirements of intraoperative real-time guidance.
[0042] Furthermore, the method of this invention can dynamically update network parameters based on intraoperative feedback and utilize a lightweight knowledge distillation structure to achieve incremental learning, thereby continuously improving registration accuracy and system stability without increasing computational burden.
[0043] In summary, this invention establishes a lesion identification and guidance system that spans the entire preoperative-intraoperative process by combining multimodal image fusion, deep learning recognition, and servo-guided control. It can achieve precise lesion localization, dynamic model registration, and adaptive instrument guidance, significantly improving the safety and precision of surgery and has broad clinical application prospects.
[0044] Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Anyone skilled in the art can make various modifications and alterations without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be determined by the claims.
Claims
1. A comprehensive method for preoperative lesion identification and intraoperative guidance, characterized in that, Includes the following steps: Step 1: Acquire structural data of deep tissues in the target area using optical coherence tomography (OCT). Step 2: Deeply fuse the spatially registered microscope images with OCT 3D data to construct a 3D fusion model and a 3D coordinate system that simultaneously contains rich texture information and fine structural information; Step 3: Apply a pre-trained deep segmentation network to perform automated segmentation and spatial localization of the lesion region in the 3D fusion model; the final output of this step is the precise contour of the lesion and its detailed information in the 3D coordinate system. Step 4: Perform a comprehensive 3D reconstruction of the patient's preoperative imaging data to generate a preoperative planning model; Step 5: Real-time acquisition of microscope video stream and OCT B-scan images during the surgery; Step Six: Construct an advanced neural network model for a two-stream multimodal deep network; Step 7: Accurately map the three-dimensional deformation field output by the dual-stream multimodal deep network and the real-time lesion location onto the three-dimensional coordinate system of the three-dimensional fusion model established before surgery; Ultimately, the movement of the instruments is dynamically adjusted through a servo control system, thereby enabling real-time compensation for errors that occur during the surgical procedure.
2. The comprehensive method according to claim 1, characterized in that, Step one also includes: constructing a high-resolution three-dimensional tissue volume model using a tomographic reconstruction algorithm.
3. The comprehensive method according to claim 2, characterized in that, Step one further includes: using a joint optimization algorithm based on feature points and surface curvature to accurately register the microscope image; the joint optimization algorithm accurately maps the microscope image to the surface of the three-dimensional tissue volume model through feature matching and photometric correction, thereby ensuring the spatial consistency of the multimodal image.
4. The comprehensive method according to claim 3, characterized in that, In step two, during the deep fusion process, cross-modal features are extracted using a deep learning network.
5. The comprehensive method according to claim 4, characterized in that, The preoperative imaging data includes OCT volumetric data and structural images of the retina.
6. The comprehensive method according to claim 5, characterized in that, Step five also includes preprocessing the real-time acquired data, including frame synchronization, image denoising, data standardization, and time calibration.
7. The comprehensive method according to claim 6, characterized in that, The steps for constructing an advanced neural network model of a dual-stream multimodal deep network include: acquiring multi-frame OCT B-Scan slice sequence data (IOCT(x,z)) from the patient's fundus; using a layer extraction algorithm of a 3D U-Net network to segment the OCT data into retinal layers, obtaining the depth position functions of the main tissue interfaces within the retina, and converting them into real-space 3D coordinates: Pxyz (idx, kdy, jdz); obtaining the 3D point cloud data of the corresponding layer; constructing a 3D distance scalar field Volume(x,y,z) to the point cloud surface using a fast nearest neighbor search algorithm; using the Marching Cubes algorithm to linearly interpolate the boundary positions to solve for the intersection points to form triangular patches, and combining them to generate a 3D mesh model of the multi-layer retinal tissue; fusing the high-frequency surface motion of the microscope video with the low-frequency but scaled depth information of the intraoperative OCT to regress a dense 3D deformation field; and remapping the preoperative retinal model to the current anatomical morphology during surgery, thereby reflecting and compensating for retinal deformation in real time.
8. The comprehensive method according to claim 7, characterized in that, Step seven also includes: by calculating the registration matrix in real time and continuously comparing it with the position of the surgical instrument tip, the system can accurately identify any deviation.