AI-assisted real-time tissue recognition and cutting path planning system for microsurgery

By unifying the calibration and registration of multimodal data and real-time tissue recognition, and combining instrument kinematics and tissue cutability constraints, an executable cutting trajectory is generated, which solves the complex problem of multimodal data calibration and registration in microsurgery and improves the safety and efficiency of the surgery.

CN122376253APending Publication Date: 2026-07-14GUANGZHOU TIANDE HEALTH TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU TIANDE HEALTH TECHNOLOGY CO LTD
Filing Date
2026-04-14
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

In current microsurgical procedures, multimodal data calibration and registration are complex, and time delays and drifts are unavoidable, resulting in a heavy burden on surgeons to integrate information and making it difficult to form real-time decisions under a unified coordinate system. Furthermore, existing intelligent auxiliary systems lack path planning with constraints on instrument kinematics and tissue cutability, making it difficult to generate executable trajectories and posing a high risk of miscutting or leaving residues.

Method used

Employing modules for multimodal data acquisition, calibration and registration, tissue identification and boundary segmentation, risk assessment, and path planning, the system generates a safe distance field and outputs a cutting trajectory under kinematic and shearability constraints by unifying the calibration and registration of data such as microscopic video, fluorescence images, and OCT depth maps, combined with instrument pose and force feedback. This allows for real-time guidance using AR overlay display and robot interface.

Benefits of technology

It significantly reduces misjudgments caused by fragmented multi-source information, improves the separability of tissue layers and thin structures, reduces vascular/nerve boundary rupture and adhesion, generates executable cutting trajectories, improves surgical safety and efficiency, and reduces misleading information through risk maps and AR prompts, thereby enhancing clinical interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122376253A_ABST
    Figure CN122376253A_ABST
Patent Text Reader

Abstract

The present application relates to an AI-assisted real-time tissue recognition and cutting path planning system for microsurgery, which includes multi-modal acquisition, registration fusion, tissue segmentation, risk assessment and path planning output modules. Microscopic videos, fluorescence images, OCT depth maps, etc. are collected, and heart rate, blood pressure, instrument pose and force feedback are synchronously acquired; after camera calibration, registration and time alignment, a fusion input is formed. The tissue recognition CNN uses spatio-temporal aggregation and cross-modal attention fusion, and sets a differentiable safety distance field layer between fusion and decoding, generates a safety distance field based on instrument embedding and OCT depth gradient, as attention bias and skip-connection gating, outputs class map, boundary map, confidence and uncertainty. Based on the confidence and uncertainty, structure, mis-cut and residual risk maps are generated; under the constraints of kinematics and cuttability, trajectories containing path points, depth and speed are output, and superimposed acoustic and light warnings or robot instructions are output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical device technology, specifically to an AI-assisted real-time tissue recognition and cutting path planning system for microsurgery. Background Technology

[0002] Microsurgery is widely used in neurosurgery, ophthalmology, otolaryngology, microvascular anastomosis, and precise tumor resection. Its typical characteristics include a confined operating space, a high density of critical structures (such as blood vessels, nerves, and functional areas), and thin, easily deformable tissue layers, placing extremely high demands on cutting boundaries, cutting depth, and instrument stability. Surgeons typically rely on the magnified field of view provided by a microscope for judgment and manipulation. However, in actual surgery, bleeding, reflections, smoke, and tissue traction leading to deformation and obstruction are common. The appearance of tissues varies significantly under different lighting and angles, making the determination of boundaries between lesions and normal tissues, and between critical structures and surrounding tissues, highly subjective and carrying a high risk of miscutting or leaving residual tissue. To enhance contrast and structural information, technologies such as fluorescence imaging, narrow-band imaging, and OCT have been introduced to observe perfusion, metabolic, or chromatographic structures. Simultaneously, surgical robots and navigation systems can provide instrument pose measurement and auxiliary control. However, the aforementioned systems often suffer from complex calibration and registration between multimodal data, and unavoidable latency and drift, leading to frequent switching between different display sources by the surgeon, a heavy burden of information integration, and difficulty in forming a real-time decision-making basis under a unified coordinate system.

[0003] In recent years, intraoperative tissue recognition and semantic segmentation technologies based on convolutional neural networks have begun to be used to assist in locating lesion areas or indicating suspicious boundaries. However, existing solutions are mostly based on single microscopic images or single-frame inputs, and do not adequately consider cross-modal fusion and temporal consistency. They are easily affected by noise interference caused by transient occlusion, bleeding, and instrument occlusion, resulting in jitter and false detections. Furthermore, most methods only output category or boundary results, lacking explicit representation of model confidence and uncertainty, making it difficult to provide "interpretable cautionary hints" in high-risk areas, and easily leading to overconfident misguidance. In addition, existing intelligent assistance is mostly limited to the "recognition / hint" level, lacking a path planning mechanism that couples the recognition results with instrument kinematic constraints (accessible space, maximum curvature, speed limit, safe distance) and tissue cutability constraints (hardness difference, allowable cutting depth, abnormal contact and occlusion state). It cannot output trajectory parameters that can be directly executed or constrained by the robot, making it difficult to achieve closed-loop correction and local replanning during surgery. Furthermore, while physiological monitoring information (such as heart rate and blood pressure) and instrument force feedback are widely collected in clinical practice, they are often disconnected from visual recognition, risk alerts, and trajectory planning, failing to form a unified risk assessment framework. In summary, there is an urgent need for a system capable of unified calibration and registration of multimodal intraoperative data, real-time tissue identification and boundary segmentation, and quantified risk expression. This system should also generate executable cutting trajectories under the constraints of instrument kinematics and tissue cutability, and be output via augmented reality and robotic interfaces to improve the safety and efficiency of microsurgery. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention discloses an AI-assisted real-time tissue recognition and cutting path planning system for microsurgery, comprising modules for multimodal acquisition, registration and fusion, tissue segmentation, risk assessment, and path planning output. It acquires microscopic videos, fluorescence images, OCT depth maps, etc., and simultaneously obtains heart rate, blood pressure, instrument pose, and force feedback; these are fused into an input after camera calibration, registration, and temporal alignment. The tissue recognition CNN employs spatiotemporal aggregation and cross-modal attention fusion, and sets a differentiable safe distance field layer between fusion and decoding. A safe distance field is generated based on instrument embedding and OCT depth gradient, serving as attention bias and jump-gating, outputting a category map, boundary map, confidence level, and uncertainty. Based on the confidence level and uncertainty, it generates structure, miscutting, and residual risk maps; under kinematic and cutability constraints, it outputs a trajectory containing path points, depth, and velocity, and outputs AR-overlaid audio-visual alarms or robot commands.

[0005] AI-assisted real-time tissue recognition and cutting path planning system for microsurgery, including: Multimodal data acquisition unit: Acquires multimodal data of the same surgical field within a time period. The multimodal data includes microscopic video images and auxiliary data. The auxiliary data includes fluorescence imaging, OCT depth map, heart rate, blood pressure, surgical instrument pose data and force feedback data. Calibration and registration unit: Performs intrinsic and extrinsic parameter calibration on the camera and registers the fluorescence imaging and OCT depth map to the coordinate system of the microscopic video image; aligns multimodal data by timestamp to form a fused input vector; Tissue Recognition and Boundary Segmentation Unit: Inputs the fused input vector into the trained improved tissue recognition convolutional neural network model, and outputs pixel tissue category map, boundary map, corresponding confidence score and uncertainty; Risk map generation unit: Generates a risk map based on confidence level and model uncertainty estimation. The risk map represents structural risk, miscutting risk, and residual risk. Cutting path planning unit: Under the constraints of instrument kinematics and tissue cutability, it performs cutting path planning for the target resection area and outputs an executable trajectory containing path point sequence, cutting depth and speed parameters; Output unit: Outputs the executable trajectory to the screen as an augmented reality overlay, and outputs audio-visual prompts to indicate risk information. When connected to the surgical robot, it outputs trajectory instructions and constraint instructions to the robot controller for auxiliary control.

[0006] Preferably, the multimodal data acquisition unit includes a microscopic imaging acquisition module, a fluorescence imaging module, an OCT acquisition module, a physiological monitoring interface, and an instrument tracking and force feedback acquisition module; wherein, the microscopic video images and fluorescence imaging are synchronized with frame-level timestamps through hardware triggering, the OCT depth map is acquired according to a preset scanning cycle and aligned with the timestamps; heart rate and blood pressure are acquired by the monitor through the communication interface and aligned with the image frames according to the sampling time; the instrument posture is acquired by the electromagnetic tracker, and the force feedback is acquired by the micro-force sensor and output synchronously with the same time reference.

[0007] Preferably, the calibration and registration unit includes a calibration module and a registration and fusion module; wherein the calibration module uses a calibration plate to obtain the intrinsic parameters of the microscope camera and solves the rigid body transformation between the microscope camera, the fluorescence camera, and the OCT probe based on the extrinsic parameters; the registration and fusion module performs distortion correction on the fluorescence image and projects it onto the microscope image coordinate system according to the rigid body transformation, resamples the OCT depth map to the microscope image resolution based on the depth-pixel mapping table, and aligns and stitches the modal data according to the timestamp of a unified time base to form a fused input vector.

[0008] Preferably, the improved tissue recognition convolutional neural network model includes: a multi-branch encoder, which extracts multi-scale features from microscopic video images, fluorescence imaging images, and OCT depth maps, and encodes surgical instrument pose and force feedback to obtain instrument interaction embedding; a spatiotemporal aggregation module, used to enhance the temporal consistency of continuous frame microscopic features; a cross-modal attention fusion module, used to perform attention fusion of microscopic features with fluorescence and OCT features at multiple scale levels; a differentiable interactive safe distance field layer set between the cross-modal attention fusion module and the decoder, used to generate a pixel-level safe distance field based on the instrument interaction embedding and OCT depth gradient, and use the safe distance field as attention bias and skip connection weight to modulate the decoded features; a boundary consistency decoder, including a semantic decoding path and a boundary refinement path, wherein the boundary features output by the boundary refinement path guide the semantic decoding path to achieve boundary alignment; and a multi-head output structure, including a tissue category output head, a boundary output head, a safe distance field regression head, a confidence output head, and an uncertainty output head, used to output pixel-level tissue category map, boundary map, safe distance field, corresponding confidence, and uncertainty, respectively.

[0009] Preferably, the differentiable interactive safety distance field layer embeds the device interaction into a device proximity probability map via a multilayer perceptron and fuses it with the boundary intensity map generated by the OCT depth gradient to obtain a safety distance field, wherein the safety distance field is calculated based on the soft distance from the pixel to the nearest high-risk boundary; the safety distance field is normalized by Sigmoid and injected as an attention bias term into the cross-modal attention weight calculation, and used as a jump connection gating coefficient to weighted suppress or enhance the feature channels from the encoder to the decoder.

[0010] Preferably, the risk map generation unit fuses the pixel confidence and uncertainty output of the tissue category map to obtain a pixel risk value; wherein the risk value of the vascular structure region is set as structural risk, pixels located outside the planned cutting boundary and with a risk value higher than the threshold are marked as miscutting risk, pixels located within the target resection area and with a risk value lower than the threshold or uncertainty higher than the threshold are marked as residual risk, and a risk map containing three types of risk labels and their risk levels is output.

[0011] Preferably, the cutting path planning unit sets the instrument kinematic constraints as follows: the reachable workspace constraint, maximum curvature constraint, maximum velocity constraint, and minimum safe distance constraint between the instrument tip and the structural warning zone in the microscopic coordinate system; and sets the tissue cutability constraints as: the tissue hardness threshold obtained based on OCT depth map and force feedback estimation, the upper limit of allowable cutting depth, and the pause threshold under bleeding obstruction; the cutting path planning unit generates an initial path based on the boundary of the target resection area, and iteratively optimizes the path points using a risk map as the cost field to obtain a path point sequence; assigns a cutting depth to each path point according to the upper limit of allowable cutting depth, and performs time parameterization on the path points according to the maximum velocity constraint of the instrument, outputting an executable trajectory containing the path point sequence, cutting depth, and velocity parameters.

[0012] Preferably, the output unit converts the executable trajectory into pixel trajectory lines in the microscopic video coordinate system and displays them synchronously overlaid on the risk map, wherein the cutting order and depth level are distinguished by color; when the structural risk or miscutting risk in the risk map exceeds the threshold, a buzzer and indicator light alarm are triggered; when connected to the surgical robot, the trajectory point sequence and corresponding speed and depth parameters are encapsulated into control frames and sent to the robot controller, and at the same time, minimum safe distance and speed upper limit constraints are issued for confined auxiliary control.

[0013] Preferably, the cutting path planning unit adopts a two-level planning mechanism of global coarse planning and local real-time replanning. The global coarse planning generates a main path covering the entire area based on the boundary of the target resection area and the risk map. When the local real-time replanning detects that the instrument pose deviates from the threshold, the force feedback is abnormal, or the risk map update exceeds the threshold, it recalculates the next path segment with the instrument tip as the center within a preset local window and replaces the corresponding path segment to output a continuous executable trajectory.

[0014] Preferably, the boundary consistency decoder further includes an elongated structure connectivity constraint branch for extracting the skeleton connectivity graph of blood vessels and nerves based on fluorescence imaging and OCT depth maps, and applying the skeleton connectivity graph as a regularization constraint term to the boundary output head and the safe distance field regression head to suppress critical structure boundary breaks and improve the segmentation consistency of the critical structure neighborhood.

[0015] Compared with the prior art, the technical solution of the present invention has the following beneficial effects: (1) This invention calibrates and aligns microscopic video, fluorescence imaging, and OCT depth maps in a unified coordinate system with timestamp alignment, forming a consistent fusion input and significantly reducing misjudgments caused by fragmented multi-source information. The improved tissue recognition convolutional neural network introduces spatiotemporal aggregation and cross-modal attention fusion, which can suppress single-frame jitter and enhance the separability of tissue layers and thin structures under complex conditions such as bleeding, reflection, smoke, and instrument occlusion. Through boundary consistency decoding and slender structure connectivity constraints, it reduces the breakage and adhesion of blood vessel / nerve boundaries, and improves the segmentation accuracy and stability of key structural neighborhoods.

[0016] (2) The embedding position of the differentially interactive safe distance field layer in this application brings segmentation improvement by "controlling risks from the source": This invention sets the "differentiable interactive safe distance field layer" between the cross-modal attention fusion module and the decoder, placing it in a key channel before the "high semantic fusion features" enter "pixel-level reconstruction / boundary refinement". This layer utilizes the interactive embedding formed by the device pose and force feedback, and combines the tissue layer boundary strength extracted by OCT depth gradient to generate a pixel-level safe distance field, and directly modulates the decoding features using attention bias and skip-connection gating. Since the modulation occurs before decoding, it can suppress the over-response of high-risk neighborhoods during the feature reconstruction stage, enhance the boundary alignment and thin-layer structure resolution near key structures, thereby significantly reducing missegmentation, boundary drift and breakage in the vascular / nerve neighborhood.

[0017] (3) Under kinematic constraints such as instrument reachability, maximum curvature, speed limit, and minimum safe distance, the system generates an executable trajectory containing path point sequences, cutting depth, and speed parameters by combining tissue stiffness thresholds, allowable cutting depth limits, and occlusion pause thresholds based on OCT and force feedback estimation, thus avoiding trajectory suggestions that are "unreachable / uncuttable / unsafe." A global coarse planning and local real-time replanning mechanism is adopted to quickly update path segments when there are pose deviations, force anomalies, or significant changes in the risk map, maintaining continuous operability. The trajectory is intuitively guided through AR overlay and audio-visual cues. When connected to the robot, trajectory and constraint commands can be issued to achieve confined-domain assisted control, improving surgical efficiency and consistency.

[0018] (4) By using the safety distance field as one of the multiple outputs, and outputting it simultaneously with confidence and uncertainty, the system can convert "how close to the instrument contact area / critical structure and whether it can be safely cut" into pixel-level continuous quantities, thereby constructing three types of risk maps: structural risk, miscutting risk, and residual risk. Compared with schemes that rely solely on confidence thresholds, this invention automatically increases the modulation intensity and uncertainty level of the safety distance field when there is instrument contact, tensile deformation, or abnormal force, making the risk warning more consistent with the actual risk distribution during surgery, reducing the misleading effect of "high confidence misjudgment" by the model on the surgeon, and improving clinical interpretability and monitorability. Attached Figure Description

[0019] Figure 1 This is a flowchart of an AI-assisted real-time tissue recognition and cutting path planning system for microsurgery according to the present invention. Figure 2 This is a structural diagram of a microsurgical AI-assisted real-time tissue recognition and cutting path planning system module. Detailed Implementation

[0020] Those skilled in the art will understand that, in order to make the above-mentioned objects, features, and beneficial effects of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This application illustrates an AI-assisted real-time tissue recognition and cutting path planning system for microsurgery, comprising the following steps: Multimodal data acquisition unit: Acquires multimodal data of the same surgical field within a time period. The multimodal data includes microscopic video images and auxiliary data. The auxiliary data includes fluorescence imaging, OCT depth map, heart rate, blood pressure, surgical instrument pose data and force feedback data. In some embodiments, a multimodal data acquisition unit is set up in a neurosurgical microsurgical tumor resection surgery, which includes a microscopic imaging acquisition module, a fluorescence imaging module, an OCT acquisition module, a physiological monitoring interface, and an instrument pose and force feedback acquisition module. Each module achieves synchronous acquisition through a unified clock and trigger line, and outputs a timestamped data stream to the fusion computing terminal.

[0021] For microscopic video image acquisition, the main optical path of the microscope is connected to a 4K microscope camera (which can be a CMOS camera). The camera is connected to an edge computing terminal via an acquisition card (such as a PCIe acquisition card). The camera outputs raw frames at 30 fps or 60 fps. Single-frame data includes: frame number FrameID, exposure parameters, resolution, and hardware timestamp t_cam. The trigger input of the microscope camera is connected to a unified trigger controller, which generates a timestamp with each frame trigger.

[0022] Fluorescence imaging acquisition: The fluorescence imaging module shares the field of view with the microscope, and a beam splitter or spectroscope is used to guide the fluorescence channel to an independent fluorescence camera. The fluorescence camera is also connected to the acquisition card / synchronization interface, and acquisition is triggered by a unified trigger controller, ensuring a one-to-one correspondence or a fixed doubling / divisional relationship between fluorescence frames and microscopic frames (e.g., every two microscopic frames correspond to one fluorescence frame). The fluorescence frame carries a timestamp t_fluor, excitation intensity, and gain information for subsequent illumination normalization.

[0023] OCT depth map acquisition: The OCT probe is connected to the scanning module via a side-path fiber. The OCT controller performs B-scan or volumetric scan at a preset scan cycle. Each scan outputs an OCT depth map (depth-pixel array) and scan parameters (scan line number, step size, depth calibration coefficient), along with a hardware timestamp t_oct. To ensure alignment with the image frame, the trigger controller outputs a synchronization pulse to the OCT controller. Upon receiving the pulse, the OCT begins a scan or locks the scan cycle to the trigger clock.

[0024] Heart rate and blood pressure are acquired. The physiological monitoring interface reads heart rate (HR) and blood pressure (BP, systolic / diastolic) in real time from the monitor via wired communication (e.g., USB / serial / Ethernet), with a sampling frequency of, for example, 1 Hz to 5 Hz. Each physiological data record includes: HR, BP_sys, BP_dia, and a sampling timestamp t_vital (converted by a unified clock or timestamped by the interface driver). The acquisition unit buffers the physiological data and establishes a correspondence with the timestamp of the most recent image frame for subsequent risk threshold modulation or alert strategies.

[0025] Surgical instrument pose and force feedback acquisition: Instrument pose is acquired by an electromagnetic tracker: an electromagnetic sensor is fixed to the instrument handle, and an electromagnetic transmitter is fixed beside the operating table and calibrated in the field; the tracker outputs the instrument pose (position x, y, z and attitude quaternion q or Euler angles) at 100 Hz, with a timestamp t_pose. Force feedback is acquired by a micro-force sensor mounted on the tip or proximal end of the instrument, which can be a triaxial or hexaaxial force / torque sensor, outputting Fx, Fy, Fz (and Mx, My, Mz) at 200 Hz, with a timestamp t_force. The acquisition unit synchronously outputs pose and force data based on the same time reference, and retains the original sampling rate for subsequent contact detection and tissue stiffness estimation. 6) Unified synchronization and data output: a unified trigger controller provides the master clock and trigger signal: the microscope camera and fluorescence camera receive frame triggers; the OCT receives scan triggers or phase-locked signals; pose / force / physiological data are timestamped by the acquisition unit software using the same clock source (e.g., PTP or hardware counter). The acquisition unit ultimately outputs a set of multimodal record packets indexed by timestamps for subsequent calibration, registration, and fusion input vector construction.

[0026] Calibration and registration unit: Performs intrinsic and extrinsic parameter calibration on the camera and registers the fluorescence imaging and OCT depth map to the coordinate system of the microscopic video image; aligns multimodal data by timestamp to form a fused input vector; In some embodiments, based on multimodal data acquisition, a calibration and registration unit is set up, which consists of a calibration module, a registration and fusion module and a time alignment module. It runs on an edge computing terminal or workstation and outputs a fused input vector for organizing and recognizing convolutional neural networks.

[0027] Distortion calibration is performed within the microscope camera. Preoperatively, with the microscope working distance and magnification fixed, a calibration plate with a fine grid and dot array is placed at the center of the microscopic field of view, and microscopic image frames are acquired from multiple angles and positions. The calibration module estimates the focal length, principal point, and distortion parameters of the microscope camera based on the known geometry of the calibration plate, and generates a distortion correction mapping table for the microscope camera. If the magnification change exceeds a threshold during the procedure, the system prompts for recalibration or calls up the calibration parameter set for the corresponding magnification.

[0028] The calibration of extrinsic parameters between the fluorescence camera and the microscopic camera involves two channels sharing the same field of view but with different imaging pathways. Preoperatively, a dual-mode calibration target with both visible light and fluorescence response characteristics is placed in the field of view, and microscopic and fluorescence images are acquired separately. The calibration module first performs distortion correction on both images, then detects corresponding marker points in both images and solves the rigid transformation relationship from the fluorescence camera coordinate system to the microscopic camera coordinate system, obtaining a set of stable cross-camera extrinsic parameters. These extrinsic parameters are used to map the fluorescence image to the microscopic image coordinate system.

[0029] Spatial calibration and mapping between OCT depth maps and microscopic images: The OCT probe and the microscopic field of view have a fixed installation relationship. Preoperatively, a calibration object (e.g., a stepped block or microstructure calibration sheet) with a stepped height or a known height distribution is placed within the microscopic field of view, and microscopic images and OCT scan data are acquired simultaneously. The registration and fusion module establishes a mapping table of "OCT scan coordinates - microscopic pixel coordinates" based on the OCT scan parameters and the geometric relationship of the calibration object, and generates a lookup table to resample the OCT depth map to the resolution and field of view of the microscopic image. During the procedure, each frame or scan of the OCT depth map is projected / resampled to the microscopic coordinate system according to this mapping table, resulting in a depth layer of the same scale and field of view as the microscopic image.

[0030] Real-time registration and quality self-checking: The intraoperative registration fusion module first performs distortion correction on the input microscopic and fluorescence frames, then maps the fluorescence frame to the microscopic coordinate system according to extrinsic parameters and clips it to the microscopic field of view; the OCT depth map is resampled to the same resolution as the microscopic frame according to the mapping table. To avoid registration failure caused by instrument occlusion or drift, the system periodically calculates the consistency of key points or mutual information between fluorescence and microscopic images as a registration quality indicator; when the indicator is below the threshold, it enters a conservative mode: reducing the fluorescence weight or prompting recalibration, and recording the "registration confidence mark" in the fusion input.

[0031] Multimodal timestamp alignment is implemented, with the time alignment module using the timestamps of the microscopic video frames as the primary time axis. For each frame of microscopic image, the system selects the fluorescence frame and OCT depth map with the closest timestamp from the buffer. If the fluorescence or OCT sampling frequency is lower than the microscopic frame rate, nearest neighbor matching or short window interpolation is used to generate aligned data, and the alignment error is marked. Low-frequency physiological data such as heart rate and blood pressure are aligned with the nearest image frame according to the sampling time and kept updated at zero order. Instrument pose and force feedback are interpolated or the nearest value is taken at the corresponding frame timestamp based on high-frequency sampling to obtain the instrument state vector corresponding to that frame.

[0032] The fusion input vector construction and output process involves organizing data from the same time point into a fusion input vector (or fusion input feature package) after spatial registration and temporal alignment. This fusion input vector includes at least: a microscopic image, a registered fluorescence image, a resampled OCT depth map, and numerical features such as device pose, force feedback, heart rate, and blood pressure aligned to that frame. It may also include registration confidence markers and alignment error markers. The fusion input vector is then fed into the tissue identification and boundary segmentation unit for subsequent pixel-level tissue category maps, boundary maps, confidence scores, and uncertainty assessments.

[0033] Tissue Recognition and Boundary Segmentation Unit: Inputs the fused input vector into the trained improved tissue recognition convolutional neural network model, and outputs pixel tissue category map, boundary map, corresponding confidence score and uncertainty; In some embodiments, after completing multimodal acquisition, calibration and registration and forming a fusion input vector, the tissue identification and boundary segmentation unit runs in real-time at the frame level on an edge computing terminal (GPU / AI accelerator) to output pixel-level tissue category map, boundary map, confidence and uncertainty for the surgical field of view at each moment, and simultaneously outputs a safe distance field for subsequent risk map generation and path planning.

[0034] 1) Input and Output: For any timestamp t, the input is a fused input vector, which includes at least: a microscopic video image frame, a fluorescence imaging image registered to the microscopic coordinate system, an OCT depth map resampled to the same resolution, and surgical instrument pose data and force feedback data time-aligned with that frame. The network output includes: pixel-level tissue category map, pixel-level boundary map, pixel-level safe distance field, pixel-level confidence level, and pixel-level uncertainty level.

[0035] 2) Improve the overall structure of the tissue recognition convolutional neural network. The improved model in this embodiment is cascaded in the following order: multi-branch encoder, spatiotemporal aggregation module, cross-modal attention fusion module, (between fusion and decoding) differentiable interactive safe distance field layer, boundary consistency decoder, and multi-head output structure.

[0036] The multi-branch encoder (multi-scale feature extraction and instrument interaction embedding) includes the following branches: Microscopic video image coding branch: Extracts multi-scale features (low-level texture / edges, mid-level morphology, high-level semantics) from microscopic images, forming a multi-level feature pyramid. Fluorescence imaging coding branch: Extracts enhanced regions and gradient change features from fluorescence images, forming multi-scale features aligned with the microscopic branch. OCT depth map coding branch: Extracts hierarchical structure and interface change features from OCT depth maps, forming multi-scale depth features.

[0037] Instrument interaction encoding: Input instrument pose and force feedback into a lightweight encoding network (e.g., several fully connected / one-dimensional convolutional layers), and output instrument interaction embedding to express the state of instrument approach, contact intensity, motion changes and potential deformation risk.

[0038] The spatiotemporal aggregation module (enhancing the temporal consistency of microscopic features) introduces a spatiotemporal aggregation module at the mid-to-high-level features of the microscopic branch to enhance the temporal consistency of microscopic features across several consecutive frames: when transient anomalies such as bleeding, reflection, smoke, or instrument occlusion occur, this module suppresses single-frame jitter, making subsequent segmentation boundaries more continuous and stable in the temporal dimension.

[0039] The cross-modal attention fusion module (multi-scale attention fusion) uses microscopic features as the backbone at each scale level, fusing fluorescence and OCT features through an attention mechanism to obtain a fused feature pyramid. This fusion method is "scale-corresponding and spatially adaptive," used to introduce contrast and hierarchical information provided by fluorescence and OCT when microscopic texture is weak or severely occluded, thereby improving tissue differentiation capabilities.

[0040] The implementation and location of the differentiable interactive safe distance field layer are as follows: This layer is positioned between the cross-modal attention fusion module and the boundary consistency decoder, specifically in the critical channel where "fusion semantics have been formed" but "pixel-level decoding has not yet occurred," giving it a direct modulation effect on the decoding stage. Specifically, the implementation is as follows: Device proximity probability map generation: Device interaction is embedded into the input multilayer perceptron, outputting a device proximity probability map at the same scale as the microscopic field of view, used to characterize the device tip neighborhood and its affected area. Boundary intensity map acquisition: A boundary intensity map is generated from the OCT depth gradient, used to characterize the interlayer interfaces of tissues and the locations of potential high-risk boundaries. Safe distance field fusion generation: The device proximity probability map and the boundary intensity map are fused to obtain a pixel-level safe distance field, which represents the soft distance from the pixel to the nearest high-risk boundary.

[0041] The subsequent network is modulated in two ways: Attention bias: The normalized safe distance field is injected into the cross-modal attention weight calculation as an attention bias term, so that the region near the high-risk boundary is more cautious in attention allocation and reduces overconfident activation; Skip gating: The safe distance field is used as the skip gating coefficient to weight and suppress or enhance the skip feature channel from encoder to decoder, so that the key structure neighborhood highlights the boundary-related features and suppresses the background interference features.

[0042] The boundary consistency decoder comprises two parallel paths: a semantic decoding path that progressively upsamples and fuses skip connections to recover pixel-level semantics and outputs tissue category-related decoded features; and a boundary refinement path that focuses on recovering fine-grained features of edges / thin structures and outputs refined boundary features. These boundary features then guide the semantic decoding path in reverse (e.g., through feature fusion or gating) to align semantics with boundaries, reducing the situation of "correct category but coarse / drifted boundaries." The boundary consistency decoder may also include a slender structure connectivity constraint branch, using fluorescence and OCT-based extraction of vascular and / or neural skeletal connectivity maps as regularization constraints to suppress boundary breaks.

[0043] The multi-head output structure sets up multiple outputs on the shared decoding features: Tissue Category Output Head: outputs pixel-level tissue category map; Boundary Output Head: outputs pixel-level boundary map; Safe Distance Field Regression Head: outputs pixel-level safe distance field; Confidence Output Head: outputs pixel-level confidence; Uncertainty Output Head: outputs pixel-level uncertainty, used to characterize low-confidence regions of the model under conditions such as occlusion, reflection, and bleed.

[0044] Compared to conventional fusion segmentation networks without a "differentiable interactive safety distance field layer," the key improvements in this embodiment are reflected in the following aspects: Safer neighborhoods for critical structures: Because the safety distance field layer is located between fusion and decoding, it can perform risk-sensitive modulation on features before pixel reconstruction, resulting in clearer boundaries and fewer adhesions and cross-boundary missegments in the neighborhoods of critical structures such as blood vessels / nerves; More robust to device interactions: Device pose and force feedback influence the safety distance field through device interaction embedding, enabling the network to automatically increase caution when the device approaches / contacts or experiences abnormal forces, reducing the high-confidence output of mis-cut regions; More usable risk representation: The network outputs not only categories / boundaries but also safety distance fields, confidence levels, and uncertainties, allowing subsequent risk map generation to more accurately distinguish between structural risks, mis-cut risks, and residual risks, and providing a more reliable cost field and safety distance basis for path planning.

[0045] Risk map generation unit: Generates a risk map based on confidence level and model uncertainty estimation. The risk map represents structural risk, miscutting risk, and residual risk. In some embodiments, the tissue identification and boundary segmentation unit outputs a pixel-level tissue category map, boundary map, confidence map, and uncertainty map for each frame t (and may also output a safe distance field). The risk map generation unit is deployed on the same edge computing terminal, receives the above outputs in real time, and generates a risk map containing three types of risk labels and risk levels, which is used by the output unit for prompting and the cutting path planning unit as a cost field.

[0046] For each frame t, the risk map generation unit receives the following input data: (1) Pixel-level tissue category map: including at least key structural categories such as target excised tissue, normal tissue, and vascular structures; (2) Pixel-level boundary map: reflecting the location and edge strength of tissue boundaries; (3) Pixel-level confidence map: reflecting the confidence level of the model in the classification results; (4) Pixel-level uncertainty map: reflecting the low confidence areas of the model under conditions such as occlusion, reflection, hemorrhage, and deformation; (5) Optional input: Pixel-level safe distance field (if output by the network), used to enhance the continuous expression of risk in the neighborhood of key structures.

[0047] Pixel risk value construction (confidence-uncertainty fusion): The risk map generation unit first performs unified scale normalization on the confidence and uncertainty, and generates pixel risk values ​​according to preset fusion rules: when the confidence decreases or the uncertainty increases, the pixel risk value increases; if both "low confidence and high uncertainty" are met, the pixel is marked as a "pixel of key concern". To improve stability, the risk map generation unit performs small-range temporal smoothing on the pixel risk values: a sliding window is used to count the same pixel position over several consecutive frames to suppress risk flickering caused by single-frame noise.

[0048] The generation of critical structural risks involves extracting vascular structures from the tissue category map as critical structural regions. These regions are then expanded by a preset safety distance to create a structural warning zone. Within this warning zone, pixel risk values ​​are increased to a high-risk level, forming a structural risk layer. If the input includes a safety distance field, the structural risk level is further increased in areas with smaller safety distances (indicating proximity to high-risk boundaries), expanding the structural risk from a "binary warning zone" to a "continuous risk gradient," facilitating subsequent path planning to avoid critical structures.

[0049] The system generates a risk map for erroneous cutting. After obtaining the planned cutting boundary or the boundary of the target cut area, the risk map generation unit considers the area outside the boundary as an area that should not be cut. When the risk value of a pixel outside the boundary exceeds a first threshold (e.g., indicating a conflict between the identification result and the boundary position or proximity to a critical structure), the pixel is marked as an erroneous cutting risk. To reduce false alarms, the system requires that erroneous cutting risks only escalate to a higher level of erroneous cutting risk and trigger an audio-visual warning after they have appeared continuously in adjacent frames for more than a preset number of frames or formed continuous patches in space.

[0050] Residual risk generation: The risk map generation unit identifies pixels within the target excision area as areas to be excised. When a pixel's risk value is below a second threshold or its uncertainty is above a third threshold, the model is considered to lack certainty regarding whether the target tissue has been adequately covered, or there is a situation where "non-target tissue is still present within the boundary / target tissue has not been fully identified," and this pixel is marked as residual risk. For residual risks, the system can further combine the boundary map: if the residual risk is close to the boundary and the boundary strength is weak, the residual risk level is increased, prompting the surgeon to review or triggering the planning unit to perform local replanning.

[0051] The risk map output and interface: The risk map generation unit outputs a risk map with the same resolution as the microscopic video. The risk map includes at least: a structural risk layer (risk levels of critical structural areas and their warning zones); a miscutting risk layer (high-risk pixels and their levels outside the planned cutting boundary); and a residual risk layer (low-confidence / high-uncertainty pixels and their levels within the target removal area). A corresponding threshold trigger flag and maximum risk level are provided for each type of risk. This risk... Figure 1 On the one hand, it sends data to the output unit for AR overlay display and alarm triggering; on the other hand, it serves as the cost field input for the cutting path planning unit, enabling the planned path to automatically avoid structural risks and areas of accidental cutting, and to provide coverage compensation or suggest review for residual risk areas.

[0052] Cutting path planning unit: Under the constraints of instrument kinematics and tissue cutability, it performs cutting path planning for the target resection area and outputs an executable trajectory containing path point sequence, cutting depth and speed parameters; In some embodiments, the system has obtained the target resection area boundary, risk map (including structural risk, miscutting risk, and residual risk), and OCT depth map aligned with the microscopic coordinate system, instrument pose, and force feedback. The cutting path planning unit runs on an edge computing terminal in frame-level or near real-time mode, outputs an executable trajectory, and performs local replanning when necessary.

[0053] The input and coordinates are unified, and the cutting path planning unit receives: (1) the boundary of the target resection area (from preoperative annotation or intraoperative identification result projection); (2) the risk map and structural warning zone (from the risk map generation unit); (3) the OCT depth map and the upper limit of the allowable cutting depth (determined by the depth map and surgical parameters); (4) the instrument pose and force feedback (used for current state estimation and cutability judgment); and (5) the safety distance field (used for smoother safety distance cost). All the above data are expressed in the microscopic video coordinate system, and the instrument tip pose is mapped to the same coordinate system through extrinsic parameters to ensure that the path points can be directly connected with the robot control.

[0054] The determination of instrument kinematic constraints involves the planning unit pre-reading the instrument / robot kinematic parameters and surgical settings to form the following constraints: Accessible workspace constraint: Based on instrument length, joint range of motion, entry point (e.g., fixed puncture point), and microscopic field of view, the accessible area of ​​the instrument tip in the microscopic coordinate system is constructed; for handheld instruments, the accessible area is limited by the surgeon's allowed operating range and the microscope's field of view boundary. Maximum curvature constraint: Limits the turning amplitude of adjacent path segments to ensure a smooth trajectory that conforms to the habits of delicate microscopic operations. Maximum speed constraint: Limits the upper limit of trajectory execution speed; if connected to a robot, acceleration / jerk can be further limited to ensure stability. Minimum safe distance constraint: Using the structural warning zone boundary as a reference, a preset safe distance must be maintained between the trajectory point and the critical structural risk area, and this safe distance can be dynamically increased according to the risk level.

[0055] The determination of tissue cutability constraints involves a planning unit that combines OCT and force feedback to estimate the cutability status in real time: Tissue hardness threshold: The relationship between force feedback and instrument pose changes is used to determine whether "high resistance / slippage / jamming" occurs, corresponding to the highly reflective hard tissue areas displayed on the OCT. If the threshold is exceeded, the tissue is deemed uncut or requires reduced speed / layered cutting. Maximum allowable cutting depth: The maximum cutting depth is set based on the OCT depth map to identify tissue interlayer interfaces or key layers (e.g., blood vessel walls, layers near nerve bundles); a smaller maximum limit is used for areas near key structures. Occlusion pause threshold: When the risk map shows a continuous increase in uncertainty or bleeding / smoke occlusion causing decreased visibility, a pause is triggered or a conservative mode is entered, outputting only the next step suggestion without outputting continuous trajectory segments.

[0056] The initial path generation (global coarse planning) planning unit first generates a covering initial path within the target cut area. Specifically, it generates a main path skeleton based on the boundary of the target cut area (e.g., generating a loop path along the inner side of the boundary, or generating a parallel scan line covering the area); it trims the initial path so that it falls entirely within the reachable workspace and automatically avoids structural warning zones; if the target area is large, it sorts the path segments in the order of "from outside to inside" or "from low risk to high risk" to form a main path sequence.

[0057] The risk-map-based iterative optimization (safety cost field optimization) constructs a cost field using the risk map (and an optional safety distance field) after obtaining the initial path. High costs are imposed on path points crossing structural risks or accidental cutting risks, pushing the path away from high-risk areas during optimization. Coverage incentives are applied to residual risk areas, ensuring the path more fully covers suspected residual areas while meeting safety constraints. Smoothing constraints are applied to curvature changes and path length to ensure the optimization result meets maximum curvature requirements. The final path point sequence is obtained after optimization, and the corresponding risk level and safety distance margin are recorded at each path point for subsequent velocity modulation.

[0058] Cutting Depth and Speed ​​Parameter Assignment (Trajectory Executability) Depth Assignment: For each path point, the tissue layer information at that location is obtained by looking up the OCT depth map, and the cutting depth is set in conjunction with the upper limit of the allowable cutting depth. In areas close to structural warning zones or where the safety distance is insufficient, the cutting depth is automatically reduced and layered cutting is recommended. Speed ​​Assignment and Time Parameterization: Using the maximum speed constraint as the upper limit, speed modulation is performed based on the risk level and cutability status: the speed is reduced in high-risk or high-uncertainty areas; further speed reduction or pausing occurs when force feedback is abnormal or the hardness threshold is close; subsequently, the path point sequence is time parameterized, outputting the expected arrival time or speed command for each path point, forming an executable trajectory.

[0059] During the local real-time replanning process, the planning unit continuously monitors events such as instrument pose deviation from the threshold, abnormal force feedback, risk map update exceeding the threshold, and occlusion pause threshold triggering. Once triggered, the system establishes a local window centered on the instrument tip and performs rapid replanning only on the "next path segment": re-trimming the reachable area, updating the cost field, and replacing the corresponding path segment to ensure trajectory continuity and real-time availability.

[0060] The output data format of the executable trajectory output by the cutting path planning unit includes at least: a sequence of path points (a sequence of points in a microscopic coordinate system or a robot working coordinate system); the cutting depth parameter for each path point; the velocity parameter or time parameter for each path point; and optional constraint information (minimum safe distance, speed limit, pause flag), which is used for the output unit display and the robot controller to perform limited / speed-limited execution.

[0061] Output unit: Outputs the executable trajectory to the screen as an augmented reality overlay, and outputs audio-visual prompts to indicate risk information. When connected to the surgical robot, it outputs trajectory instructions and constraint instructions to the robot controller for auxiliary control.

[0062] In some embodiments, the cutting path planning unit outputs an executable trajectory (including a sequence of path points, cutting depth, and speed parameters), and the risk map generation unit outputs a risk map of structural risk, accidental cutting risk, and residual risk. The output unit is connected to a microscopic display terminal, a prompting terminal, and a robot controller to present "how to cut, where the danger is, and how to avoid it" in an intuitive way and / or to assist in control execution. The hardware components and connection methods include the following output units: (1) Image overlay rendering module: running on the edge computing terminal GPU or workstation, receiving the microscopic video stream and overlaying trajectory and risk information on it; (2) Display terminal: including one of the following: external microscope display, surgeon's main display screen, or eyepiece display (HUD / eyepiece overlay module); (3) Audio-visual prompt terminal: including a buzzer / speaker and indicator light bar (red, yellow, and green) or foot pedal prompt; (4) Robot communication interface module: connected to the surgical robot controller via Ethernet / serial port / real-time bus, supporting control frame transmission and status readback; (5) Interactive input module: foot switch, touch screen, or handle button, used for surgeon confirmation, pause, switching layers, and accepting / rejecting trajectory suggestions. The above modules are connected through a unified time base and buffer queue to ensure that the display frame, risk map, and trajectory information are presented synchronously at the same timestamp.

[0063] The AR overlay display process (trajectory to pixel trajectory line) output unit uses the microscopic video frame as the base image and performs overlay display according to the following steps: S21: Read the current frame microscopic video image and obtain the executable trajectory and risk map under the corresponding timestamp; S22: Convert the path point sequence of the executable trajectory from the microscopic coordinate system to the screen pixel coordinate system (if the trajectory is in the robot coordinate system, it is first mapped to the microscopic coordinate system through the calibrated coordinate transformation, and then mapped to the pixel coordinate system); S23: Connect the path point sequence to generate a pixel trajectory line and overlay it on the microscopic video screen; S24: Map the cutting depth to the "layer marker" of the trajectory, for example, with different line types / virtual markers. Solid / line width represents different depth levels; the cutting sequence is mapped to trajectory direction arrows or serial number markers; the speed parameter is mapped to the display intensity or flashing frequency of the trajectory segment to indicate "slow cut / fast cut" suggestions; S25: the risk map is displayed in a semi-transparent thermal overlay form, where structural risk areas are displayed as warning bands, erroneous cutting risk areas are displayed with highlighted outlines, and residual risk areas are indicated by patches; and the overlay transparency and coverage are updated when the risk level changes; S26: the real-time position of the instrument tip (obtained by pose tracking) is overlaid on the screen as a cursor, providing the relationship between "current position - suggested trajectory - risk area" to facilitate quick decision-making by the operator.

[0064] To avoid screen overload, the output unit provides "display mode switching": for example, three modes: display only trajectory, display only risk, and trajectory + risk overlay; and supports switching the display on / off and transparency according to the surgeon's preference. The audio-visual prompt strategy (risk threshold triggering and graded prompts) continuously monitors the relationship between the risk map and the instrument tip position, and executes graded prompts: Level 1 prompt (early warning): When the instrument tip enters the outer edge buffer zone of the structural warning band or the next segment of the trajectory approaches a high-risk area, the indicator light changes from green to yellow, and a short prompt sound is played; Level 2 prompt (strong warning): When the instrument tip enters a high-level area of ​​structural risk / miscutting risk, or the risk value of the corresponding area in the risk map exceeds the threshold for several frames, the indicator light turns red and triggers a continuous buzzer or voice prompt (such as "approaching a blood vessel / miscutting risk"); Pause prompt: When the occlusion pause threshold is triggered (such as a large increase in uncertainty or severe bleeding occlusion), the output unit prompts "It is recommended to pause / clear the field of view" and can send deceleration or hold commands to the robot interface. The threshold and duration of frame rate can be configured by the surgical procedure: for example, in neurosurgery, the structural risk threshold can be set more conservatively, while in minimally invasive ophthalmic surgery, the speed threshold can be made more sensitive.

[0065] When connected to the surgical robot, the output unit does not directly "cut in place of the surgeon" but provides auxiliary control, typically including three types: "domain limitation, speed limitation, and follow-up prompts." Control frames are encapsulated and transmitted. The output unit encapsulates the trajectory point sequence into control frames at a fixed frequency. Each frame contains at least: the target path point (position and attitude, or only position); the expected velocity and cutting depth parameters corresponding to that point; the risk level label for that point (for secondary verification on the controller side); and a timestamp and sequence number (for packet loss detection and playback synchronization). The control frames are sent to the robot controller via the robot communication interface and wait for the controller to read back "receive confirmation / execution status / current pose."

[0066] Constraint commands are issued (domain / speed limits), and the output unit simultaneously issues constraint commands, typically including: minimum safe distance constraints: setting prohibited entry areas or soft constraint areas with the boundary of the structural warning zone as a reference; the robot controller implements this as a virtual wall / virtual clamp, automatically increasing damping or restricting entry when approaching the boundary; speed limit constraints: dynamically setting the speed limit according to the risk level, automatically reducing speed when entering a high-risk area; pause / hold constraints: when the occlusion pause threshold is triggered or the force feedback is abnormal, the robot is instructed to enter hold mode or retreat to a safe point. Human-robot collaborative control mode: the output unit supports at least two collaborative modes: prompt mode: the robot does not automatically execute the trajectory, but only displays the trajectory and provides audio and visual prompts; auxiliary restriction mode: the robot allows manual operation by the operator, but the controller limits the domain / speed according to the constraint commands to prevent accidental entry into high-risk areas; semi-automatic following mode: after the operator confirms, the robot follows the trajectory point sequence at a low speed, and the operator can interrupt it at any time via a foot switch.

[0067] Synchronization, Delay, and Safety Mechanisms: To ensure safety, the output unit employs the following mechanisms: Timestamp Consistency Check: If the difference between the timestamp of the trajectory and risk map and the current video frame exceeds a threshold, the overlay update stops and a "data delay" message is displayed; Confidence / Uncertainty Linked Degradation: When uncertainty increases significantly, the system automatically reduces the weight of the trajectory display, highlighting only structural warning strips and risk warnings to avoid misleading users with erroneous trajectories; Emergency Interruption: The surgeon can disable the robot's auxiliary output with a single press of a foot pedal or button, retaining only the displayed prompts; Log Recording: The system records the display status, alarm triggering, and control frame transmission sequences for postoperative review.

[0068] Through the aforementioned output units, the surgeon can simultaneously view the suggested trajectory, risk distribution, and current instrument position within the same microscopic view, reducing the need for switching between multiple screens; audio-visual prompts provide immediate risk feedback; when connected to a robot, trajectory and constraint commands can achieve limited-area / limited-speed auxiliary control, thereby reducing the probability of accidental cutting in the vicinity of critical structures and improving operational consistency and efficiency.

[0069] The microsurgical AI-assisted real-time tissue recognition and cutting path planning system of this invention mainly consists of a multi-source data acquisition terminal, an edge computing terminal, and an output and execution terminal, and achieves closed-loop interaction of data and control commands through wired / wireless communication interfaces. Figure 2As shown, the acquisition end includes a microscope camera for acquiring microscope video images. The microscope camera is connected to the edge computing terminal via a video acquisition interface, which can be any one or more of HDMI / SDI / USB3.0 / Gigabit Ethernet, and the acquisition card can perform frame capture and timestamp writing. A fluorescence camera and an OCT probe are used to acquire fluorescence imaging images and OCT depth maps. The fluorescence camera can be connected to the edge computing terminal via USB3.0 / Ethernet / CameraLink interfaces, and the OCT probe is connected to the edge computing terminal via its controller, with connection methods including Ethernet / USB / PCIe, and outputs a depth data stream with scan sequence number and timestamp. A sensor module is used to measure heart rate, blood pressure, device pose, and force feedback data. Heart rate / blood pressure can be output from the monitor to the edge computing terminal via serial port / USB / Ethernet; device pose can be output from an electromagnetic / optical tracker to the edge computing terminal; force feedback is provided by a micro-force sensor via a signal conditioning module and an A / D conversion module to the control processing unit or the edge computing terminal. The edge computing terminal includes: Processor: at least a GPU and a PLC (or MCU / real-time controller). The GPU is used for highly parallel computing tasks such as organizing and recognizing convolutional neural network inference, risk map generation, and trajectory calculation; the PLC is used for real-time control of acquisition triggering / synchronization management, real-time safety logic, peripheral linkage, and robot interface. The GPU and PLC can communicate via PCIe / Ethernet / shared memory / industrial bus to achieve the division of labor and cooperation between "AI computing and real-time control". Memory (ROM, RAM): ROM is used to store programs, model parameters, and system configuration; RAM is used to cache multimodal data frames, aligned fused input vectors, risk maps, trajectory queues, and other online data. Communication interfaces (Bluetooth, Wi-Fi, and wired interfaces): used to connect the display terminal, audio-visual prompt terminal, and surgical robot controller; the wired interface can be Ethernet / USB / RS485 / CAN / EtherCAT, etc., for low-latency control frame transmission; Bluetooth / Wi-Fi is used for non-critical links (such as status monitoring, log upload, and remote maintenance). Encoding / decoding module: Used for encoding and decoding microscopic and fluorescence videos, bitstream buffering, frame synchronization and format conversion (e.g., YUV / RGB conversion, resolution scaling), and providing a unified video frame format for AR overlay rendering.

[0070] The output and execution terminals include: a display terminal: a screen or eyepiece HUD for displaying microscopic video and overlaying the planned trajectory and risk map; the display terminal connects to the edge computing terminal via HDMI / DP / Ethernet, etc.; an audio-visual prompt terminal: a buzzer / speaker and indicator lights for outputting graded prompts for structural risks, accidental cutting risks, and residual risks; the audio-visual prompt terminal connects to a PLC or communication interface and can be controlled via GPIO / serial port / Ethernet port; and a surgical robot controller for receiving trajectory and constraint commands output from the edge computing terminal and transmitting robot status / execution feedback; it connects to the edge computing terminal via low-latency wired communication (such as Ethernet / EtherCAT / CAN).

[0071] The system's operation can be summarized in the following steps: synchronous acquisition and input buffering of microscopic video frames from the microscope camera; output of fluorescence frames from the fluorescence camera; output of depth maps from the OCT controller; and output of heart rate, blood pressure, pose, and force feedback according to their respective sampling periods. After the above data enters the edge computing terminal, it is first written into the buffer queue by the encoding / decoding module and the acquisition interface module, and the timestamp or frame number is recorded uniformly.

[0072] Calibration, registration, and time alignment (completed within the terminal): The edge computing terminal calls the calibration parameters to correct the distortion of the microscope camera and maps / registers the fluorescence image and OCT depth map to the microscope video coordinate system. At the same time, using the microscope frame timestamp as the main time axis, the fluorescence frame, OCT depth map, heart rate and blood pressure, pose, and force feedback are aligned to form a fused input vector at the same moment, which is stored in RAM for subsequent calculations.

[0073] AI inference and risk assessment (primarily GPU-based): The GPU reads the fused input vector, executes improved tissue identification convolutional neural network inference, and outputs tissue category map, boundary map, confidence level, and uncertainty (and can also output a safe distance field); then it generates risk maps of structural risk, accidental cutting risk, and residual risk, and passes the risk level and trigger flag to the PLC for real-time alarm control.

[0074] The trajectory generation and constraint encapsulation cutting path planning module generates an executable trajectory (path point sequence, cutting depth, and velocity parameters) based on the instrument kinematic constraints and tissue cutability constraints, combined with a risk map. The trajectory is encapsulated as a trajectory queue, and constraint information such as minimum safe distance and maximum velocity limit is generated simultaneously.

[0075] The AR overlay display and audio-visual prompt output terminal receives and displays "microscopic video frames + overlay layers (trajectory lines / risk hot zones / warning tapes / prompt text)"; when the risk exceeds the threshold, the PLC controls the audio-visual prompt terminal to output a buzzer / voice alarm and indicator light alarm, enabling the surgeon to perceive the risk in real time.

[0076] Robot Interface Output (Optional) and Closed-Loop Feedback: When the system is connected to the surgical robot, the edge computing terminal sends trajectory command frames (including points, speed, and depth) and constraint commands (domain limit / speed limit / safe distance) to the robot controller through the communication interface. The robot controller sends back the execution status and feedback data, and the edge computing terminal updates the trajectory queue or triggers local replanning accordingly, and updates the display and prompt information synchronously, forming a closed-loop collaboration.

[0077] In some embodiments, the multimodal data acquisition unit includes a microscopic imaging acquisition module, a fluorescence imaging module, an OCT acquisition module, a physiological monitoring interface, and an instrument tracking and force feedback acquisition module; wherein, the microscopic video images and fluorescence imaging are synchronized with frame-level timestamps through hardware triggering, the OCT depth map is acquired according to a preset scanning cycle and aligned with the timestamps; heart rate and blood pressure are acquired by the monitor through the communication interface and aligned with the image frames according to the sampling time; the instrument pose is acquired by the electromagnetic tracker, and the force feedback is acquired by the micro-force sensor and output synchronously with the same time reference.

[0078] In some embodiments, the calibration and registration unit includes a calibration module and a registration and fusion module; wherein the calibration module uses a calibration plate to obtain the intrinsic parameters of the microscope camera and solves the rigid body transformation between the microscope camera, the fluorescence camera, and the OCT probe based on the extrinsic parameters; the registration and fusion module performs distortion correction on the fluorescence image and projects it onto the microscope image coordinate system according to the rigid body transformation, resamples the OCT depth map to the microscope image resolution based on the depth-pixel mapping table, and aligns and stitches the modal data according to the timestamp of a unified time base to form a fused input vector.

[0079] In some embodiments, the improved tissue recognition convolutional neural network model includes: a multi-branch encoder, which extracts multi-scale features from microscopic video images, fluorescence imaging images, and OCT depth maps, and encodes surgical instrument pose and force feedback to obtain instrument interaction embedding; a spatiotemporal aggregation module, which enhances the temporal consistency of microscopic features in consecutive frames; a cross-modal attention fusion module, which performs attention fusion on microscopic features, fluorescence, and OCT features at multiple scale levels; a differentiable interactive safe distance field layer disposed between the cross-modal attention fusion module and the decoder, which generates a pixel-level safe distance field based on the instrument interaction embedding and OCT depth gradient, and uses the safe distance field as an attention bias and skip connection weight to modulate the decoded features; a boundary consistency decoder, including a semantic decoding path and a boundary refinement path, wherein the boundary features output by the boundary refinement path guide the semantic decoding path to achieve boundary alignment; and a multi-head output structure, including a tissue category output head, a boundary output head, a safe distance field regression head, a confidence output head, and an uncertainty output head, which output pixel-level tissue category map, boundary map, safe distance field, corresponding confidence, and uncertainty, respectively.

[0080] In some embodiments, the differentially interactive safe distance field layer maps the device interaction embedding to a device proximity probability map via a multilayer perceptron and fuses it with the boundary intensity map generated by the OCT depth gradient to obtain a safe distance field, wherein the safe distance field is calculated based on the soft distance from the pixel to the nearest high-risk boundary; the safe distance field is normalized by Sigmoid and injected as an attention bias term into the cross-modal attention weight calculation, and is used as a jump connection gating coefficient to weighted suppress or enhance the feature channels from the encoder to the decoder.

[0081] In some embodiments, the risk map generation unit fuses the pixel confidence and uncertainty output of the tissue category map to obtain a pixel risk value; wherein the risk value of the vascular structure region is set as structural risk, pixels located outside the planned cutting boundary and with a risk value higher than the threshold are marked as miscutting risk, pixels located within the target resection area and with a risk value lower than the threshold or uncertainty higher than the threshold are marked as residual risk, and a risk map containing three types of risk labels and their risk levels is output.

[0082] In some embodiments, the cutting path planning unit sets the instrument kinematic constraints as follows: the reachable workspace constraint, maximum curvature constraint, maximum velocity constraint, and minimum safe distance constraint between the instrument tip and the structural warning zone in the microscopic coordinate system; and sets the tissue cutability constraints as follows: the tissue hardness threshold obtained based on OCT depth map and force feedback estimation, the upper limit of allowable cutting depth, and the pause threshold under bleeding obscuration; the cutting path planning unit generates an initial path based on the boundary of the target resection area, and iteratively optimizes the path points using a risk map as a cost field to obtain a path point sequence; it assigns a cutting depth to each path point according to the upper limit of allowable cutting depth, and performs time parameterization on the path points according to the maximum velocity constraint of the instrument, outputting an executable trajectory containing the path point sequence, cutting depth, and velocity parameters.

[0083] In some embodiments, the output unit converts the executable trajectory into pixel trajectory lines in the microscopic video coordinate system and displays them synchronously overlaid on the risk map, wherein the cutting order and depth level are distinguished by color; when the structural risk or erroneous cutting risk in the risk map exceeds the threshold, a buzzer and indicator light alarm are triggered; when connected to the surgical robot, the trajectory point sequence and corresponding speed and depth parameters are encapsulated into control frames and sent to the robot controller, and at the same time, minimum safe distance and speed upper limit constraints are issued for confined auxiliary control.

[0084] In some embodiments, the cutting path planning unit adopts a two-level planning mechanism of global coarse planning and local real-time replanning. The global coarse planning generates a main path covering the entire area based on the boundary of the target resection area and the risk map. When the local real-time replanning detects that the instrument pose deviates from the threshold, the force feedback is abnormal, or the risk map update exceeds the threshold, it recalculates the next path segment with the instrument tip as the center within a preset local window and replaces the corresponding path segment to output a continuous executable trajectory.

[0085] In some embodiments, the boundary consistency decoder further includes an elongated structure connectivity constraint branch for extracting skeleton connectivity maps of blood vessels and nerves based on fluorescence imaging and OCT depth maps, and applying the skeleton connectivity map as a regularization constraint term to the boundary output head and the safe distance field regression head to suppress critical structure boundary breaks and improve the segmentation consistency of the critical structure neighborhood.

[0086] Compared with the prior art, the technical solution of the present invention has the following beneficial effects: (1) This invention calibrates and aligns microscopic video, fluorescence imaging, and OCT depth maps in a unified coordinate system with timestamp alignment, forming a consistent fusion input and significantly reducing misjudgments caused by fragmented multi-source information. The improved tissue recognition convolutional neural network introduces spatiotemporal aggregation and cross-modal attention fusion, which can suppress single-frame jitter and enhance the separability of tissue layers and thin structures under complex conditions such as bleeding, reflection, smoke, and instrument occlusion. Through boundary consistency decoding and slender structure connectivity constraints, it reduces the breakage and adhesion of blood vessel / nerve boundaries, and improves the segmentation accuracy and stability of key structural neighborhoods.

[0087] (2) The embedding position of the differentially interactive safe distance field layer in this application brings segmentation improvement by "controlling risks from the source": This invention sets the "differentiable interactive safe distance field layer" between the cross-modal attention fusion module and the decoder, placing it in a key channel before the "high semantic fusion features" enter "pixel-level reconstruction / boundary refinement". This layer utilizes the interactive embedding formed by the device pose and force feedback, and combines the tissue layer boundary strength extracted by OCT depth gradient to generate a pixel-level safe distance field, and directly modulates the decoding features using attention bias and skip-connection gating. Since the modulation occurs before decoding, it can suppress the over-response of high-risk neighborhoods during the feature reconstruction stage, enhance the boundary alignment and thin-layer structure resolution near key structures, thereby significantly reducing missegmentation, boundary drift and breakage in the vascular / nerve neighborhood.

[0088] (3) Under kinematic constraints such as instrument reachability, maximum curvature, speed limit, and minimum safe distance, the system generates an executable trajectory containing path point sequences, cutting depth, and speed parameters by combining tissue stiffness thresholds, allowable cutting depth limits, and occlusion pause thresholds based on OCT and force feedback estimation, thus avoiding trajectory suggestions that are "unreachable / uncuttable / unsafe." A global coarse planning and local real-time replanning mechanism is adopted to quickly update path segments when there are pose deviations, force anomalies, or significant changes in the risk map, maintaining continuous operability. The trajectory is intuitively guided through AR overlay and audio-visual cues. When connected to the robot, trajectory and constraint commands can be issued to achieve confined-domain assisted control, improving surgical efficiency and consistency.

[0089] (4) By using the safety distance field as one of the multiple outputs, and outputting it simultaneously with confidence and uncertainty, the system can convert "how close to the instrument contact area / critical structure and whether it can be safely cut" into pixel-level continuous quantities, thereby constructing three types of risk maps: structural risk, miscutting risk, and residual risk. Compared with schemes that rely solely on confidence thresholds, this invention automatically increases the modulation intensity and uncertainty level of the safety distance field when there is instrument contact, tensile deformation, or abnormal force, making the risk warning more consistent with the actual risk distribution during surgery, reducing the misleading effect of "high confidence misjudgment" by the model on the surgeon, and improving clinical interpretability and monitorability. Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products, and therefore this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.

[0090] While the present invention has been disclosed above, it is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of the invention; therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.

Claims

1. An AI-assisted real-time tissue recognition and cutting path planning system for microsurgery, characterized in that, include: Multimodal data acquisition unit: Acquires multimodal data of the same surgical field within a time period. The multimodal data includes microscopic video images and auxiliary data. The auxiliary data includes fluorescence imaging, OCT depth map, heart rate, blood pressure, surgical instrument pose data and force feedback data. Calibration and registration unit: Performs intrinsic and extrinsic parameter calibration on the camera and registers the fluorescence imaging and OCT depth map to the coordinate system of the microscopic video image; aligns multimodal data by timestamp to form a fused input vector; Tissue Recognition and Boundary Segmentation Unit: Inputs the fused input vector into the trained improved tissue recognition convolutional neural network model, and outputs pixel tissue category map, boundary map, corresponding confidence score and uncertainty; Risk map generation unit: Generates a risk map based on confidence level and model uncertainty estimation. The risk map represents structural risk, miscutting risk, and residual risk. Cutting path planning unit: Under the constraints of instrument kinematics and tissue cutability, it performs cutting path planning for the target resection area and outputs an executable trajectory containing path point sequence, cutting depth and speed parameters; Output unit: Outputs the executable trajectory to the screen as an augmented reality overlay, and outputs audio-visual prompts to indicate risk information. When connected to the surgical robot, it outputs trajectory instructions and constraint instructions to the robot controller for auxiliary control.

2. The AI-assisted real-time tissue recognition and cutting path planning system for microsurgery according to claim 1, characterized in that, The multimodal data acquisition unit includes a microscopic imaging acquisition module, a fluorescence imaging module, an OCT acquisition module, a physiological monitoring interface, and an instrument tracking and force feedback acquisition module. Microscopic video images and fluorescence imaging are synchronized at the frame level via hardware triggering. OCT depth maps are acquired according to a preset scanning cycle and aligned with the timestamps. Heart rate and blood pressure are acquired by the monitor via the communication interface and aligned with the image frames according to the sampling time. Instrument posture is acquired by an electromagnetic tracker, and force feedback is acquired by a micro-force sensor and output synchronously with the same time reference.

3. The AI-assisted real-time tissue recognition and cutting path planning system for microsurgery according to claim 1, characterized in that, The calibration and registration unit includes a calibration module and a registration and fusion module. The calibration module uses a calibration plate to obtain the intrinsic parameters of the microscope camera and solves the rigid body transformation between the microscope camera, the fluorescence camera, and the OCT probe based on the extrinsic parameters. The registration and fusion module performs distortion correction on the fluorescence image and projects it onto the microscope image coordinate system according to the rigid body transformation. It resamples the OCT depth map to the microscope image resolution based on the depth-pixel mapping table and aligns and stitches the modal data according to the timestamp of a unified time base to form a fused input vector.

4. The AI-assisted real-time tissue recognition and cutting path planning system for microsurgery according to claim 1, characterized in that, The improved tissue recognition convolutional neural network model includes: a multi-branch encoder, which extracts multi-scale features from microscopic video images, fluorescence imaging images, and OCT depth maps, and encodes surgical instrument pose and force feedback to obtain instrument interaction embeddings; a spatiotemporal aggregation module, used to enhance the temporal consistency of continuous frame microscopic features; a cross-modal attention fusion module, used to perform attention fusion of microscopic features with fluorescence and OCT features at multiple scale levels; a differentiable interactive safe distance field layer set between the cross-modal attention fusion module and the decoder, used to generate a pixel-level safe distance field based on the instrument interaction embeddings and OCT depth gradients, and use the safe distance field as attention bias and skip connection weights to modulate the decoded features; a boundary consistency decoder, including a semantic decoding path and a boundary refinement path, wherein the boundary features output by the boundary refinement path guide the semantic decoding path to achieve boundary alignment; and a multi-head output structure, including a tissue category output head, a boundary output head, a safe distance field regression head, a confidence output head, and an uncertainty output head, used to output pixel-level tissue category maps, boundary maps, safe distance fields, corresponding confidence, and uncertainties, respectively.

5. The AI-assisted real-time tissue recognition and cutting path planning system for microsurgery according to claim 4, characterized in that, The differentiable interactive safety distance field layer embeds the device interaction into a device proximity probability map via a multilayer perceptron and fuses it with the boundary intensity map generated by the OCT depth gradient to obtain the safety distance field. The safety distance field is calculated based on the soft distance from the pixel to the nearest high-risk boundary. After being normalized by Sigmoid, the safety distance field is injected as an attention bias term into the calculation of cross-modal attention weights and used as a jump connection gating coefficient to weighted suppress or enhance the feature channels from the encoder to the decoder.

6. The AI-assisted real-time tissue recognition and cutting path planning system for microsurgery according to claim 1, characterized in that, The risk map generation unit fuses the pixel confidence and uncertainty outputs of the organization category map to obtain the pixel risk value; The risk value of the vascular structure region is set as structural risk. Pixels located outside the planned cutting boundary and with a risk value higher than the threshold are marked as miscutting risk. Pixels located within the target resection area and with a risk value lower than the threshold or uncertainty higher than the threshold are marked as residual risk. A risk map containing three types of risk labels and their risk levels is output.

7. The AI-assisted real-time tissue recognition and cutting path planning system for microsurgery according to claim 1, characterized in that, The cutting path planning unit sets the instrument kinematic constraints as follows: the reachable workspace constraint, maximum curvature constraint, maximum speed constraint, and minimum safe distance constraint between the instrument tip and the structural warning zone in the microscopic coordinate system; and sets the tissue cutability constraints as follows: the tissue hardness threshold, the upper limit of allowable cutting depth, and the pause threshold under bleeding obscuration, all based on the OCT depth map and force feedback estimation. The cutting path planning unit generates an initial path based on the boundary of the target excision area, and iteratively optimizes the path points using a risk map as a cost field to obtain a path point sequence. It assigns a cutting depth to each path point according to the upper limit of the allowed cutting depth, and performs time parameterization on the path points according to the maximum speed constraint of the instrument, outputting an executable trajectory containing the path point sequence, cutting depth, and speed parameters.

8. The AI-assisted real-time tissue recognition and cutting path planning system for microsurgery according to claim 1, characterized in that, The output unit converts the executable trajectory into pixel trajectory lines in the microscopic video coordinate system and displays them synchronously overlaid on the risk map, wherein the cutting order and depth level are distinguished by color. When the structural risk or accidental cutting risk in the risk diagram exceeds the threshold, a buzzer and indicator light alarm are triggered. When connected to the surgical robot, the trajectory point sequence and corresponding speed and depth parameters are encapsulated into a control frame and sent to the robot controller. At the same time, minimum safe distance and speed upper limit constraints are issued for confined auxiliary control.

9. The AI-assisted real-time tissue recognition and cutting path planning system for microsurgery according to claim 1, characterized in that, The cutting path planning unit adopts a two-level planning mechanism of global coarse planning and local real-time replanning. The global coarse planning generates a main path covering the entire area based on the boundary of the target resection area and the risk map. When the local real-time replanning detects that the instrument pose deviates from the threshold, the force feedback is abnormal, or the risk map update exceeds the threshold, it recalculates the next path segment with the instrument tip as the center within a preset local window and replaces the corresponding path segment to output a continuous executable trajectory.

10. The AI-assisted real-time tissue recognition and cutting path planning system for microsurgery according to claim 1, characterized in that, The boundary consistency decoder further includes an elongated structure connectivity constraint branch, used to extract the skeleton connectivity graph of blood vessels and nerves based on fluorescence imaging and OCT depth maps, and to apply the skeleton connectivity graph as a regularization constraint term to the boundary output head and the safe distance field regression head, so as to suppress the boundary breakage of key structures and improve the segmentation consistency of the neighborhood of key structures.