AI-based neurosurgery auxiliary robot vision positioning system
The AI-based neurosurgical robot system addresses the limitations of single-modal imaging and lack of feedback in existing systems by integrating multi-modal imaging and deep learning for real-time path planning and correction, enhancing surgical precision and stability.
Patent Information
- Application Number
- CN202510729843.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing neurosurgical robot systems have problems in terms of intraoperative visual recognition and navigation, which are difficult to fully identify the deep structure and surface blood vessels of brain tissue, and lack dynamic feedback regulation mechanisms, resulting in the risk of operational lag and error accumulation.
Using AI-based neurosurgery-assisted robot visual positioning system, integrating structured light and near-infrared multimodal vision acquisition equipment, combining three-dimensional point cloud reconstruction and preoperative map spatial registration, identifying key anatomical features through multi-channel convolutional neural networks, and combining deep reinforcement learning and path optimization algorithms to achieve real-time structural change monitoring and attitude deviation correction.
It significantly improves the safety, stability and automation level of neurosurgery, improves the recognition accuracy of brain tissue structure and intraoperative response capabilities, and reduces the risk of surgery.
Smart Images

Figure CN120304948A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual positioning, and in particular to a visual positioning system for a neurosurgical operation assisting robot based on AI. Background Art
[0002] In neurosurgery, doctors usually rely on preoperative images, intraoperative visual perception and their own experience to locate lesions and judge the surgical path. However, due to the complex structure, dynamic changes and spatial visibility limitations of the brain tissue, traditional methods are difficult to meet the requirements of fine operation in terms of real-time performance, accuracy and stability. In recent years, with the continuous development of robot-assisted surgery technology, intelligent systems integrating artificial intelligence and medical image processing technology are gradually becoming the research focus for improving the accuracy and safety of neurosurgery.
[0003] The existing surgical robot systems still have the following deficiencies in intraoperative visual recognition and navigation: on the one hand, the traditional image processing module only relies on single-modal images, making it difficult to achieve a comprehensive recognition of the deep structures and surface blood vessels of the brain tissue, resulting in a lack of sufficient structural basis for intraoperative path planning. On the other hand, most systems lack an AI-based dynamic feedback adjustment mechanism and cannot identify intraoperative structural displacement or surgical instrument offset problems in real time, posing risks of operation lag and error accumulation. Summary of the Invention
[0004] The purpose of the present invention is to provide a visual positioning system for a neurosurgical operation assisting robot based on AI in view of the deficiencies of the prior art. The system aims to achieve a high-precision model construction of brain tissue structures by integrating structured light and near-infrared multi-modal visual acquisition devices, combining three-dimensional point cloud reconstruction and preoperative atlas spatial registration. By loading a multi-channel convolutional neural network to identify key anatomical features and combining deep reinforcement learning and path optimization algorithms, an obstacle avoidance path is dynamically generated. The system is configured with a real-time feedback adjustment module and a target confirmation mechanism to continuously monitor structural changes and pose deviations during the operation and automatically correct the execution path, thereby significantly improving the safety, stability and automation level of neurosurgery.
[0005] To this end, the present application provides a visual positioning system for a neurosurgical operation assisting robot based on AI, including the following modules:
[0006] A visual acquisition module for deploying and initializing a structured light scanning device and a near-infrared imaging device, performing multi-modal image acquisition of the surgical area, and performing standardization processing, spatio-temporal alignment and fusion of the acquired images.
[0007] A registration and fusion module for configuring an intraoperative three-dimensional point cloud model and a three-dimensional mesh, configuring a preoperative intracranial structure feature atlas and spatial registration, and performing registration error judgment and an adaptive optimization mechanism.
[0008] A feature recognition module, which is used to load and fuse an anatomical space model, normalize it into a network input tensor, load a multi-channel convolutional neural network model, perform structure recognition, perform semantic mapping between the recognition result and a standard neuroanatomical atlas, and configure a structure index binding table and a confidence judgment mechanism.
[0009] A path planning module, which is used to calibrate the target tissue area, configure path input parameters, configure a deep reinforcement learning policy network, generate a candidate path sequence, perform path optimization and continuous trajectory generation, configure a path fallback mechanism, and dynamically trigger path replanning.
[0010] A control and regulation module, which is used to configure an intraoperative action perception feedback mechanism and a state prediction model, perform path deviation recognition and a dynamic adjustment mechanism, configure an actuator closed-loop control strategy, and generate a feedback status report.
[0011] An identification and confirmation module, which is used to configure a target confirmation comparison set, perform target error evaluation, perform tolerance judgment, generate a standard control instruction, and issue the instruction to access the intraoperative doctor confirmation interface.
[0012] In some specific embodiments, the visual acquisition module specifically includes:
[0013] Deploy and initialize a structured light scanning device and a near-infrared imaging device, and perform multi-modal image acquisition of the surgical area.
[0014] Fix and install the structured light scanning device and the near-infrared imaging device on the end effector of the neurosurgical operation assistance robot, and perform device parameter initialization and spatial calibration operations on the two types of imaging devices respectively.
[0015] The structured light scanning device uses a structured light projector with stripe coding projection ability and a binocular camera module to form a structured light image acquisition system.
[0016] The near-infrared imaging device uses a short-wave near-infrared camera with a working band of 800nm to 1100nm, which has dynamic gain adjustment and automatic exposure control functions, and is used for near-infrared image acquisition of blood vessel paths and nerve trajectories.
[0017] Start the visual acquisition and processing module, control the structured light image acquisition system to project a stripe coding light pattern onto the surgical area, and use the binocular camera module to synchronously acquire reflected image data from multiple perspectives.
[0018] Synchronously start the near-infrared imaging system, perform real-time infrared image acquisition of the same surgical area, obtain surface nerve and blood vessel structure image information, and the acquisition process is completed under the control of a unified timestamp.
[0019] Perform standardization processing, spatio-temporal alignment and fusion on the acquired images.
[0020] Perform image preprocessing operations on the acquired structured light image and infrared image respectively. The preprocessing includes: distortion correction processing, illumination equalization processing, artifact removal processing, and edge enhancement processing.
[0021] For the distortion correction processing, according to the device calibration results, use the back-projection method to comprehensively correct the image distortion caused by lens distortion. For the illumination equalization processing, perform histogram equalization on the image brightness distribution to enhance the image contrast. For the artifact removal processing, use a combined algorithm of median filtering and edge-preserving smoothing filtering to remove the noise and light spots generated by acquisition interference. For the edge enhancement processing, perform superposition processing of the Sobel edge extraction algorithm and the non-linear enhancement algorithm on the structure boundary to improve the clarity of the key tissue contour.
[0022] Input the preprocessed structured light image and infrared image into the spatio-temporal synchronization module to perform image alignment and fusion operations, including: spatial registration processing, time synchronization processing, and image fusion processing.
[0023] For the spatial registration processing, use the rigid registration algorithm based on feature points to map the structured light image and infrared image to a unified spatial coordinate system. For the time synchronization processing, precisely align the two types of images through frame-level timestamps. For the image fusion processing, use a multi-channel fusion strategy to construct a fused image tensor, where the structured light image is used to provide three-dimensional depth information, and the infrared image is used to provide soft tissue recognition features.
[0024] The fused result will be output as a standard multi-modal image data block, and the fused result has multi-level feature dimensions such as depth, infrared, and texture.
[0025] In some specific embodiments, the registration and fusion module specifically includes:
[0026] Configure the intraoperative three-dimensional point cloud model and three-dimensional mesh.
[0027] Input the structured light image acquired by the structured light image acquisition system into the multi-view geometric reconstruction module to perform depth estimation operations based on disparity calculation, and restore the three-dimensional point coordinates through the disparity information between binocular image pairs.
[0028] Calculate the three-dimensional space coordinate value corresponding to each pixel point.
[0029] Generate an intraoperative three-dimensional point cloud dataset according to the calculation results, and perform point cloud cropping and boundary constraint processing, and retain the valid point cloud data within the spatial range of the surgical area.
[0030] Input the cropped intraoperative three-dimensional point cloud into the meshing processing unit in the reconstruction, registration and fusion module, and use the Poisson surface reconstruction algorithm to generate a continuous triangular mesh model for the point cloud data.
[0031] The triangular mesh model is standardized and output in STL format, and the Laplacian smoothing algorithm is executed for surface optimization of non-boundary points to suppress depressions and cusps caused by edge discontinuities or noise during the reconstruction process.
[0032] Configure the preoperative intracranial structure feature atlas and spatial registration, and perform registration error judgment and adaptive optimization mechanism.
[0033] Call the preoperative magnetic resonance imaging (MRI) images and computed tomography (CT) images, perform three-dimensional registration and resampling on the preoperative images in the image processing module, and set the resampling interval to 0.5 mm.
[0034] Use a convolutional neural network model based on the UNet structure to perform multi-class semantic segmentation on the preoperative MRI images, extract the ventricular, cortical, and skull boundary structures, and generate an intracranial structure label map.
[0035] Encode the segmentation results into three-dimensional structure labels, construct a preoperative structure feature index table, and uniformly embed the ventricular, cortical, and skull boundary labels into the spatial reference coordinate system of the preoperative images to form a standard preoperative intracranial structure feature atlas.
[0036] Perform a spatial registration operation on the constructed intraoperative three-dimensional point cloud model and the preoperative intracranial structure feature atlas, and use a rigid registration method based on feature point matching to complete the coordinate alignment process.
[0037] The rigid registration process is as follows: Use the SURF algorithm to extract the key structural feature points in the intraoperative three-dimensional point cloud and the preoperative atlas. After configuring the set of matching point pairs, use the RANSAC algorithm to eliminate the mismatched point pairs, generate an initial pose transformation estimate, and use the ICP algorithm for fine registration to solve the optimal rotation matrix R and translation vector t, satisfying the objective function of minimizing the registration error. The registration error is quantified and calculated using the root mean square error (RMSE).
[0038] When the registration error meets the condition of ≤1.0 mm, the system fuses the intraoperative three-dimensional point cloud model and the preoperative intracranial structure feature atlas under a unified spatial reference coordinate system to generate a fused anatomical space model.
[0039] The fused anatomical space model includes: the intraoperative brain tissue surface structure layer, the preoperative internal anatomical structure layer, and the structure index mapping layer.
[0040] In some specific embodiments, the feature recognition module specifically includes:
[0041] Load the fused anatomical space model and standardize it into a network input tensor, load a multi-channel convolutional neural network model, and perform structure recognition.
[0042] Taking the fused anatomical space model output by the reconstruction registration and fusion module as input data, the fused anatomical space model includes the intraoperative brain tissue surface structure layer, the preoperative internal anatomical structure layer, and the structure index mapping layer.
[0043] Through the data preprocessing module, the fused anatomical space model is formatted into a standard input tensor, and the tensor dimensions are set to [C, H, W, D].
[0044] C represents the number of image channels, including the depth information channel, the texture information channel, and the structure label channel.
[0045] H, W, and D respectively represent the height, width, and depth of the tensor in the spatial coordinate system, with the unit of pixel.
[0046] The tensor values are normalized to the [0, 1][0, 1][0, 1] interval, and the structure index mapping information is bound.
[0047] Call the multi-channel convolutional neural network model deployed in the feature recognition and extraction module. The model structure is based on the ResNet backbone network, and the feature channels for anatomical structure classification are extended.
[0048] After inputting the standard tensor, the output includes the structure label map and the recognition confidence heat map. The structure label map annotates the anatomical structure categories of brain sulci, gyri, ventricles, vascular bifurcation points, tumor boundaries, and skull edges. The recognition confidence heat map has a value range of [0, 1][0, 1][0, 1] at each position, which is used to represent the credibility score of the structure label result.
[0049] Execute the semantic mapping between the recognition result and the standard neuroanatomical atlas, and configure the structure index binding table and the confidence judgment mechanism.
[0050] Input the structure label map into the anatomical semantic mapping module, and conduct a structural semantic comparison with the standard neuroanatomical atlas database. Map all recognition labels to the medical semantic standard structure names uniformly, and establish a standard term mapping table for the brain sulci, gyri, ventricles, cortical regions, and vascular path structure types identified in the label map.
[0051] If the offset of the recognition result from the center coordinates of the standard structure label exceeds 2.5 mm, the system records the recognition deviation mark.
[0052] According to the structure label map and the recognition confidence heat map, construct the structure index binding table, bind the structure name, spatial coordinates, and recognition confidence value to each recognized structure point. If the recognition confidence of a certain structure point is lower than 85%, the feature recognition and extraction module automatically triggers the local area re-recognition process, re-intercepts the image sub-block of this area, and performs the structure recognition inference operation.
[0053] Generate a set of structure annotation points from all the structure recognition results judged by the confidence threshold. The set of structure annotation points includes: structure label, three-dimensional coordinates, structure number, and confidence score. The set data format is output in a combined manner using the PLY format and the JSON format of the structure attribute index table.
[0054] In some specific embodiments, the path planning module specifically includes:
[0055] Calibrate the target tissue area and configure the path input parameters.
[0056] Calibrate the target tissue area in the set of structure annotation points, including: the tumor core, the lesion edge, and the resection target structure set by the doctor before surgery. Extract the three-dimensional spatial coordinates, structure labels, and recognition confidence information of the target tissue from the structure index binding table.
[0057] Configure the set of input parameters required for path planning, respectively: the spatial coordinates of the target tissue, the set of spatial coordinates of the structures identified as high-risk structures in the structure index binding table, the current position and attitude information of the end effector of the neurosurgical assistant robot, including: position vector and attitude angle, the structural arm length parameter, motion range constraint, attitude change speed, and execution step accuracy of the neurosurgical assistant robot.
[0058] Configure the deep reinforcement learning policy network and generate a candidate path sequence.
[0059] Call the deep reinforcement learning policy network in the positioning path planning module. The policy network includes a state input layer, a policy generation layer, and a value evaluation layer.
[0060] The state input layer is used to input the current motion state of the neurosurgical assistant robot, the position of the target tissue, and the distribution of dangerous structures. The policy generation layer uses the Actor network to generate candidate path points. The value evaluation layer uses the Critic network to calculate the cumulative return value of the path policy. The reinforcement learning goal is to maximize the cumulative reward function R.
[0061] The closer the end position of the robot is to the target tissue, the more positive reward is given. If the path is close to dangerous structures or violates the operation constraints, negative punishment is given. If the path is smooth and the attitude stability is high, gain reward is given.
[0062] Output the path point sequence {P1, P2, …, P n}, where the path point P i includes: spatial position vector (X i , Y i , Z i ), the corresponding motion direction vector recommended execution speed value V i , with the unit of millimeters per second.
[0063] Perform path optimization and continuous trajectory generation, configure a path fallback mechanism, and dynamically trigger path replanning.
[0064] Input the path point sequence output by the deep reinforcement learning policy network into the path optimization module, and use a multi-level path optimization algorithm combination to achieve fine correction.
[0065] The multi-level path optimization algorithm includes: heuristic graph search algorithm A, dynamic path correction using the DLite algorithm, and B-spline curve interpolation method.
[0066] The heuristic graph search algorithm A is used for initial path screening to eliminate path segments that cross high-risk structural areas in the path. The dynamic path correction uses the DLite algorithm to enable the path to respond in real time to intraoperative environmental changes. The continuous path generation uses the B-spline curve interpolation method to smooth the path point sequence and generate a continuous executable robot execution trajectory.
[0067] The final path sequence is cached in the path planning buffer, and a path topology index table is established for path state tracking and dynamic instruction synchronization control.
[0068] The path planning module continuously receives the structural annotation point coordinate status information and the real-time motion status of the end effector transmitted by the feedback control adjustment module, and performs path stability judgment.
[0069] When any condition is met, the path fallback mechanism is triggered, including: the spatial position change amount of any structural annotation point exceeds 2.0 mm, a new high-risk structural label recognition result appears in the path planning area, the current position of the end effector of the neurosurgical operation assistance robot deviates from the current path segment by more than the preset threshold, or the attitude angle deviation exceeds 10 degrees.
[0070] In some specific embodiments, the control adjustment module specifically includes:
[0071] Configure an intraoperative action perception feedback mechanism and a state prediction model, and execute a path offset recognition and dynamic adjustment mechanism.
[0072] Configure a real-time data interface between the feedback control adjustment module and the visual acquisition and processing module, and perform periodic acquisition and synchronous update of information. The periodic acquisition and synchronous update of information are: the spatial position, attitude angle, and execution speed information of the end effector of the neurosurgical operation assistance robot, the current path point sequence and the path topology index table, the structured light image and the near-infrared image output by the visual acquisition and processing module, and the set of structural annotation points and the structural label map generated by the feature recognition and extraction module.
[0073] Set the acquisition frequency to 1 Hz, and manage the acquisition data through a unified time stamp.
[0074] Call the state prediction sub-module in the feedback control adjustment module, use the Kalman filter to predict the spatial position of the end effector of the neurosurgical operation assistance robot, and identify abnormal offsets in combination with the structural annotation point set.
[0075] The structural offset identification logic is as follows: Compare the spatial coordinates of the current structural annotation points with the predicted coordinates, calculate the offset. If the offset of any structural annotation point exceeds 1.2 mm, record it as a structural offset event, attach the structural category information to the offset event, and include it in the path correction judgment parameter set.
[0076] Take the structural offset event as the input, and calculate the path offset degree D in combination with the path point sequence p .
[0077] When the maximum path offset D p ≥1.2 mm, the system immediately activates the path rollback mechanism. The feedback control adjustment module sends a path update request to the positioning path planning module and regenerates the target path segment.
[0078] Configure the actuator closed-loop control strategy and generate a feedback status report.
[0079] According to the path correction result, update the motion instruction parameters of the end effector of the neurosurgical operation assistance robot to make the motion trajectory consistent with the latest path. The execution control adopts a closed-loop logic, and the control period is the same as the feedback acquisition frequency, both are 1 Hz.
[0080] The closed-loop control logic includes: If the path does not deviate, keep running the current trajectory. If the path offset is between 1.2 mm and 2.0 mm, fine-tune the motion direction and speed parameters of the end effector to perform dynamic deviation correction operations. If the path offset ≥2.0 mm, trigger the path rollback mechanism and pause the execution, and resume the operation after the path replanning is completed.
[0081] In some specific embodiments, the recognition and confirmation module specifically includes:
[0082] Configure the target confirmation comparison set, perform target error evaluation, and perform tolerance judgment.
[0083] Call the annotation point coordinates (X T , Y T , Z T ) calibrated as the target structure from the structural annotation point set, extract its structural label and confidence information, and load the corresponding reference attitude angle values α T , β T , γ T .
[0084] At the same time, the current state of the end effector output by the feedback control adjustment module is called to obtain the real-time spatial position (X E , Y E , Z E ) and the attitude angle values (α E , β E , γ E ). The target structure and the current execution state are used to form a target confirmation comparison set, which is used as the input for the joint evaluation of position error and attitude error.
[0085] Calculate the Euclidean distance between the current position of the actuator and the coordinates of the target structure.
[0086] At the same time, calculate the attitude angle error, and use the average angular difference of three axes to define the attitude error D att .
[0087] Generate a standard control instruction and send the instruction to access the intraoperative doctor confirmation interface.
[0088] The control instruction is encoded by the recognition and confirmation output module according to the control bus protocol and written into the control instruction buffer for scheduling by the motion execution module of the neurosurgical operation assistance robot.
[0089] In summary, the AI-based visual positioning system for neurosurgical operation assistance robots provided by this application realizes high-precision recognition and dynamic modeling of brain tissue structures by constructing a multi-modal visual acquisition and preprocessing module, a three-dimensional tissue model construction and registration module, a visual-anatomical feature recognition and extraction module, an AI-assisted positioning and path optimization module, an intraoperative dynamic perception and feedback control module, and a target position confirmation and instruction output module. It can perform real-time recognition, path planning, and position correction of key anatomical structures during the operation, significantly improving the accuracy, intelligent level, and intraoperative response ability of neurosurgical operations, effectively reducing the surgical risk, and enhancing the positioning reliability and automatic control efficiency. Description of the Drawings
[0090] Figure 1 is the overall framework diagram of the AI-based visual positioning system for neurosurgical operation assistance robots provided by the embodiments of this application.
[0091] Figure 2 is the overall flow chart of the AI-based visual positioning system for neurosurgical operation assistance robots provided by the embodiments of this application. Detailed Embodiments
[0092] Please refer to Figure 1 , which shows the process of an embodiment of the AI-based visual positioning system for neurosurgical operation assistance robots according to the present disclosure
[0093] As Figure 1As shown, the AI-based visual positioning system for neurosurgical robotic assistance includes the following modules:
[0094] The visual acquisition module is used for the deployment and initialization of the structured light scanning device and the near-infrared imaging device, and performs multi-modal image acquisition of the surgical area, and conducts standardization processing, spatio-temporal alignment, and fusion of the acquired images.
[0095] The registration and fusion module is used to configure the intraoperative three-dimensional point cloud model and three-dimensional mesh, configure the preoperative intracranial structure feature atlas and spatial registration, and perform registration error judgment and adaptive optimization mechanism.
[0096] The feature recognition module is used to load the fused anatomical space model and standardize it into a network input tensor, load the multi-channel convolutional neural network model, and perform structure recognition, perform semantic mapping of the recognition result and the standard neuroanatomical atlas, and configure the structure index binding table and confidence judgment mechanism.
[0097] The path planning module is used to calibrate the target tissue area, configure the path input parameters, configure the deep reinforcement learning policy network, generate a candidate path sequence, perform path optimization and continuous trajectory generation, configure the path fallback mechanism, and dynamically trigger path replanning.
[0098] The control and adjustment module is used to configure the intraoperative action perception feedback mechanism and state prediction model, and perform path deviation recognition and dynamic adjustment mechanism, configure the actuator closed-loop control strategy, and generate a feedback status report.
[0099] The recognition and confirmation module is used to configure the target confirmation comparison set, perform target error evaluation, conduct tolerance judgment, generate standard control instructions, and send the instructions to access the intraoperative doctor confirmation interface.
[0100] In some specific embodiments, the visual acquisition module specifically includes:
[0101] The deployment and initialization of the structured light scanning device and the near-infrared imaging device, and the multi-modal image acquisition of the surgical area.
[0102] The structured light scanning device and the near-infrared imaging device are fixedly installed on the end effector of the neurosurgical robotic assistant, and device parameter initialization and spatial calibration operations are respectively performed on the two types of imaging devices.
[0103] The structured light scanning device uses a structured light projector with stripe coding projection ability and a binocular camera module to form a structured light image acquisition system.
[0104] The near-infrared imaging device uses a short-wave near-infrared camera with a working wavelength range of 800 nm to 1100 nm, and has functions of dynamic gain adjustment and automatic exposure control, and is used for collecting near-infrared images of blood vessel paths and nerve trajectories.
[0105] Start the visual acquisition and processing module, control the structured light image acquisition system to project a stripe-coded light pattern onto the surgical area, and use the binocular camera module to synchronously collect reflected image data from multiple perspectives.
[0106] Synchronously start the near-infrared imaging system to perform real-time infrared image acquisition on the same surgical area, obtain surface nerve and blood vessel structure image information, and the acquisition process is completed under the control of a unified timestamp.
[0107] Perform normalization processing, spatio-temporal alignment, and fusion on the acquired images.
[0108] Perform image preprocessing operations on the acquired structured light images and infrared images respectively. The preprocessing includes: distortion correction processing, illumination equalization processing, artifact removal processing, and edge enhancement processing.
[0109] For distortion correction processing, according to the device calibration results, use the back-projection method to comprehensively correct the image distortion caused by lens distortion. For illumination equalization processing, perform histogram equalization on the image brightness distribution to enhance the image contrast. For artifact removal processing, use a combined algorithm of median filtering and edge-preserving smoothing filtering to remove noise and light spots generated by acquisition interference. For edge enhancement processing, perform superposition processing of the Sobel edge extraction algorithm and the non-linear enhancement algorithm on the structure boundary to improve the clarity of the key tissue contour.
[0110] Input the preprocessed structured light images and infrared images into the spatio-temporal synchronization module to perform image alignment and fusion operations, including: spatial registration processing, time synchronization processing, and image fusion processing.
[0111] For spatial registration processing, use a rigid registration algorithm based on feature points to map the structured light image and the infrared image to a unified spatial coordinate system. For time synchronization processing, precisely align the two types of images through frame-level timestamps. For image fusion processing, use a multi-channel fusion strategy to construct a fused image tensor, where the structured light image is used to provide three-dimensional depth information, and the infrared image is used to provide soft tissue recognition features.
[0112] The fused result will be output as a standard multi-modal image data block, and the fused result has multi-level feature dimensions such as depth, infrared, and texture.
[0113] In some specific embodiments, the registration and fusion module specifically includes:
[0114] Configure the intraoperative three-dimensional point cloud model and three-dimensional mesh.
[0115] Input the structured light image collected by the structured light image acquisition system into the multi-view geometric reconstruction module, perform depth estimation operations based on disparity calculation, and restore the three-dimensional point coordinates through the disparity information between binocular image pairs.
[0116] Calculate the three-dimensional space coordinate values corresponding to each pixel point, specifically:
[0117]
[0118] In the formula: Z is the depth coordinate of the pixel point in the camera coordinate system, in millimeters; X is the horizontal coordinate of the pixel point in the camera coordinate system, in millimeters; y is the vertical coordinate of the pixel point in the camera coordinate system, in millimeters; f is the focal length of the camera, in pixels; B is the baseline length of the binocular cameras in the structured light image acquisition system, in millimeters; d is the disparity value of the pixel point in the binocular images, in pixels; u is the horizontal pixel position of the pixel point in the image; v is the vertical pixel position of the pixel point in the image; c x is the pixel coordinate of the image center in the horizontal direction; c y is the pixel coordinate of the image center in the vertical direction.
[0119] Generate an intraoperative three-dimensional point cloud dataset based on the calculation results, perform point cloud cropping and boundary constraint processing, and retain the valid point cloud data within the spatial range of the surgical area.
[0120] Input the cropped intraoperative three-dimensional point cloud into the meshing processing unit in the reconstruction registration and fusion module, and use the Poisson surface reconstruction algorithm to generate a continuous triangular mesh model for the point cloud data.
[0121] The triangular mesh model is standardized and output in STL format, and the Laplacian smoothing algorithm is executed for the surface optimization of non-boundary points to suppress the depressions and sharp points caused by edge discontinuities or noise during the reconstruction process.
[0122] Configure the preoperative intracranial structure feature atlas and spatial registration, and perform registration error judgment and adaptive optimization mechanism.
[0123] Call the preoperative magnetic resonance imaging (MRI) images and computed tomography (CT) images, perform three-dimensional registration and resampling on the preoperative images in the image processing module, and set the resampling interval to 0.5 mm.
[0124] Use a convolutional neural network model based on the UNet structure to perform multi-class semantic segmentation on the preoperative MRI images, extract the ventricular system, cortical area, and skull boundary structures, and generate an intracranial structure label map.
[0125] Perform three-dimensional structural label encoding on the segmentation results, construct a preoperative structural feature index table, and uniformly embed the ventricle, cortical region, and skull boundary labels into the spatial reference coordinate system of the preoperative image to form a standard preoperative intracranial structural feature atlas.
[0126] Perform a spatial registration operation on the constructed intraoperative three-dimensional point cloud model and the preoperative intracranial structural feature atlas, and use a rigid registration method based on feature point matching to complete the coordinate alignment process.
[0127] The rigid registration process is as follows: Use the SURF algorithm to extract the key structural feature points in the intraoperative three-dimensional point cloud and the preoperative atlas. After configuring the set of matching point pairs, use the RANSAC algorithm to remove the mismatched point pairs to generate an initial pose transformation estimate. Then use the ICP algorithm for fine registration to solve the optimal rotation matrix R and translation vector t, which satisfy the objective function of minimizing the registration error. The registration error is quantitatively calculated using the root mean square error RMSE. Specifically:
[0128]
[0129] In the formula: RMSE is the root mean square error in the registration process of the preoperative and intraoperative models, with the unit of millimeters; N is the number of effective feature points participating in the registration. is the three-dimensional coordinate of the i-th feature point in the intraoperative three-dimensional point cloud, with the unit of millimeters. is the three-dimensional coordinate corresponding to the i-th intraoperative point in the preoperative structural atlas, with the unit of millimeters. R is the three-dimensional rotation matrix for rigid registration, t is the three-dimensional translation vector for rigid registration, and · represents the Euclidean distance.
[0130] Set the registration error tolerance threshold to 1.0 mm. When the RMSE value of the registration result exceeds this tolerance threshold, the system will automatically start the registration optimization mechanism and perform adaptive reconstruction operations, including: increasing the density of SURF feature point extraction, performing local resampling and sparse region interpolation on the intraoperative point cloud, adding an image gray histogram similarity judgment link to enhance the stability of point pair matching, and performing the initial matching and ICP fine registration processes again until the RMSE meets the requirements.
[0131] When the registration error satisfies the condition of ≤ 1.0 mm, the system fuses the intraoperative three-dimensional point cloud model and the preoperative intracranial structural feature atlas under the unified spatial reference coordinate system to generate a fused anatomical space model.
[0132] The fused anatomical space model includes: the intraoperative brain tissue surface structure layer, the preoperative internal anatomical structure layer, and the structure index mapping layer.
[0133] The triangular mesh surface of the intraoperative brain tissue surface structure layer is reconstructed from the structured light image, the preoperative internal anatomical structure layer is the structure label map generated by segmenting the MRI image, and the structure index mapping layer records the structure label, source identifier and corresponding spatial coordinates.
[0134] In some specific embodiments, the feature recognition module specifically includes:
[0135] Load the fused anatomical space model and standardize it into a network input tensor, load the multi-channel convolutional neural network model, and perform structure recognition.
[0136] Use the fused anatomical space model output by the reconstruction registration fusion module as the input data. The fused anatomical space model includes the intraoperative brain tissue surface structure layer, the preoperative internal anatomical structure layer and the structure index mapping layer.
[0137] Through the data preprocessing module, format the fused anatomical space model into a standard input tensor, and set the tensor dimension to [C, H, W, D].
[0138] C represents the number of image channels, including the depth information channel, texture information channel and structure label channel.
[0139] H, W, and D respectively represent the height, width and depth of the tensor in the spatial coordinate system, and the unit is pixel.
[0140] Normalize the tensor values to the interval [0, 1][0, 1][0, 1], and bind the structure index mapping information.
[0141] Call the multi-channel convolutional neural network model deployed in the feature recognition extraction module. The model structure is based on the ResNet backbone network and extends the feature channels for anatomical structure classification.
[0142] After inputting the standard tensor, the output includes the structure label map and the recognition confidence heat map. The structure label map is: label the anatomical structure categories of brain sulci, gyri, ventricles, vascular bifurcation points, tumor boundaries and skull edges. The recognition confidence heat map is: the value range of each position is [0, 1][0, 1][0, 1], which is used to represent the credibility score of the structure label result.
[0143] Perform the semantic mapping of the recognition result and the standard neuroanatomical atlas, and configure the structure index binding table and confidence judgment mechanism.
[0144] Input the structure label map into the anatomical semantic mapping module, and perform structural semantic comparison with the standard neuroanatomical atlas database. Map all recognition labels to the medical semantic standard structure names uniformly, and establish a standard term mapping table for the identified brain sulci, gyri, ventricles, cortical areas and vascular path structure types in the label map.
[0145] If the offset of the recognition result from the center coordinates of the standard structure label exceeds 2.5 mm, the system records the recognition deviation mark.
[0146] According to the structure label map and the recognition confidence heat map, construct a structure index binding table, bind the structure name, spatial coordinates, and recognition confidence value to each recognized structure point. If the recognition confidence of a certain structure point is lower than 85%, the feature recognition extraction module automatically triggers the local area re-recognition process, re-intercepts the image sub-block of this area, and performs the structure recognition inference operation.
[0147] Generate a set of structure annotation points for all structure recognition results that pass the confidence threshold judgment. The set of structure annotation points includes: structure label, three-dimensional coordinates, structure number, and confidence score. The set data format is output in combination with the PLY format and the JSON format of the structure attribute index table.
[0148] In some specific embodiments, the path planning module specifically includes:
[0149] Calibrate the target tissue area and configure the path input parameters.
[0150] Calibrate the target tissue area in the set of structure annotation points, including: the tumor core, the lesion edge, and the resection target structure set by the doctor before the operation. Extract the three-dimensional spatial coordinates, structure label, and recognition confidence information of the target tissue from the structure index binding table.
[0151] Configure the set of input parameters required for path planning, respectively: the spatial coordinates of the target tissue, the set of spatial coordinates of the structures recognized as high-risk structures in the structure index binding table, the current position and pose information of the end effector of the neurosurgery-assisted robot, including: position vector and attitude angle, the structural arm length parameter, motion range constraint, attitude change speed, and execution step accuracy of the neurosurgery-assisted robot.
[0152] Configure the deep reinforcement learning policy network and generate a candidate path sequence.
[0153] Call the deep reinforcement learning policy network in the positioning path planning module. The policy network includes a state input layer, a policy generation layer, and a value evaluation layer.
[0154] The state input layer is used to input the current motion state of the neurosurgery-assisted robot, the target tissue position, and the distribution of dangerous structures. The policy generation layer uses the Actor network to generate candidate path points. The value evaluation layer uses the Critic network to calculate the cumulative return value of the path policy. The reinforcement learning goal is to maximize the cumulative reward function R, specifically:
[0155]
[0156] Where: R is the cumulative reward value of the path planning strategy, representing the comprehensive score during the entire path execution process, used to measure the quality of the strategy; γ is the reward decay factor, with a value range of [0, 1], used to reduce the impact of late rewards on the total strategy value; t is the time step number of the path execution, representing the current time point in the path planning, and r t is the immediate reward value obtained after performing the action at the t-th step, dynamically calculated based on the degree of proximity to the target, obstacle avoidance success, and pose stability of the path.
[0157] The closer the end position of the robot is to the target tissue, the positive reward is given; if the path is close to dangerous structures or violates operation constraints, negative punishment is given; if the path is smooth and the pose stability is high, gain rewards are given.
[0158] Output the path point sequence {P1, P2, …, P n}, where the path point P i includes: the spatial position vector (X i , Y i , Z i ), the corresponding motion direction vector the recommended execution speed value V i , with the unit of millimeters per second.
[0159] Perform path optimization and continuous trajectory generation, configure the path fallback mechanism, and dynamically trigger path replanning.
[0160] Input the path point sequence output by the deep reinforcement learning policy network into the path optimization module, and use a multi-level path optimization algorithm combination to achieve fine correction.
[0161] The multi-level path optimization algorithm includes: the heuristic graph search algorithm A, the dynamic path correction using the DLite algorithm, and the B-spline curve interpolation method.
[0162] The heuristic graph search algorithm A is used for initial path screening, eliminating path segments that cross high-risk structure areas in the path; the dynamic path correction uses the DLite algorithm to enable the path to respond in real time to intraoperative environmental changes; the continuous path generation uses the B-spline curve interpolation method to smooth the path point sequence and generate a continuous executable robot execution trajectory.
[0163] The final path sequence is cached in the path planning buffer, and a path topology index table is established for path state tracking and dynamic instruction synchronization control.
[0164] The path planning module continuously receives the structure annotation point coordinate status information and the real-time motion status of the end effector transmitted by the feedback control adjustment module, and performs path stability judgment.
[0165] When any one of the conditions is met, the path fallback mechanism is triggered, including: the spatial position change of any structural annotation point exceeds 2.0 mm, a new high-risk structure label recognition result appears within the path planning area, the current position of the end effector of the neurosurgical operation assistance robot deviates from the current path segment by more than a preset threshold, or the attitude angle deviation exceeds 10 degrees.
[0166] The fallback mechanism traces back to the previous stable path node according to the path topology index table.
[0167] In some specific embodiments, the control and adjustment module specifically includes:
[0168] Configure the intraoperative action perception feedback mechanism and the state prediction model, and execute the path deviation recognition and dynamic adjustment mechanism.
[0169] Configure the real-time data interface between the feedback control and adjustment module and the visual acquisition and processing module, and perform periodic acquisition and synchronous update of information. The periodic acquisition and synchronous update of information include: the spatial position, attitude angle and execution speed information of the end effector of the neurosurgical operation assistance robot, the current path point sequence and the path topology index table, the structured light image and the near-infrared image output by the visual acquisition and processing module, and the set of structural annotation points and the structure label map generated by the feature recognition and extraction module.
[0170] Set the acquisition frequency to 1 Hz, and manage the acquired data through a unified time stamp.
[0171] Call the state prediction sub-module in the feedback control and adjustment module, use the Kalman filter to predict the spatial position of the end effector of the neurosurgical operation assistance robot, and combine the set of structural annotation points for abnormal deviation recognition.
[0172] The structure deviation recognition logic is: compare the spatial coordinates of the current structural annotation points with the predicted coordinates, calculate the deviation amount. If the deviation amount of any structural annotation point exceeds 1.2 mm, it is recorded as a structure deviation event, attach the structure category information to the deviation event, and include it in the path correction judgment parameter set.
[0173] Take the structure deviation event as the input, and calculate the path deviation degree D in combination with the path point sequence p , specifically:
[0174]
[0175] In the formula: D p is the maximum path deviation, representing the maximum spatial deviation between the actual coordinates and the predicted coordinates among all structural annotation points, with the unit of millimeter. It is an operation to take the maximum value of the offsets among the 1st to the Nth structural annotation points. N is the number of structural annotation points, representing the total number of key structural points used for path offset judgment. is the actual spatial coordinate of the structural annotation point i, which is output in real time by the visual acquisition and processing module and the feature recognition and extraction module. is the predicted spatial coordinate of the structural annotation point i, which is generated by the Kalman filter. '.' is the Euclidean distance operator, representing the spatial straight-line distance between two three-dimensional coordinate points, with the unit of millimeters.
[0176] When the maximum path offset D p ≥ 1.2 mm, the system immediately activates the path backtracking mechanism. The feedback control and adjustment module sends a path update request to the positioning path planning module and regenerates the target path segment.
[0177] Configure the actuator closed-loop control strategy and generate a feedback status report.
[0178] According to the path correction result, update the motion instruction parameters of the end effector of the neurosurgical operation assistance robot to make the motion trajectory consistent with the latest path. The execution control adopts closed-loop logic, and the control period is the same as the feedback acquisition frequency, both being 1 Hz.
[0179] The closed-loop control logic includes: if the path does not deviate, keep running the current trajectory; if the path offset is between 1.2 mm and 2.0 mm, fine-tune the motion direction and speed parameters of the end effector to perform dynamic deviation correction; if the path offset ≥ 2.0 mm, trigger the path backtracking mechanism and pause the execution, and resume the operation after the path replanning is completed.
[0180] The feedback control and adjustment module generates a feedback status report at the end of each control period. The report content includes: the current spatial position and attitude angle of the end effector, the predicted offsets and structure labels of each structural annotation point, the path offset status, and the execution progress and remaining path point information of the current path segment.
[0181] In some specific embodiments, the recognition and confirmation module specifically includes:
[0182] Configure the target confirmation comparison set, perform target error evaluation, and conduct tolerance judgment.
[0183] Call the coordinates (X T , Y T , Z T ) of the annotation points calibrated as the target structure from the structural annotation point set, extract their structure labels and confidence information, and load the corresponding reference attitude angle values α T , β T , γ T .
[0184] At the same time, call the current state of the end effector output by the feedback control adjustment module to obtain the real-time spatial position (X E , Y E , Z E ) and the attitude angle values (α E , β E , γ E ), and form a target confirmation comparison set with the target structure and the current execution state as the input for the joint evaluation of position error and attitude error.
[0185] Perform Euclidean distance calculation on the current position of the actuator and the coordinates of the target structure. The error calculation is specifically as follows:
[0186]
[0187] In the formula: D pos is the spatial position error value, representing the three-dimensional straight-line distance between the current position of the end effector of the neurosurgical operation assistance robot and the position of the target structure, with the unit of millimeter. X T , Y T , Z T are the three-dimensional space coordinates of the target structure provided by the set of structure annotation points. X E , Y E , Z E are the current three-dimensional space coordinates of the end effector provided in real time by the feedback control adjustment module.
[0188] At the same time, calculate the attitude angle error. Use the average angular difference of three axes to define the attitude error D att , specifically as follows:
[0189]
[0190] In the formula: D att is the attitude angle error value, representing the average difference between the reference attitude of the target structure and the current attitude of the end effector, with the unit of angle. α T , β T , γ T are the reference attitude angles of the target structure on three axes (X, Y, Z) given by the structure index binding table. α E , β E , γ E are the current attitude angles of the end effector on three axes (X, Y, Z) provided in real time by the feedback control adjustment module. |·| is the absolute value symbol used to calculate the difference between two angles, and / 3 represents taking the average of the differences of three axes.
[0191] Set the tolerance judgment condition as: If D pos ≤1.5mm and D attIf the angle is less than 5°, the target position is determined to be confirmed successfully, otherwise the actuator enters the end position fine-tuning state and notifies the feedback control adjustment module to regenerate the end adjustment path.
[0192] Generate standard control instructions and send them to access the intraoperative doctor confirmation interface.
[0193] After the target position is confirmed successfully, the recognition and confirmation output module converts the spatial information of the structure annotation point, the end point of the path point sequence and the posture state data into a standard control instruction format. The control instruction fields include: the three-dimensional coordinates of the execution target, the attitude angle of the actuator end, the operation type, the recommended execution speed value, and the control path point index number and channel identification code.
[0194] The control instructions are encoded by the identification and confirmation output module according to the control bus protocol and written into the control instruction buffer area for scheduling by the neurosurgery auxiliary robot motion execution module.
[0195] After the recognition and confirmation output module completes the control instruction encoding, it sends the control frame to the execution system through the robot control bus, starts the final path segment control logic, and activates the doctor's interactive interface in the recognition and confirmation output module. The interactive interface displays the following content: the current target structure name, spatial coordinates and recognition confidence, real-time preview of the actuator's current position and posture status, system judgment error and set tolerance interval, and target confirmation, manual correction and pause execution operation options.
[0196] Complete manual confirmation in the intraoperative confirmation interface or perform terminal fine-tuning based on image information. If there is no manual intervention, the system automatically enters the execution control state, the identification confirmation output module records the target state locking time and switches to the control scheduling standby state.
[0197] In the above content, in actual application, first, the structured light image acquisition system and the near-infrared image acquisition system are integrated at the end of the neurosurgery auxiliary robot, and the equipment parameter initialization and spatial calibration are completed. The structured light image acquisition system completes multi-angle reflection image acquisition by projecting stripe coding patterns and combining with the binocular camera module. The near-infrared image acquisition system completes the real-time acquisition of vascular path and nerve direction images in the 800nm to 1100nm band. The structured light images and infrared images collected by the system are pre-processed to complete distortion correction, illumination balance, artifact removal and edge enhancement processing, and are input into the time and space synchronization module in a unified resolution format for image registration and fusion.
[0198] Next, the fused image data generates a three-dimensional point cloud model of the intraoperative brain tissue via a multi-view geometry reconstruction algorithm. The Poisson surface reconstruction algorithm is used to generate a dense triangular mesh surface, and boundary constraint and normalization processing are performed. At the same time, preoperative MRI and CT images are called to construct a preoperative intracranial structure feature map. The system completes the spatial registration of the intraoperative and preoperative models after SURF feature point matching and RANSAC outlier rejection. Coordinate alignment is achieved through ICP fine registration. When the registration error meets the requirement of not exceeding 1.0 mm, a fused anatomical space model is generated and integrated into a unified coordinate reference system.
[0199] Subsequently, after the fused anatomical space model is normalized into a multi-channel network input tensor, it is input into a multi-channel convolutional neural network model with a ResNet structure. The model outputs a structure label map and a recognition confidence heat map, and the structure labels are semantically normalized and mapped through a neuroanatomical atlas. The system automatically constructs a structure index binding table, which includes structure labels, spatial coordinates, and confidence values. If the recognition confidence is lower than 85%, the local region re-identification mechanism is automatically triggered. Finally, the system outputs a set of structure annotation points and uses it as the basic data for path planning and target recognition.
[0200] After that, the three-dimensional coordinates, structure labels, and pose angle information of the target tissue region are identified in the set of structure annotation points. Combining the current pose state of the end effector and the positions of high-risk structures identified in the structure index binding table, it is input into a deep reinforcement learning policy network. The network generates a sequence of candidate path points and completes path optimization through the A* search algorithm and the DLite dynamic path correction algorithm. The sequence of path points generates a continuous execution trajectory by B-spline interpolation. The trajectory data is cached in the path planning buffer, and a path topology index table is established for synchronization control and path backtracking mechanisms.
[0201] Then, the feedback control adjustment module collects the spatial state of the end effector, visual images, and the set of structure annotation points per second, and performs state prediction through a Kalman filter. The system calculates the structure offset and determines whether the path offset degree exceeds 1.2 mm. If it exceeds the limit, a path backtracking request is sent to the path planning module. The system backtracks to the previous stable path point and regenerates a local path under the control of the path backtracking mechanism. At the same time, the feedback status report records the end effector pose state, structure offset data, and the execution progress of the current path segment, providing a decision basis for subsequent path control.
[0202] Finally, the system constructs a target confirmation comparison set from the structural marking points calibrated as targets and the real-time state of the end effector, calculates the spatial error value and the average difference in the three-axis attitude. If the confirmation conditions within the set tolerance range are met, the target structure coordinates, attitude angles, path point numbers and execution speed values are encapsulated into the standard control instruction format and sent to the execution system via the control bus. At the same time, the target structure attributes and confirmation options are displayed on the doctor's interaction interface. If no manual correction instruction is triggered, the system automatically enters the path execution state.
Claims
1. An AI-based visual positioning system for a neurosurgical operation assisting robot, characterized in that, It includes the following steps: The visual acquisition module is used for the deployment and initialization of the structured light scanning device and the near-infrared imaging device, and performs multi-modal image acquisition of the surgical area, and performs standardization processing, spatio-temporal alignment and fusion on the acquired images; The registration and fusion module is used for configuring the intraoperative three-dimensional point cloud model and three-dimensional grid, configuring the preoperative intracranial structure feature map and spatial registration, and performing registration error judgment and adaptive optimization mechanism; The feature recognition module is used for loading the fused anatomical space model and standardizing it into a network input tensor, loading a multi-channel convolutional neural network model, and performing structure recognition, performing semantic mapping of the recognition result and the standard neuroanatomical atlas, and configuring a structure index binding table and a confidence judgment mechanism; The path planning module is used for calibrating the target tissue area, configuring path input parameters, configuring a deep reinforcement learning policy network, generating a candidate path sequence, performing path optimization and continuous trajectory generation, configuring a path fallback mechanism, and dynamically triggering path replanning; The control and adjustment module is used for configuring an intraoperative action perception feedback mechanism and a state prediction model, and performing path deviation recognition and dynamic adjustment mechanism, configuring an actuator closed-loop control strategy, and generating a feedback status report; The recognition and confirmation module is used for configuring a target confirmation comparison set, performing target error evaluation, performing tolerance judgment, generating a standard control instruction, and sending the instruction to access the intraoperative doctor confirmation interface; 2. The AI-based visual positioning system for a neurosurgical operation assisting robot according to claim 1, characterized in that, The visual acquisition module specifically includes: The deployment and initialization of the structured light scanning device and the near-infrared imaging device, and the multi-modal image acquisition of the surgical area; The structured light scanning device and the near-infrared imaging device are fixedly installed on the end effector of the neurosurgical operation assistance robot, and device parameter initialization and spatial calibration operations are respectively performed on the two types of imaging devices; The structured light scanning device uses a structured light projector with stripe coding projection ability and a binocular camera module to form a structured light image acquisition system; The near-infrared imaging device uses a short-wave near-infrared camera with a working band of 800nm to 1100nm, and has dynamic gain adjustment and automatic exposure control functions for near-infrared image acquisition of blood vessel paths and nerve directions; The visual acquisition and processing module is started to control the structured light image acquisition system to project a stripe coding light map onto the surgical area, and the binocular camera module is used to synchronously acquire reflected image data from multiple perspectives; The near-infrared imaging system is started synchronously to perform real-time infrared image acquisition of the same surgical area, and the surface nerve and blood vessel structure image information is obtained. The acquisition process is completed under the control of a unified timestamp; The acquired images are subjected to standardization processing, spatio-temporal alignment and fusion; Image preprocessing operations are respectively performed on the acquired structured light images and infrared images. The preprocessing includes: distortion correction processing, illumination equalization processing, artifact removal processing and edge enhancement processing; The distortion correction process comprehensively corrects the image distortion caused by lens distortion using the back-projection method according to the device calibration results. The illumination equalization process performs histogram equalization on the image brightness distribution to enhance the image contrast. The artifact removal process removes the noise and light spots generated by acquisition interference using a combined algorithm of median filtering and edge-preserving smoothing filtering. The edge enhancement process performs superposition processing of the Sobel edge extraction algorithm and the non-linear enhancement algorithm on the structural boundaries to improve the clarity of the key tissue contours; The preprocessed structured light image and infrared image are input into the spatio-temporal synchronization module to perform image alignment and fusion operations, including: spatial registration processing, temporal synchronization processing, and image fusion processing; For spatial registration processing, a rigid registration algorithm based on feature points is used to map the structured light image and the infrared image to a unified spatial coordinate system. For temporal synchronization processing, the two types of images are precisely aligned through frame-level timestamps. For image fusion processing, a multi-channel fusion strategy is adopted to construct a fused image tensor, where the structured light image is used to provide three-dimensional depth information, and the infrared image is used to provide soft tissue recognition features; The fused result will be output as a standard multi-modal image data block, and the fused result has multi-level feature dimensions such as depth, infrared, and texture.
3. The AI-based visual positioning system for a neurosurgical operation assistance robot according to claim 1, wherein The registration and fusion module specifically includes: Configure the intraoperative three-dimensional point cloud model and three-dimensional mesh; Input the structured light image collected by the structured light image acquisition system into the multi-view geometric reconstruction module to perform depth estimation operations based on disparity calculation, and restore the three-dimensional point coordinates through the disparity information between binocular image pairs; Calculate the three-dimensional space coordinate values corresponding to each pixel point, specifically: Where: Z is the depth coordinate of the pixel point in the camera coordinate system, with the unit of millimeter; X is the horizontal coordinate of the pixel point in the camera coordinate system, with the unit of millimeter; Y is the vertical coordinate of the pixel point in the camera coordinate system, with the unit of millimeter; f is the focal length of the camera, with the unit of pixel; B is the baseline length of the binocular cameras in the structured light image acquisition system, with the unit of millimeter; d is the disparity value of the pixel point in the binocular images, with the unit of pixel; u is the horizontal pixel position of the pixel point in the image; v is the vertical pixel position of the pixel point in the image; c x is the pixel coordinate of the image center in the horizontal direction; c y is the pixel coordinate of the image center in the vertical direction; Generate an intraoperative three-dimensional point cloud dataset according to the calculation results, and perform point cloud cropping and boundary constraint processing, and retain the valid point cloud data within the spatial range of the surgical area; Input the cropped intraoperative three-dimensional point cloud into the meshing processing unit in the reconstruction, registration and fusion module, and use the Poisson surface reconstruction algorithm to generate a continuous triangular mesh model for the point cloud data; The triangular mesh model is output in the STL format, and the Laplacian smoothing algorithm is executed for surface optimization of non-boundary points to suppress the depressions and cusps caused by discontinuous edges or noise during the reconstruction process; Configure the preoperative intracranial structure feature map and spatial registration, and perform registration error judgment and adaptive optimization mechanism; Call the preoperative magnetic resonance imaging image and computed tomography image, perform three-dimensional registration and resampling on the preoperative image in the image processing module, and set the resampling interval to 0.5mm; Use a convolutional neural network model based on the UNet structure to perform multi-class semantic segmentation on the preoperative MRI image, extract the ventricle, cortical area, and skull boundary structures, and generate an intracranial structure label map; Perform three-dimensional structure label encoding on the segmentation results, construct a preoperative structure feature index table, and uniformly embed the ventricle, cortical area, and skull boundary labels into the spatial reference coordinate system of the preoperative image to form a standard preoperative intracranial structure feature map; Perform spatial registration operations on the constructed intraoperative three-dimensional point cloud model and the preoperative intracranial structure feature map, and complete the coordinate alignment process using a rigid registration method based on feature point matching; The rigid registration process is as follows: The SURF algorithm is used to extract the key structural feature points in the intraoperative three-dimensional point cloud and the preoperative atlas. After configuring the set of matching point pairs, the RANSAC algorithm is used to eliminate the mismatched point pairs, generating an initial pose transformation estimate. The ICP algorithm is used for fine registration to solve the optimal rotation matrix R and translation vector t, satisfying the objective function of minimizing the registration error. The registration error is quantitatively calculated using the root mean square error RMSE, specifically as follows: Where: RMSE is the root mean square error during the preoperative and intraoperative model registration process, with the unit of millimeters; N is the number of effective feature points participating in the registration, is the three-dimensional coordinate of the i-th feature point in the intraoperative three-dimensional point cloud, with the unit of millimeters, is the three-dimensional coordinate corresponding to the i-th intraoperative point in the preoperative structural atlas, with the unit of millimeters; R is the three-dimensional rotation matrix for rigid registration, t is the three-dimensional translation vector for rigid registration, and · represents the Euclidean distance; Set the registration error tolerance threshold to 1.0 mm. When the RMSE value of the registration result exceeds this tolerance threshold, the system will automatically start the registration optimization mechanism and perform adaptive reconstruction operations, including: increasing the density of SURF feature point extraction, performing local resampling and sparse region interpolation on the intraoperative point cloud, adding an image grayscale histogram similarity judgment link to enhance the stability of point pair matching, and performing the initial matching and ICP fine registration process again until the RMSE meets the requirements; When the registration error satisfies the condition of ≤ 1.0 mm, the system fuses the intraoperative three-dimensional point cloud model and the preoperative intracranial structure feature atlas under a unified spatial reference coordinate system to generate a fused anatomical space model; The fused anatomical space model includes: the intraoperative brain tissue surface structure layer, the preoperative internal anatomical structure layer, and the structure index mapping layer.
4. The AI-based visual positioning system for neurosurgical operation assisting robot according to claim 1, wherein The feature recognition module specifically includes: Loading the fused anatomical space model and normalizing it into a network input tensor, loading a multi-channel convolutional neural network model, and performing structure recognition; Using the fused anatomical space model output by the reconstruction registration fusion module as input data. The fused anatomical space model includes the intraoperative brain tissue surface structure layer, the preoperative internal anatomical structure layer, and the structure index mapping layer; Through the data preprocessing module, format the fused anatomical space model into a standard input tensor with the tensor dimension set to C represents the number of image channels, including the depth information channel, the texture information channel, and the structure label channel; H, W, and D respectively represent the height, width, and depth of the tensor in the spatial coordinate system, with the unit of pixel; Normalize the tensor values to the [0,1][0,1][0,1] interval and bind the structure index mapping information; Call the multi-channel convolutional neural network model deployed in the feature recognition extraction module. The model structure is based on the ResNet backbone network, and the feature channels for anatomical structure classification are extended; After inputting the standard tensor, the output includes a structure label map and a recognition confidence heat map. The structure label map is: annotating the anatomical structure categories of brain sulci, gyri, ventricles, vascular bifurcation points, tumor boundaries, and skull edges. The recognition confidence heat map is: the value range at each position is [0,1][0,1][0,1], used to represent the credibility score of the structure label result; Perform the semantic mapping of the recognition result and the standard neuroanatomical atlas, and configure the structure index binding table and the confidence judgment mechanism; Input the structure label map into the anatomical semantic mapping module and perform a structural semantic comparison with the standard neuroanatomical atlas database. Map all recognized labels to the medical semantic standard structure names, and establish a standard term mapping table for the recognized brain sulci, gyri, ventricles, cortical regions, and vascular path structure types in the label map; If the offset of the recognition result from the center coordinates of the standard structure label exceeds 2.5 mm, the system records the recognition deviation mark; According to the structure label map and the recognition confidence heat map, construct a structure index binding table, bind the structure name, spatial coordinates, and recognition confidence value to each recognized structure point. If the recognition confidence of a certain structure point is lower than 85%, the feature recognition and extraction module automatically triggers the local area re-recognition process, re-intercepts the image sub-block of this area, and executes the structure recognition and inference operation; Generate a set of structure annotation points for all structure recognition results that pass the confidence threshold judgment. The set of structure annotation points includes: structure label, three-dimensional coordinates, structure number, and confidence score. The set data format is jointly output in PLY format and the JSON format of the structure attribute index table.
5. The AI-based visual positioning system for a neurosurgical operation assisting robot according to claim 1, wherein The path planning module specifically includes: Calibrate the target tissue area and configure the path input parameters; Calibrate the target tissue area in the set of structure annotation points, including: the tumor core, the lesion edge, and the resection target structure set by the doctor before the operation. Extract the three-dimensional spatial coordinates, structure label, and recognition confidence information of the target tissue from the structure index binding table; The set of input parameters required for path planning is configured respectively as: the spatial coordinates of the target tissue, the set of spatial coordinates of the structures recognized as high-risk structures in the structure index binding table, the current position and attitude information of the end effector of the neurosurgical operation assistance robot, including: position vector and attitude angle, the structure arm length parameter, motion range constraint, attitude change speed, and execution step accuracy of the neurosurgical operation assistance robot; Configure the deep reinforcement learning policy network and generate a candidate path sequence; Call the deep reinforcement learning policy network in the positioning path planning module. The policy network includes a state input layer, a policy generation layer, and a value evaluation layer; The state input layer is used to input the current motion state of the neurosurgical operation assistance robot, the target tissue position, and the distribution of dangerous structures. The policy generation layer uses the Actor network to generate candidate path points. The value evaluation layer uses the Critic network to calculate the cumulative return value of the path policy. The reinforcement learning goal is to maximize the cumulative reward function R; The closer the end position of the robot is to the target tissue, the positive reward is given. If the path is close to the dangerous structure or violates the operation constraints, the negative penalty is given. If the path is smooth and has high attitude stability, the gain reward is given; Output path point sequence Among them, path point P i includes: spatial position vector corresponding motion direction vector recommended execution speed value V i , with the unit of millimeters per second; Execute path optimization and continuous trajectory generation, configure the path backtracking mechanism, and dynamically trigger path replanning; Input the path point sequence output by the deep reinforcement learning policy network into the path optimization module, and use a multi-level path optimization algorithm combination to achieve fine correction; The multi-level path optimization algorithm includes: the heuristic graph search algorithm A, the dynamic path correction uses the DLite algorithm, and the B-spline curve interpolation method; The heuristic graph search algorithm A is used for initial path screening, eliminating the path segments that cross the high-risk structure area in the path. The dynamic path correction uses the DLite algorithm to make the path respond to the real-time changes in the intraoperative environment. The continuous path generation uses the B-spline curve interpolation method to smooth the path point sequence and generate a continuous executable robot execution trajectory; The final path sequence is cached in the path planning buffer, and a path topology index table is established for path state tracking and dynamic instruction synchronization control; The path planning module continuously receives the structural annotation point coordinate status information and the real-time motion status of the end effector transmitted by the feedback control and adjustment module, and performs path stability judgment; When any of the conditions is met, the path backtracking mechanism is triggered, including: the spatial position change of any structural annotation point exceeds 2.0 mm, a new high-risk structure label recognition result appears within the path planning area, the current position of the end effector of the neurosurgical operation assistance robot deviates from the current path segment by more than the preset threshold, or the attitude angle deviation exceeds 10 degrees.
6. The AI-based visual positioning system for a neurosurgical operation assistance robot according to claim 1, wherein, The control and adjustment module specifically includes: Configure the intraoperative action perception feedback mechanism and the state prediction model, and execute the path deviation recognition and dynamic adjustment mechanism; Configure the real-time data interface between the feedback control and adjustment module and the visual acquisition and processing module, and perform periodic acquisition and synchronous update of information. The periodic acquisition and synchronous update of information are: the spatial position, attitude angle, and execution speed information of the end effector of the neurosurgical operation assistance robot, the current path point sequence and the path topology index table, the structured light image and the near-infrared image output by the visual acquisition and processing module, and the set of structural annotation points and the structure label map generated by the feature recognition and extraction module; Set the acquisition frequency to 1 Hz, and manage the acquisition data through a unified time stamp; Call the state prediction sub-module in the feedback control and adjustment module, use the Kalman filter to predict the spatial position of the end effector of the neurosurgical operation assistance robot, and combine the set of structural annotation points to identify abnormal offsets; The structure offset recognition logic is: compare the current spatial coordinates of the structural annotation points with the predicted coordinates, calculate the offset. If the offset of any structural annotation point exceeds 1.2 mm, it is recorded as a structure offset event, add the structure category information to the offset event, and include it in the path correction judgment parameter set; Taking the structure offset event as the input, calculate the path offset degree D in combination with the path point sequence p ; When the maximum path offset D p ≥ 1.2 mm, the system immediately activates the path fallback mechanism. The feedback control adjustment module sends a path update request to the positioning path planning module and regenerates the target path segment; Configure the actuator closed-loop control strategy and generate a feedback status report; According to the path correction result, update the motion instruction parameters of the end effector of the neurosurgical operation assistance robot to make the motion trajectory consistent with the latest path. The execution control uses a closed-loop logic, and the control period is the same as the feedback acquisition frequency, both being 1 Hz; The closed-loop control logic includes: if the path does not deviate, keep the current trajectory running; if the path offset is between 1.2 mm and 2.0 mm, fine-tune the motion direction and speed parameters of the end effector to perform dynamic deviation correction operations; if the path offset ≥ 2.0 mm, trigger the path backtracking mechanism and suspend the execution, and resume the operation after the path replanning is completed; The feedback control and adjustment module generates a feedback status report at the end of each control cycle. The report content includes: the current spatial position and attitude angle of the end effector, the predicted offsets and structure labels of each structural annotation point, the path offset status, and the execution progress of the current path segment and the remaining path point information.
7. The AI-based visual positioning system for a neurosurgical operation assisting robot according to claim 1, characterized in that, The recognition and confirmation module specifically includes: Configure the target confirmation comparison set, perform target error evaluation, and perform tolerance judgment; Call the coordinates of the marked points calibrated as the target structure from the set of structure marking points Extract its structure label and confidence information, and load the corresponding reference attitude angle values α T , β T , γ T ; Simultaneously call the current state of the end effector output by the feedback control adjustment module to obtain the real-time spatial position and the attitude angle value Construct a target confirmation comparison set with the target structure and the current execution state as the input for the joint evaluation of position error and attitude error; Perform Euclidean distance calculation on the current position of the actuator and the target structure coordinates. The error calculation is specifically: Where: D pos is the spatial position error value, representing the three-dimensional straight-line distance between the current position of the end effector of the neurosurgical operation assisting robot and the position of the target structure, with the unit of millimeter, X T , Y T , Z T are the three-dimensional spatial coordinates of the target structure, provided by the set of structure annotation points, X E , Y E , Z E are the current three-dimensional spatial coordinates of the end effector, provided in real time by the feedback control adjustment module; Meanwhile, calculate the attitude angle error, and define the attitude error D using the average angular difference of three axes att ; Set the tolerance judgment condition as: If D pos ≤ 1.5 mm and D att ≤ 5°, it is determined that the target position confirmation is successful; otherwise, the actuator enters the fine-tuning state of the end position and notifies the feedback control adjustment module to regenerate the end adjustment path. Generate a standard control instruction and send the instruction to access the intraoperative doctor confirmation interface; After the target position is successfully confirmed, the recognition and confirmation output module combines the spatial information of the structural annotation points, the end point of the path point sequence, and the pose state data and converts them into the standard control instruction format. The control instruction fields include: the execution target three-dimensional coordinates, the end-effector pose angle, the operation type, the recommended execution speed value, and the control path point index number and channel identification code; The control instruction is encoded by the recognition and confirmation output module according to the control bus protocol and written into the control instruction buffer for scheduling by the motion execution module of the neurosurgical operation assistant robot; After the recognition and confirmation output module completes the control instruction encoding, it sends the control frame to the execution system through the robot control bus, starts the control logic of the final path segment, and at the same time activates the doctor interaction interface in the recognition and confirmation output module. The content displayed on the interaction interface includes: the current target structure name, spatial coordinates, and recognition confidence, the real-time preview of the current position and pose state of the end-effector, the system-determined error and the set tolerance interval, and the operation options of target confirmation, manual correction, and pause execution.
Citation Information
Cited By
Dental operation plan generation method and device, computer device and storage device
CN120581147A
Method and device for adjusting X-ray imaging equipment to standard positive and lateral positions of spine centrum
CN120643236A
CT interventional puncture positioning method and puncture positioning system based on artificial intelligence
CN120694731A
AI visual positioning method and system for robot automatic assembly
CN120839797A
Precision machining system and method for pore structure of three-component parallel composite fiber spinneret plate
CN120985420A