Temporal bone surgery navigation system in mixed reality environment
By constructing a temporal bone surgical navigation system using mixed reality technology, the problems of information fragmentation and insufficient risk identification in traditional temporal bone surgery have been solved. It enables accurate identification of the intraoperative space and real-time display of risk areas, significantly reducing the incidence of surgical complications.
Patent Information
- Application Number
- CN202511126593.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-18
AI Technical Summary
Traditional temporal bone surgery suffers from problems such as information fragmentation, large spatial cognition errors, and unclear identification of risk structures, leading to serious complications such as facial nerve injury, inner ear injury, or skull base hemorrhage. Existing systems lack dynamic assessment and navigation optimization of the distance between surgical instruments and high-risk anatomical areas.
A temporal bone surgical navigation system in a mixed reality environment is adopted, including a 3D reconstruction module, a risk analysis module, a risk distance calculation module, and an intelligent navigation optimization module. It constructs a virtual anatomical model through multimodal medical image fusion, calculates the distance between instruments and risk areas in real time, optimizes the movement trajectory of instruments, and provides multimodal feedback perception.
It enables precise identification of the intraoperative space and intuitive display of risk areas, reducing the probability of misoperation, improving the safety and efficiency of surgery, and reducing the risk of facial nerve injury and bleeding.
Smart Images

Figure CN120959891A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of instrument path planning technology, and in particular to a temporal bone surgical navigation system in a mixed reality environment. Background Technology
[0002] With the continuous development of medical imaging technology, neurosurgical techniques, and information technology, Mixed Reality (MR), an advanced interactive technology that integrates the features of Virtual Reality (VR) and Augmented Reality (AR), is gradually being applied in the field of clinical medicine, particularly showing great promise in high-risk, delicate craniomaxillary and otological surgeries. The temporal bone region, as one of the most complex and densely packed areas of the skull, integrates critical life structures such as the auditory organs, vestibular system, facial nerve, and jugular bulb, placing extremely high demands on precise intraoperative manipulation. Therefore, in temporal bone surgery, achieving accurate identification of surgical field structures, safe planning of surgical instrument operation paths, and risk avoidance of critical tissues have become pressing challenges that need to be overcome.
[0003] Traditional temporal bone surgery relies primarily on preoperative two-dimensional image analysis, intraoperative microscopic or endoscopic visual assistance, and the surgeon's extensive experience. However, due to complex anatomical structures, limited spatial perception, and narrow surgical fields, traditional navigation methods often suffer from information fragmentation, large spatial perception errors, and unclear identification of risk structures. These problems can lead to serious complications such as facial nerve injury, inner ear injury, or skull base hemorrhage during actual surgery, directly impacting postoperative functional recovery and patient safety. Some navigation systems attempt to improve positioning accuracy by fusing preoperative image reconstruction with intraoperative tracking, but most remain limited to two-dimensional interface displays, lacking real-time interactive support for surgeon's operational perception and failing to effectively address the coordination between intraoperative spatial positioning and risk structure identification. Furthermore, existing systems generally lack dynamic assessment and navigation optimization functions for the distance between surgical instruments and high-risk anatomical areas, still relying on surgeon judgment and posing a risk of operational errors. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention proposes a temporal bone surgical navigation system in a mixed reality environment, thereby resolving at least one of the aforementioned technical issues.
[0005] To achieve the above objectives, the present invention provides a temporal bone surgical navigation system in a mixed reality environment. The mixed reality environment temporal bone surgical navigation system includes a three-dimensional reconstruction module, a risk analysis module, a risk distance calculation module, and an intelligent navigation optimization module. The three-dimensional reconstruction module is used to acquire multimodal medical images, perform voxel-level deep fusion and three-dimensional point cloud geometric reconstruction, and construct a virtual anatomical model of the temporal bone. The risk analysis module is used to analyze the damage risk propagation path of the temporal bone virtual anatomical model and obtain a visualized risk anatomical model. The risk distance calculation module is used to detect the real-time pose parameters of surgical instruments, calculate the risk area distance of the visualized risk anatomy model, and extract the risk area distance parameters. The intelligent navigation optimization module is used to adjust the movement trajectory parameters of the device and optimize the intelligent device navigation based on the distance parameters of the risk area.
[0006] The beneficial effects of this invention are as follows: By fusing multimodal images such as CT (providing bone details) and MRI (providing soft tissue contrast), more comprehensive information on the temporal bone and adjacent tissue structures can be obtained. Voxel-level deep fusion makes the model more accurate in spatial resolution and density, and the geometric reconstruction of three-dimensional point clouds can realistically reflect the individual anatomical characteristics of the patient. A risk propagation map is constructed through the spatial relationships between physiological structures, lesion distribution, and neurovascular pathways to predict the damage chain that may be caused by misoperation. Risk areas are visually presented in the three-dimensional model using color, transparency, etc., which helps surgeons quickly identify key high-risk areas (such as the facial nerve and semicircular canals). Customized risk maps are provided for different lesion types and individual structural differences, assisting in the development of precise approach plans. The three-dimensional position and orientation of instruments are acquired in real time through trackers or visual sensors, and the minimum distance between them and risk structures can be dynamically calculated. Once an instrument approaches a high-risk area (such as the facial nerve or jugular bulb), the system can issue an early warning or impose operational restrictions to reduce the probability of accidental injury. Digital mapping of the intraoperative environment is achieved, providing accurate basic data for subsequent navigation and control strategies. Based on real-time risk distance feedback, the system automatically optimizes the instrument's movement trajectory (e.g., adjusting the drill path or clamping angle) to maintain a safe path. In a mixed reality environment, operators can perceive approaching risks through visual / tactile / auditory multimodal feedback, enhancing their operational confidence and reaction speed. Attached Figure Description
[0007] Figure 1 This is a schematic diagram of the structure of a temporal bone surgical navigation system in a mixed reality environment according to the present invention; Detailed Implementation It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0008] This application provides a temporal bone surgical navigation system in a mixed reality environment. The execution entities of the temporal bone surgical navigation system in the mixed reality environment include, but are not limited to, the following: mechanical equipment, data processing platform, cloud server node, network upload device, etc., which can be considered as general computing nodes of this application. The data processing platform includes, but is not limited to, at least one of the following: an audio-image management system, an information management system, and a cloud data management system.
[0009] Please see Figure 1 This invention provides a temporal bone surgical navigation system in a mixed reality environment, comprising the following steps: The three-dimensional reconstruction module is used to acquire multimodal medical images, perform voxel-level deep fusion and three-dimensional point cloud geometric reconstruction, and construct a virtual anatomical model of the temporal bone. The risk analysis module is used to analyze the damage risk propagation path of the temporal bone virtual anatomical model and obtain a visualized risk anatomical model. The risk distance calculation module is used to detect the real-time pose parameters of surgical instruments, calculate the risk area distance of the visualized risk anatomy model, and extract the risk area distance parameters. The intelligent navigation optimization module is used to adjust the movement trajectory parameters of the device and optimize the intelligent device navigation based on the distance parameters of the risk area.
[0010] In the embodiments of the present invention, see Figure 1 The diagram below illustrates the steps of a temporal bone surgical navigation system in a mixed reality environment according to the present invention. In this example, the steps of the temporal bone surgical navigation system in a mixed reality environment include: The three-dimensional reconstruction module is used to acquire multimodal medical images, perform voxel-level deep fusion and three-dimensional point cloud geometric reconstruction, and construct a virtual anatomical model of the temporal bone. In this embodiment, anatomical information is extracted from multimodal medical images (CT and MRI), and a high-precision virtual anatomical model of the temporal bone is constructed through registration, fusion, and reconstruction algorithms. The collected data mainly includes: (1) high-resolution temporal bone CT images, slice thickness 0.5 mm, matrix 512×512, bone window width 1500, window level 500; (2) T2-weighted cochlear MRI, slice thickness 0.8 mm, matrix 256×256, specifically for neural tissue recognition. All image data are uniformly converted to NIfTI format to improve processing efficiency. Rigid + non-rigid joint registration of CT and MRI is performed using the ANTs framework, and the registration error is controlled within 0.6 mm. Then, the multimodal fusion network (MMF-Net) is used to perform voxel-level feature fusion of CT and MRI to achieve complementary enhancement of bony structure and neural tissue information. In terms of segmentation processing, 3D U-Net is used to segment the structure, and the network adopts channel attention mechanism to enhance the segmentation effect of small structures (such as facial nerve). Each voxel is assigned multiple labels (e.g., bone tissue, facial nerve, auditory nerve, blood vessels, etc.) for subsequent reconstruction and classification. The Marching Cubes algorithm is used to extract isosurfaces from the segmentation mask to construct 3D surface models of each structure. Mesh simplification is performed using Quadric Error Metrics (QEM) to ensure rendering performance on the MR platform (model face count controlled to within 200,000). The spatial distribution of structural points is represented using point cloud methods (PCL library), forming a virtual temporal bone anatomical model with spatial attributes, structural labels, and texture information.
[0011] The risk analysis module is used to analyze the damage risk propagation path of the temporal bone virtual anatomical model and obtain a visualized risk anatomical model. In this embodiment, a virtual anatomical model of the temporal bone is used to assess potential intraoperative injury paths, identify high-risk areas, and construct a visual risk anatomical model with dynamic feedback capabilities. This module first extracts the spatial course of all key neurovascular structures based on previous structural segmentation and geometric modeling results, and uniformly extracts 10-15 key structural points along the center line of each structure. An adjacency graph is constructed, where each node is a structural point and edges are its spatial neighbors. Irrelevant connections are filtered by setting spatial distance thresholds (≤3 mm) and directional angles (≤45°), resulting in a structural adjacency matrix. A risk propagation model is then introduced, defining a propagation probability function by evaluating the likelihood of injury spread between structural points due to their location, orientation, and relationship to safety thresholds. This function integrates the inverse spatial distance ratio, the cosine function of the angle, and a threshold mutation term to construct a weighted directed graph. A multi-source shortest path algorithm (such as a hybrid optimization of Floyd-Warshall and Dijkstra) is used to calculate the potential injury path, including first-, second-, and third-order propagation relationships, starting from any structural point. Each structural point is assigned a risk score, representing the sum of its propagation weights as a source or receiver. Based on this, the risk level of each structural point is mapped onto a 3D model for visualization. Low-risk areas are displayed in green, high-risk areas in red, with a transitional yellow. This risk visualization model supports interactive integration with mixed reality terminals (such as HoloLens 2), allowing surgeons to view the propagation path and risk level of structures through gestures, providing intuitive support for surgical path selection and damage avoidance.
[0012] The risk distance calculation module is used to detect the real-time pose parameters of surgical instruments, calculate the risk area distance of the visualized risk anatomy model, and extract the risk area distance parameters. In this embodiment, the risk distance calculation module aims to monitor the real-time spatial relationship between surgical instruments and high-risk structures, dynamically identifying whether surgical operations touch or approach high-risk areas, thereby achieving risk warning. First, this module integrates an optical tracking system (such as NDI Polaris Vega) to perform 6-DOF pose detection on the instruments during surgery (spatial resolution 0.25 mm, refresh rate 60 Hz). Reflective markings on the instruments are mapped to the model space through a preset calibration matrix, achieving spatial alignment with the virtual anatomical model. After the pose information is synchronized, the system performs real-time distance calculation between the current instrument tip position information and the structural point cloud in the anatomical model. The calculation method uses the KD-Tree spatial indexing algorithm, controlling the distance calculation delay to within 15 ms to ensure real-time performance. When the spatial distance is less than the corresponding safe distance threshold for each structure (e.g., 1.8 mm for the facial nerve, 2.0 mm for the internal carotid artery), the system will issue a color change and sound prompt in the visualization interface and record the distance parameters at the current time point. To support concurrent distance assessment across multiple structures, the system simultaneously maintains a "risk distance vector field," representing the shortest distance between the instrument position and each risk structure, along with its directional derivative. This facilitates subsequent judgment of whether the instrument's movement trend is "oriented" towards high-risk structures. These distance parameters are not only used to inform the system but are also fed into the next module for trajectory optimization. In clinical trial simulations, the system can monitor the distance between the instrument and seven structures in real time on average, with an error controlled within ±0.3 mm, sufficient to meet the micrometer-level precision requirements of otological surgery.
[0013] The intelligent navigation optimization module is used to adjust the movement trajectory parameters of the device and optimize the intelligent device navigation based on the distance parameters of the risk area.
[0014] In this embodiment, based on real-time sensing of the instrument's pose and distance to the risk area, an intelligent algorithm analyzes and adjusts the instrument's operation path to prevent the operator from accidentally entering a high-risk area, thus achieving intelligent assisted navigation. First, by receiving distance data streams from the risk distance calculation module, the system performs trajectory prediction analysis based on distance change trends. If the instrument tip moves towards a high-risk structure (judgment criteria: negative distance derivative, directional angle <30°), the system activates a navigation intervention mechanism. Intervention is divided into two types: prompt navigation and path correction. Prompt navigation displays a recommended path (such as a green semi-transparent channel) through the MR interface and provides a red dashed line warning when approaching a high-risk boundary. Path correction is activated when the operator selects "automatic path adjustment" mode. The system establishes a virtual movement cost map around the current instrument pose based on a dynamic constraint A search algorithm. The cost function comprehensively considers factors such as the distance to the risk structure, structural curvature, and historical trajectory inertia to output an optimal movement path. This path is displayed as a transparent pipe in the MR environment, guiding the operator along a safe path. To ensure the algorithm's real-time performance, the system sets the cost map resolution to 0.5 mm and the path planning time to within 100 ms. In a simulated otology experiment involving 10 cases, the proportion of instruments accidentally entering high-risk areas decreased by 70.4% after navigation optimization, and the average surgical path length decreased by 12.3%, demonstrating its significant clinical navigation assistance value. This module, as the core of the system's dynamic response, ensures that instruments remain on a navigation path that balances safety and efficiency even in complex temporal bone anatomy.
[0015] In this embodiment, the three-dimensional reconstruction module is used to acquire multimodal medical images, perform voxel-level deep fusion and three-dimensional point cloud geometric reconstruction, and construct a virtual anatomical model of the temporal bone. Specifically, it is used for: Acquire multimodal medical images, including preoperative high-resolution CT, MRI, and intraoperative real-time ultrasound images; Feature points are marked on the multimodal medical images, and spatiotemporal registration is performed to obtain image spatial registration parameters. Based on image spatial registration parameters, voxel-level deep fusion is performed, and three-dimensional point cloud geometric reconstruction is carried out to construct a digital twin model of the temporal bone. Pre-defined physical properties of the temporal bone are calibrated on the digital twin model of the temporal bone to obtain a virtual anatomical model of the temporal bone.
[0016] In this embodiment, preoperative and intraoperative multimodal medical imaging data are systematically acquired to comprehensively reflect the bony structure and soft tissue status of the temporal bone region. Preoperative high-resolution CT scans are performed using multi-slice spiral CT equipment, typically employing a 0.5 mm slice thickness and reconstruction interval to maximize the preservation of the fine anatomical structures of the temporal bone and acquire DICOM format image data. MRI scans are primarily used to visualize neural structures and soft tissues related to the temporal bone, such as the facial nerve, internal auditory canal, and semicircular canals. T1-weighted and T2-weighted sequences are used, with common parameters being TR 1500 ms and TE 80 ms, and resolution controlled within 1 mm to ensure subsequent registration accuracy. Intraoperative real-time ultrasound imaging is performed using a handheld high-frequency ultrasound probe (e.g., a 10-15 MHz linear probe), focusing on acquiring real-time structural changes in the surgical area, such as changes in the morphology of mastoid air cells, tympanic cavity structures, and displacement of blood vessels or nerve tissue. Ultrasound images are synchronously recorded using an image acquisition card and saved as timestamped sequence frames. The key to this step is ensuring high-quality acquisition of images from each modality in both space and time, laying a solid foundation for subsequent registration and fusion, while ensuring that all image data have unified patient identification and surgical area recognition labels. Due to differences in imaging mechanisms, multimodal medical images exhibit significant variations in grayscale distribution and tissue boundary representation, making direct spatial registration challenging. Therefore, this step employs a combined manual and automated feature point labeling strategy. For manual labeling, key structural points with anatomical consistency are selected, such as bony structures at the skull base (e.g., the cochlear window, external auditory canal, and internal auditory canal opening), meningeal structure edges, and the sigmoid sinus. Automatic labeling utilizes deep learning segmentation models (such as U-Net or nnUNet) to segment CT and MRI images into regions, and then the algorithm automatically extracts edge features and corner points. After labeling, preliminary rigid registration is performed using mutual information-based optimization algorithms (such as Mattes Mutual Information implemented in ITK) to obtain rotation and translation transformation matrices. Then, non-rigid registration is performed using free deformation models such as B-spline or the Demons algorithm, considering soft tissue deformation during surgery to generate the final spatiotemporal transformation field. Registration of intraoperative real-time ultrasound images requires a time synchronization mechanism to align its frame sequence with the time axes of MRI and CT images. Real-time spatial alignment is then achieved through image intensity correlation (such as Normalized CrossCorrelation) and feature point matching. The final output is a set of transformation matrices (4×4 homogeneous transformation matrices) from ultrasound, MRI, to CT space, along with registration error estimation metrics (such as TRE < 1.5 mm), serving as the spatial basis for subsequent fusion and modeling.
[0017] After registration and obtaining a unified spatial coordinate system, voxel-level fusion is required to integrate multimodal image information. Specifically, preoperative CT is used as the primary reference space, and MRI and ultrasound images are mapped to the CT space through the aforementioned transformation matrix. Voxel-level fusion employs a multi-scale fusion strategy, using CT to provide clear boundaries of bony structures, MRI to provide high-contrast signals of soft tissues, and ultrasound images to provide intraoperative structural dynamics. During fusion, a weighted average fusion method is used, with weights adaptively adjusted based on the signal-to-noise ratio and structural visibility of each modality in different regions. The fused image volume undergoes high-pass filtering and edge enhancement to improve structural clarity. In the 3D reconstruction stage, based on the fused image, the Marching Cubes algorithm is used to extract isosurfaces to generate a triangular mesh model. The extraction threshold is typically the grayscale threshold of bone tissue in CT (HU>300). The generated surface mesh is further constructed into dense, uniform point cloud data through point cloud resampling (point spacing approximately 0.2 mm) and geometric optimization (such as Laplacian smoothing). Point cloud data, through principal curvature analysis, identifies key anatomical structures, labeling the cochlear window, facial nerve canal, semicircular canals, sigmoid sinus, etc., providing an accurate foundation for subsequent simulation and navigation. This digital twin model realistically reflects the geometric features of the temporal bone structure and surrounding soft tissues in virtual space, serving as a crucial reference model in mixed reality environments.
[0018] After geometric reconstruction, the digital twin model needs to be endowed with physical properties to support subsequent functions such as mixed reality-based interaction, force feedback simulation, and surgical path planning. This step involves calibrating the model based on material properties to construct a complete virtual anatomical model. First, based on the relationship between grayscale values and bone density (HU value) in CT images, bone tissue is divided into compact bone (HU>1000), cancellous bone (HU 300-1000), and soft tissue (HU<300), each assigned a different elastic modulus and damping coefficient. For example, the elastic modulus of compact bone is set to 10 GPa, and that of cancellous bone to 1 GPa, with the damping coefficient ranging from 0.05 to 0.2 based on tissue analogy experiments. For soft tissue areas visible in MRI and ultrasound, such as the facial nerve, inner ear nerve, and meningeal boundaries, calibration is performed using soft tissue biomechanical model parameters from the literature, such as an elastic modulus range of 50 kPa to 300 kPa. During calibration, a finite element model (FEM) is used to physically simulate and initialize the model. This includes establishing a material property mapping field, constructing boundary conditions (such as sigmoid sinus wall fixation), and applying simulated stress points (such as drill bit force) to verify the model's response. Finally, the virtual anatomical model is dynamically verified against intraoperative ultrasound images to check whether the geometric changes of the anatomical structure under time and external forces are consistent with actual surgical observations, with the error controlled within 2 mm. After completing this step, the resulting virtual anatomical model not only possesses a high-fidelity geometric structure but also exhibits realistic physical property responses, making it a core component for interaction, simulation, and prediction in mixed reality surgical navigation systems.
[0019] In this embodiment, the risk analysis module is used to analyze the damage risk propagation path of the temporal bone virtual anatomical model to obtain a visualized risk anatomical model, specifically for: Semantic segmentation of key anatomical structures was performed on the virtual anatomical model of the temporal bone to extract the structures of the facial nerve, auditory nerve and internal carotid artery, and fitted into key temporal bone structural points. Geometric morphological analysis was performed on key temporal bone structural points to obtain their morphological characteristics. Spatial topological relationships are mined from the morphological features of structural points to obtain spatial adjacency relationships between structures; The structural safety distance threshold is calculated based on the spatial adjacency relationship between structures to obtain the structural safety distance threshold; Damage risk propagation path analysis is performed based on structural safety distance threshold to obtain the damage risk propagation network; Dynamic risk rendering of a virtual anatomical model of the temporal bone is performed based on a damage risk propagation network to construct a visualized risk anatomical model.
[0020] In this embodiment, neurovascular structures crucial for surgical navigation are extracted from the constructed virtual anatomical model of the temporal bone and converted into a structural point model, serving as the basis for subsequent spatial analysis and risk modeling. First, the temporal bone model is structurally segmented using a deep learning-based semantic segmentation model, and voxel classification of the fused image volume is performed using a 3D U-Net or nnU-Net structure. Training data comes from high-quality CT and MRI image datasets annotated by experts, and performance is evaluated using 5-fold cross-validation. In the segmentation labels, the facial nerve pathway, the auditory nerve path, and the course of the internal carotid artery in the temporal bone segment are treated as independent target categories. Segmentation accuracy is measured by the Dice coefficient; in the experiment, the facial nerve pathway reached 0.87, the auditory nerve 0.84, and the internal carotid artery 0.89, meeting the accuracy requirements for structural recognition in surgical navigation. After segmentation, the centerline point set of each structure is extracted, and the main structural path is obtained through a method based on skeletonization and center curve fitting, then discretized into a sequence of key structural points. The facial nerve pathway was fitted with 10-15 nodes, while the auditory nerve and internal carotid artery were each discretized into approximately 20 nodes. Each node was accompanied by its location's normal vector information and local surface curvature, providing a morphological basis for subsequent geometric analysis. The final output structural point data formed a semantically clear and spatially accurate set of structural points, constructing a high-level structural representation for the entire anatomical model. Geometric morphological analysis of the extracted key structural points comprehensively described their spatial structural characteristics, providing quantitative parameters for subsequent spatial relationship modeling and risk assessment. The analysis mainly included three types of morphological indicators: local geometric indicators, path geometric indicators, and tissue environment characteristics. Local geometric indicators, such as curvature, torsion, principal direction vector, and rate of change of tangent direction, reflected the morphological variations of the structure at the microscale. By constructing a K-neighborhood (K=12) on the point cloud structure, the least squares method was used to fit the local surface, calculating the normal direction and curvature tensor to obtain the curvature value and principal axis of curvature direction for each point. In the facial nerve pathway, particular attention was paid to the curvature changes as it traverses the skull base, as this is a high-risk area for surgical procedures. Path geometry indicators examine the overall direction of the structure, such as the total length of structural lines, integral of curvature, and average directional change angle. Taking the facial nerve as an example, its turning angle between the tympanic cavity and the petrous bone segment is a key analytical indicator, averaging 32° to 45°. Tissue environment characteristics are mainly analyzed from the perspective of the distribution and density changes of materials around the structure. For example, the density gradient information of the petrous bone around the internal carotid artery can be estimated through the rate of change of local CT grayscale values. Finally, these morphological features are vectorized to form a geometric descriptor for each structural point, providing a mathematical basis for spatial relationship modeling and risk propagation, and can be used for the design of adaptive navigation and obstacle avoidance algorithms in surgical pathways.
[0021] After constructing the morphological features of the structural points, it is necessary to analyze the topological relationships between these points in three-dimensional space to identify potential adjacency, overlap, or crossing risk areas. This step uses a method based on distance field analysis and spatial adjacency graph construction to model the mutual positional relationships between key structures in the temporal bone. First, the shortest Euclidean distance between each structural point and other structures is calculated, and an adjacency threshold (generally 2 mm, based on the safety boundary standards for minimally invasive surgery) is set to determine whether there is an adjacency relationship between points. In the experiment, a KNN (K-nearest neighbor) graph is constructed, and structural labels are introduced to form a directed adjacency graph, where nodes represent structural points and edge weights represent the shortest spatial distance. Further, a spatial relationship mesh is constructed using local Delaunay triangulation, and the spatial density and topological interweaving degree of different structural intersections are analyzed on the triangular facets, with particular attention paid to potential crossing areas between the facial nerve and the internal carotid artery and auditory nerve. In addition, the dot product of direction vectors is calculated to determine whether structural points have similar orientations, increasing alertness in anatomical intersection areas. The resulting spatial adjacency graph clearly expresses the proximity and spatial correlation between the structures, providing a precise data framework for subsequent safety distance analysis and risk path calculation.
[0022] The minimum safe operating distance between key structures was further calculated, and a safe buffer zone was constructed for surgical procedures. First, based on all connected structural pairs in the topological adjacency graph, their spatial distance distribution was statistically analyzed, and the probability density function of the distance between adjacent structures was obtained using the nonparametric kernel density estimation method (KDE). Taking the facial nerve and internal carotid artery as an example, if their spatial distance is less than 1.5 mm, it is considered a high-risk adjacency zone; this threshold was established with reference to existing clinical surgical guidelines. The minimum safe distance threshold within a 95% confidence interval was calculated for each pair of structures, forming a "structural safe distance matrix," where each element D(i,j) represents the recommended minimum operating distance between structure i and structure j. In the experimental dataset, the mean minimum safe distance between the facial nerve and auditory nerve was approximately 2.2 mm, with a standard deviation of 0.5 mm, while the safe distance threshold between the facial nerve and internal carotid artery needed to be controlled above 1.7 mm. When calculating the safe distance threshold, a curvature weight correction factor was also introduced to relax the distance restriction in high-curvature regions to account for the influence of spatial morphological complexity. Ultimately, this safety distance matrix is not only used for automatic obstacle avoidance prompts during real-time navigation, but can also be linked with mixed reality devices to achieve visual early warning of high-risk areas during surgery. Based on the established spatial relationships between structures and safety distance thresholds, this step constructs a damage risk propagation model for key skull structures, analyzing the propagation mechanism of potential injury events in three-dimensional space. A graph model is used, with key structural points as nodes and the connections between structures as edges. The risk propagation weights of the edges are defined comprehensively based on parameters such as inter-structural distance, safety thresholds, and tissue toughness. Specifically, when an operating area (such as a drill point) is located within the adjacent area of a structure, and the distance to the adjacent structure is below its safety threshold, it is considered a potential trigger for a "risk propagation event," which will propagate directionally to the connected structural points. A "risk potential function" is introduced into the model construction, combining tissue fragility (derived from tissue density and elastic modulus) and spatial distance to calculate the transmission intensity along the risk path. Through Dijkstra's algorithm and shortest propagation path analysis, the shortest high-risk path from any operating point to the key structure is identified. Model evaluation revealed multiple high-risk pathways between the facial nerve and the internal carotid artery near the tympanic segment of the facial nerve, with risk propagation values exceeding 0.75 (out of 1). The constructed injury risk propagation network can effectively predict potential injury areas under different operative pathways, serving as the core foundation for achieving dynamic intraoperative risk prediction and visual navigation prompts.
[0023] The rendering process first maps each structural point to a color gradient based on its node weight (i.e., cumulative risk value) in the risk propagation network, using a heatmap representation: red indicates high-risk areas, orange-yellow indicates medium-risk areas, and green indicates safe areas. The rendering algorithm employs GPU-accelerated volume rendering technology, combined with OpenXR and mixed reality device APIs to achieve real-time model updates and risk change responses. In the user's field of vision, as the position or angle of the operating tool changes, the risk heatmap on the model surface updates in real time, dynamically displaying the risk impact of the operation path on surrounding structures. To enhance interactivity, a transparency gradient and slight pulsating animation effect are overlaid in high-risk areas, focusing the user's attention on key areas. Combined with mixed reality platforms such as HoloLens 2, users can query the current risk value, risk source point, and recommended avoidance path for any structure via gestures or voice commands. In multi-center surgical simulation experiments, this risk model significantly improved the operator's efficiency in understanding complex structural relationships and risk control capabilities, reducing the average path planning time by 27% and the facial nerve injury simulation rate by nearly 40%. The final visualized risk anatomy model lays a key technological foundation for building an intelligent, dynamically responsive surgical navigation system.
[0024] In this embodiment, the risk distance calculation module is used to detect the real-time pose parameters of surgical instruments, calculate the risk area distance of the visualized risk anatomy model, and extract the risk area distance parameters. Specifically, it is used for: Detect the real-time pose parameters of surgical instruments and extract the six degrees of freedom motion data of the instruments; Based on the six-degree-of-freedom motion data, a virtual space coordinate transformation is performed to obtain the instrument position mapping parameters; Virtual instrument registration and correction are performed on the visualized risk anatomy model based on the instrument position mapping parameters to obtain the instrument registration anatomy model; The distance parameters between the instrument tip and the injury risk area are calculated based on the instrument registration anatomical model, and the distance parameters of the risk area are extracted.
[0025] In this embodiment, acquiring the real-time pose (position and orientation) information of surgical instruments in the intraoperative space is a crucial step in achieving interactive control and path evaluation in mixed reality surgical navigation. This step involves real-time tracking of the six degrees of freedom motion of the surgical instruments through an external positioning system, namely their three-axis translation (X, Y, Z) and three-axis rotation (Pitch, Yaw, Roll) information in three-dimensional space. Commonly used tracking technologies include optical infrared positioning systems (such as NDI Polaris Vega) and electromagnetic tracking systems (such as NDI Aurora). Optical systems offer high accuracy (error less than 0.25 mm) but require unobstructed vision, making them suitable for head surgery environments. A reflective marker ball or electromagnetic sensor is installed at the instrument tip to calibrate a fixed transformation matrix between it and the instrument tip. The system continuously samples at a frequency of approximately 60 Hz, capturing the instrument coordinate transformation matrix (4×4 homogeneous matrix) per frame and transmitting it to the navigation workstation in real time. To ensure data stability, a Kalman filter or exponential moving average is used to filter the pose sequence to remove high-frequency noise caused by vibrations, small-amplitude swaying, or system interference in the surgical environment. In addition, each type of instrument needs to be precisely calibrated once before surgery, including defining the coordinate origin (usually the instrument tip) and setting the direction vector to ensure accurate geometric mapping of the instrument in the virtual environment. This step lays the data foundation for subsequent virtual space mapping and anatomical model registration, realizing a real-time interactive closed loop between surgical instruments and digital twin models. After acquiring the six-DOF pose data of the instruments in real time, it needs to be transformed from the sensor coordinate system or positioning system coordinate system to the virtual space (i.e., the mixed reality rendering coordinate system where the digital twin model is located). This process is called virtual space coordinate transformation or spatial alignment. This step first clarifies the reference frame transformation relationship between the two coordinate systems, usually using a homogeneous coordinate transformation matrix T_world→virtual to map the instrument pose in the positioning system to the digital model in the mixed reality space. This transformation matrix comes from the preoperative multimodal image registration and equipment initialization calibration. After each frame is acquired, the instrument pose matrix T_tool is multiplied by T_world→virtual to obtain the instrument's pose T_virtual_tool in virtual space. During the conversion process, quaternion interpolation and Euler angle compensation mechanisms are employed to ensure the continuity of the rotation direction in high-frequency data, avoiding directional jumps or oscillations at the visualization end. In the experimental setup, the instrument position update frequency is approximately 30 frames per second (fps), combined with the spatial anchoring technology of a mixed reality rendering system (such as HoloLens 2) to achieve real-time display and interactive control at a 1:1 scale. After multiple simulations, the conversion error is controlled within 0.8 mm, meeting the clinical requirements for instrument path accuracy control (with an error tolerance limit of 1 mm) in temporal bone surgery.The completion of this step signifies the spatial synchronization of the physical surgical instruments with the model in the virtual environment, providing real-time support for subsequent risk area interactions and path analysis.
[0026] The instrument geometry model (CAD model or simplified mesh model) is spatially registered with the digital anatomical model to construct an instrument-registered anatomical model that can interact with the anatomical model. In this step, the three-dimensional geometry of the instrument is first modeled according to its physical prototype and functionally labeled according to instrument type (e.g., drill, suction tube, electrocautery hook). The instrument model coordinate system is rigidly registered with the coordinate system in the mixed reality anatomical model space through the aforementioned mapping parameter T_virtual_tool, so that the virtual instrument model accurately "covers" its real physical position in digital space. To further correct the interaction point position of the instrument tip, a "tip repositioning" strategy is introduced: combining preoperative instrument length calibration data and real-time feedback information during surgical operation (e.g., contact point color change or depth response), the instrument model is fine-tuned to achieve sub-millimeter registration accuracy. In the experimental evaluation, the virtual instrument is placed in the digital anatomical model for multi-point touch comparison, and the deviation of its touch position from the actual surgical instrument tip is recorded. The final average error is 0.72 mm. Once registration is complete, the virtual instrument model and the anatomical structure form a closed-loop interactive system. The system can determine their relative position, contact status, and movement trajectory in real time, and trigger feedback mechanisms in the mixed reality system (such as audio-visual prompts and vibration warnings), effectively enhancing the surgeon's operational perception and risk awareness in the mixed reality environment.
[0027] This system dynamically calculates the potential collision risk or accidental contact probability by assessing the spatial distance between the instrument tip and high-risk anatomical areas in real time. This step employs a method based on shortest path distance calculation and risk heatmap matching to extract spatial distance. First, the system extracts the current position coordinates P_tip of the instrument model's end point (usually the operating tip) in virtual space and calls the high-risk area labels (such as facial nerve, internal carotid artery, and auditory nerve) from a pre-built risk visualization anatomical model. In the 3D model, a KD-tree index structure is used to accelerate the search of the high-risk area point set, calculating the shortest Euclidean distance from P_tip to each point in the risk area, and taking the minimum value as the current risk distance d_min. To improve system response efficiency, a dynamic update threshold is set, re-triggering distance calculation when the instrument movement distance exceeds 0.1 mm or the orientation change exceeds 3°. To enhance clinical readability, the distance parameter is visualized as a real-time floating value or color intensity bar embedded in the mixed reality view, automatically turning red and triggering an audio-visual warning when the distance falls below the warning threshold (e.g., 1.5 mm). Furthermore, the system supports continuous sampling to generate "risk approach trajectories," recording the distance change curve of the instrument tip during surgery to assist in postoperative assessment and path optimization. In experimental simulations, real-time distance calculations were performed on five typical temporal bone surgical paths. The system's average response latency was 0.08 seconds, and the distance detection accuracy error was less than 0.6 mm, meeting the requirements for precise navigation support under complex skull base structure operations. This function plays a crucial protective role during navigation, providing surgeons with real-time, accurate spatial distance perception and accidental touch warning mechanisms.
[0028] In this embodiment, the intelligent navigation optimization module is used to adjust the device movement trajectory parameters and optimize intelligent device navigation based on the risk area distance parameter, specifically for: The temporal threat level is calculated based on the risk area distance parameter to generate a continuous security threat index; The color intensity rendering of risk areas is defined based on the continuous security threat index to obtain color rendering parameters; Dynamic rendering modulation of the instrument registration anatomical model is performed based on color rendering parameters to obtain a risk intensity rendering model; Multimodal early warning is based on a risk intensity rendering model, and real-time operation behavior records are collected, including equipment movement trajectory, dwell time and operation force information. Based on the real-time operation behavior recording information, the movement trajectory parameters are adjusted to construct the optimal instrument trajectory parameters; Intelligent instrument navigation optimization based on optimal instrument trajectory parameters.
[0029] In this embodiment, a time dimension is introduced to continuously model the threat state, thereby establishing a dynamic and time-series-based security threat assessment mechanism. Specifically, the system uses the spatial distance d(t) of the device tip as input, combined with its rate of change v(t) and acceleration a(t), to construct a "Continuous Threat Index" (CTI) model. The CTI is updated in each time frame (30 fps), taking into account factors such as the current distance value, the distance change trend, and the approach angle of the device tip. The threat index calculation model references clinical risk assessment standards, using a normalization method to standardize its value to between 0 and 1, where 0 represents no risk and 1 represents extremely high risk. For example, if the device continuously approaches the internal carotid artery at a relatively fast speed and remains within the safety threshold of 1.5 mm for more than 2 seconds, the CTI value will rapidly approach 1. To avoid false alarms due to instantaneous data fluctuations, a time-series window filtering mechanism (window length approximately 1.5 seconds) is introduced to achieve stable threat level judgment. The system also sets structure-specific weights based on the tissue sensitivity of different structures, such as the facial nerve area having a higher threat sensitivity coefficient than the sigmoid sinus area. In the simulation environment, CTI can achieve millisecond-level real-time response, and the assessment accuracy is more than 92% consistent with the results of manual annotation by experts. This threat index can serve as an important driving parameter for subsequent color rendering, multimodal warning, and path optimization. The threat level is presented in an intuitive way in the mixed reality visualization model, which enhances the surgeon's ability to perceive risk areas. This step defines the color rendering parameters by establishing a mapping relationship between threat level and color intensity. Specifically, a color mapping strategy based on linear gradient and perception optimization is adopted to map the CTI value (01) to the saturation and brightness parameters in the color HSV value. For example, when the CTI is below 0.3, it corresponds to the green range; 0.3-0.6 is the orange-yellow transition area; and above 0.6 is displayed as red with reduced transparency to simulate the "hot spot focusing" effect. In addition, to avoid color perception blind spots, the color intensity setting is combined with the Lab color space perception consistency correction function, so that different observers (including color-blind groups) can also effectively distinguish the risk level. When multiple high-risk structures are near the instrument tip, the system automatically overlays the color of the highest-risk structure and renders it locally in a magnified manner. The color update frequency is synchronized with CTI, updating 30 times per second with a latency of less than 60ms. In experimental verification, the rendering mapping error is less than 5%, and multiple surgical experts reported that this rendering mechanism can help identify high-risk areas during surgery more quickly, improving operational reaction time by approximately 18%. The final output color rendering parameters serve as input signals for the visual model, providing a key driver for subsequent mixed reality dynamic rendering control.
[0030] In a virtual surgical scenario, the anatomical model of the registered instruments will be dynamically modulated to form a "risk intensity rendering model" capable of expressing risk. The rendering modulation process employs a GPU-accelerated hybrid rendering algorithm based on voxel masking and surface mapping, adjusting the color, transparency, and texture details of key anatomical structures in real time to represent different levels of threat. Voxels near structural areas where the instrument tip approaches the structure are marked as "high-response areas" when the risk index increases. The system uses shader functions to locally highlight, material pulsation, and enhance the edges of these areas. For areas with a CTI exceeding 0.8, a dynamic ripple diffusion effect is added to guide the surgeon's attention to their relative position. The rendering system runs on a rendering engine integrated into a mixed reality platform (such as HoloLens 2), supporting synchronous updates with the real instrument positions, accurately reflecting the risk status of the areas pointed to or touched by the instruments on the virtual model. The system also incorporates an angle compensation mechanism to ensure consistent rendering intensity from different viewing angles, preventing distortion due to viewpoint changes. In a multi-center simulation experiment, the risk intensity rendering model effectively improved the surgeon's spatial recognition rate of high-risk structures by approximately 26% and reduced the probability of 5 cases of accidental facial nerve irritation during clinical navigation operations. This model constitutes the core visual presentation layer of the mixed reality risk feedback system.
[0031] A multimodal early warning mechanism will be built based on risk intensity rendering, and the surgeon's operation behavior will be recorded simultaneously to achieve the prevention and assessment of risk events. The system integrates visual, auditory, and tactile multimodal feedback. When the CTI exceeds a set threshold (e.g., 0.7), it automatically triggers voice prompts (e.g., "Caution: Approaching the facial nerve"), visual highlighting (red pulse effect), and tactile vibration feedback (via HoloLens or handheld devices) for comprehensive early warning. Simultaneously, the system records the instrument's movement trajectory 30 times per second, generating a 3D path line; dwell time is calculated by detecting the number of frames the instrument tip stays within a fixed radius (e.g., 1 mm); and operation force data is acquired through force sensors integrated into the instrument handle or operating table, recording the force value change curve per unit time. During the intraoperative data acquisition phase, all information is timestamped and synchronously transmitted to the backend database for analysis. In simulated operation evaluation, the system successfully captured 6 high-risk facial nerve operation events, issuing warnings 0.9 seconds in advance, with an average accuracy rate of 94%. The multimodal early warning mechanism significantly enhances the surgeon's operational awareness, while real-time behavior recording provides quantifiable evidence for postoperative path analysis and operational optimization. The instrument movement trajectory is discretized into a continuous three-dimensional coordinate sequence, and combined with behavioral tags such as dwell time and operational force to construct a trajectory feature vector. High-risk contact segments and redundant operation segments in the path are identified through overlay calculation with a structural risk intensity distribution map. The system uses a dynamic time warping algorithm to align and compare multiple surgeon operation paths, extracting the common minimum-risk operation path. Subsequently, a heuristic multi-objective optimization algorithm (such as a genetic algorithm or reinforcement learning path adjustment strategy) is introduced to iteratively optimize three core indicators—path curvature, total travel length, and dwell time in high-risk areas—without sacrificing accessibility. The optimal trajectory is defined as the operation scheme that maintains the minimum risk intensity integral curve under the shortest safe path. In evaluation tests, compared to the original path, the optimized trajectory reduces the average dwell time of surgical instruments in high-risk areas by 34%, shortens the total path length by approximately 21%, and effectively reduces the frequency of facial nerve approach.
[0032] In this embodiment, the specific steps for adjusting the movement trajectory parameters based on the real-time operation behavior recording information and constructing the optimal instrument trajectory parameters are as follows: The real-time operation behavior recording information is subjected to time-series device motion pattern mining to obtain the time-series device motion pattern; Based on the temporal movement pattern of the device, navigation decisions for risk avoidance devices are made, and a trajectory for risk avoidance devices is generated. Based on the trajectory of the risk avoidance device, a temporal movement pre-simulation of the risk intensity rendering model is performed to obtain device movement pre-simulation data. Based on the pre-simulation data of instrument movement, tissue deformation is predicted to obtain the tissue deformation and nerve compression status; Based on tissue deformation and nerve compression, the movement trajectory parameters are adjusted to construct the optimal instrument trajectory parameters.
[0033] In this embodiment, the continuously recorded instrument movement trajectory, dwell time, and force changes during surgery contain the surgeon's manipulation habits and behavioral response mechanisms in response to different anatomical environments. This step aims to extract regular and generalizable instrument movement patterns from this behavioral data through time-series data analysis techniques, forming a standardized time-series operation paradigm. First, the operational behavior data of each surgery is represented in time series form, constructing a multi-dimensional time vector flow, including position changes (X, Y, Z coordinates), posture changes (Pitch, Yaw, Roll), operation speed, and force intensity. A sliding window segmentation technique (window width set to 2 seconds, overlap rate 50%) is used to extract local behavioral units from the sequence, and a Hidden Markov Model (HMM) is applied to model the behavioral states, identifying typical action units such as "approaching high-risk structures—decelerating—pausing—fine-tuning and exiting" operation patterns. Further, temporal clustering algorithms (such as Time-series K-Means) were introduced to cluster and summarize instrument behavior trajectories in multiple surgical scenarios, forming a representative "temporal instrument motion pattern library." In experimental verification, 146 high-risk area operation behaviors extracted from 18 temporal bone surgeries were summarized into 5 typical temporal motion patterns, covering approximately 84% of operation variations. These patterns serve as an important reference basis for risk-aware navigation strategies and can be used for behavior prediction and intelligent scheduling of intraoperative navigation control systems. Based on the current real-time operation status, the motion pattern library is matched and identified to determine whether the current operation falls into a high-risk behavior paradigm (such as "continuous approach to the internal carotid artery area + no deceleration"). When the system determines a potentially risky operation, it immediately calls the "safety mode trajectory template" most similar to the current scenario and performs spatial adaptive adjustment in conjunction with the patient's individual digital anatomy model. The risk avoidance trajectory generation process comprehensively considers the following factors: first, the distribution of anatomical structures to ensure that the path avoids high-risk structures (such as the facial nerve); second, the continuity of movement and operating habits to avoid operator errors caused by trajectory changes; and third, the mechanical performance limitations of instruments, such as the maximum bending angle and operating space constraints. The system employs a cost function-based path optimization algorithm (such as a variant of Algorithm A) and introduces a "risk weight map" to minimize the overall risk intensity integral of the path as the optimization objective. The generated avoidance path is represented in space as a continuous, smooth, and reasonably avoidable three-dimensional curve. In simulation evaluation, the navigation system successfully generated a safe avoidance path using this mechanism, achieving a 96% success rate in avoiding high-risk areas of the facial nerve, while keeping the path length increase within 15%, balancing safety and surgical efficiency. The final output risk avoidance trajectory provides precise guidance for intraoperative dynamic navigation and mixed reality prompts.
[0034] By performing "temporal pre-simulation" of instrument trajectories in a virtual surgical scenario, the system analyzes the interaction between the instrument and anatomical structures and its dynamic response to risk areas. The system uses predetermined trajectory points as input to the instrument tip motion sequence, controlling the instrument model to move along a path in the virtual environment at a set speed (e.g., 5 mm / s). At each time frame, the system calculates in real-time the spatial distance, angular relationship, and CTI value changes between the instrument and surrounding structures (especially the facial nerve, auditory nerve, and internal carotid artery), and records the risk intensity curve traversed by the path in time-series form, forming "instrument movement pre-simulation data." Simultaneously, the system renders the interaction effects between the instrument model and risk structures in a 3D visual environment in real-time, including collision detection, highlighting of proximity areas, and dynamic shadows for path projection. To improve simulation efficiency, the system employs a layered pre-simulation mechanism, using frame skipping interpolation for low-risk areas and full-frame precision rendering for high-risk areas. In experimental testing, the module can complete a full time-series simulation of a 100 mm trajectory within 30 seconds and output a multimodal data package including distance curves, risk intensity change diagrams, and visual pre-show animations, providing decision support for subsequent organizational deformation and path adjustment.
[0035] This system predicts tissue deformation and nerve compression caused by device movement using a physical modeling and mechanical simulation approach. First, high-risk structures (such as the facial nerve and the adventitia of the internal carotid artery) are selected and meshed using finite element methods to establish a structure-mechanical response model. Tissue material parameters are set based on CT / MRI image grayscale and literature data; for example, the elastic modulus of the facial nerve sheath is approximately 40 kPa, and the Poisson's ratio is 0.45. Contact boundary conditions are established between the device contact surface and the structural model. Based on the changes in the device's position and orientation during the pre-simulated trajectory, the contact force field and stress distribution for each frame are calculated. The system employs a real-time finite element approximation algorithm based on quasi-static assumptions (such as Position-Based Dynamics, PBD) for rapid calculation, outputting the tissue deformation field and stress heatmap. Simultaneous simulation of multiple trajectory segments revealed a potential risk of injury when the facial nerve displacement exceeds 0.8 mm or the compressive stress exceeds 1200 Pa. Ultimately, the module outputs the maximum deformation, nerve compression value, and volume of the deformed area corresponding to each trajectory, providing a biomechanical basis for path adjustment and surgical strategy decision-making. It can also be simultaneously superimposed on the mixed reality view to achieve intraoperative visual deformation warning.
[0036] The system performs point-to-point mapping between the stress-displacement data in the deformation prediction output and the original trajectory to identify "excessively deformable nodes" and "stress concentration segments." Based on the acceptable deformation threshold for biological tissue (e.g., nerve displacement less than 0.5 mm), the system sets target constraints and constructs an objective function by combining the spatial shape, angle changes, and velocity control parameters of the trajectory. The path optimization algorithm employs a multi-objective genetic algorithm (MOGA) to simultaneously minimize path length, risk intensity integral, and maximum tissue stress. In each generation of evolution, the algorithm fine-tunes the trajectory nodes, including adjusting parameters such as path point position, posture direction, and control velocity. The biomechanical response is evaluated in real-time by a simulation engine, retaining the optimal individual. The final "optimal instrument trajectory parameters" include a set of path points, posture sequence, and velocity change graph, exhibiting high safety, good accessibility, and minimal tissue damage. In simulation tests, this strategy reduced the maximum stress on the facial nerve from 1350 Pa to 860 Pa in the original trajectory, decreased tissue displacement by 35%, and significantly improved the biocompatibility of the path. These trajectory parameters will be integrated into the navigation control module to guide the surgeon's operation in real time and adapt to path changes in complex anatomical environments, providing a physiological-level feedback mechanism for intelligent surgical control.
[0037] In this embodiment, the specific steps for optimizing intelligent medical device navigation based on optimal medical device trajectory parameters are as follows: Multi-dimensional quantitative analysis was performed on the visualized risk anatomy model and optimal instrument trajectory parameters to obtain the quantitative analysis results; Based on the quantitative analysis results, iterative training in mixed reality is performed to generate iterative training samples. Intelligent instrument navigation is optimized based on iterative training samples, and an intelligent instrument navigation model is constructed to perform temporal bone treatment instrument navigation operations.
[0038] In this embodiment, the fused visualized risk anatomy model and the final generated optimal instrument trajectory parameters are systematically evaluated. The analysis dimensions mainly include five indicators: path accuracy, operational efficiency, physiological safety, trajectory stability, and neuroprotection. First, the average Euclidean distance between the actual instrument trajectory and the recommended optimal trajectory is compared using a trajectory error analysis module as a path deviation indicator; an error within 0.8 mm is considered acceptable. Second, the total trajectory length, execution time, and operational dwell points are extracted to construct an operational efficiency evaluation matrix to measure the execution complexity of the path in the anatomical environment. Third, based on real-time intraoperative tissue stress and deformation simulation data, the average nerve pressure, maximum tissue displacement, and potential damage area volume induced during the operation are evaluated as quantitative indicators of physiological safety. Regarding trajectory stability, parameters such as the rate of change of trajectory curvature and the degree of acceleration fluctuation are introduced to assess volatility and reflect the control stability of the instrument in high-difficulty operational areas. Finally, by overlaying a heatmap of the spatial adjacency relationship between the path and high-risk structures, the success rate of avoiding high-risk areas of nerves or blood vessels during surgery is statistically analyzed, reflecting the risk control effect of the path in actual operation. In the simulation experiment, the above five-dimensional indicators were used to analyze 21 simulated navigation tasks, generating a structured quantitative result data table for subsequent training sample construction and navigation model evaluation. The construction of iterative training samples was based on the "task-path-feedback" triplet. Each sample included navigation target point information, a sequence of instrument motion trajectories (spatial pose + operational parameters), and corresponding quantitative scoring labels. The training sample generation process employed the Experience Replay mechanism in reinforcement learning, selecting samples with high operational efficiency, good physiological safety, and high risk avoidance success rate from historical operational data as positive feedback samples, while low-quality trajectories were used as negative feedback samples for comparative training. In each iteration, the system automatically identified trajectory improvement space and updated the samples based on the improved quantitative indicators. To enhance sample diversity, trajectory perturbation technology was used to controllably perturb the trajectory point position, velocity, and angle (perturbation amplitude ±1.5 mm, ±10°), expanding trajectory execution variants under different conditions. Meanwhile, in the mixed reality training environment, human interaction can generate new sample data in real time. The system records the operation behavior and navigation feedback and adds them to the sample library, forming a "human-machine collaborative optimization closed loop". In the experimental setup, the system can generate about 1,200 training samples per round, covering various typical surgical path variations and risk response modes, providing comprehensive and robust learning data support for subsequent intelligent navigation models.
[0039] A hybrid strategy of supervised learning and reinforcement learning was employed to construct an intelligent instrument navigation model adaptable to different temporal bone anatomy structures and possessing real-time response capabilities. The model is based on a deep temporal neural network (such as a Bi-LSTM + Attention structure), with inputs including the current instrument pose, target point information, and a heatmap of current anatomical risk intensity. The outputs are the recommended pose change vector for the next step and the operation speed. An adaptive loss function was used during model training, comprehensively considering path deviation, physiological safety losses, and execution efficiency penalties to ensure that the system's navigation commands balance accuracy and clinical feasibility. Simultaneously, to enhance the model's generalization ability under complex anatomical structures, a transfer learning mechanism was introduced. The pre-trained model was rapidly adapted to a digital twin model of a new patient, and an attention mechanism was used to dynamically focus on key risk structural regions. On a simulation training platform, offline data was used for batch learning, and real-time inference tests were conducted on real MR equipment. The navigation model achieved an average response time of less than 80ms and a path generation success rate of 94%. Finally, this intelligent navigation model was deployed to a surgical mixed reality terminal, combining real-time instrument tracking and anatomical model feedback to achieve personalized and intelligent navigation operation control. In practice, the model can provide path guidance, collision warnings, and risk avoidance prompts for surgeons, significantly improving surgical safety and operational efficiency, and promoting the transformation of temporal bone surgery from experience-driven to data-driven.
[0040] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.
[0041] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein are implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A temporal bone surgical navigation system in a mixed reality environment, characterized in that, The temporal bone surgical navigation system in the mixed reality environment includes a 3D reconstruction module, a risk analysis module, a risk distance calculation module, and an intelligent navigation optimization module. The three-dimensional reconstruction module is used to acquire multimodal medical images, perform voxel-level deep fusion and three-dimensional point cloud geometric reconstruction, and construct a virtual anatomical model of the temporal bone. The risk analysis module is used to analyze the damage risk propagation path of the temporal bone virtual anatomical model and obtain a visualized risk anatomical model. The risk distance calculation module is used to detect the real-time pose parameters of surgical instruments, calculate the risk area distance of the visualized risk anatomy model, and extract the risk area distance parameters. The intelligent navigation optimization module is used to adjust the movement trajectory parameters of the device and optimize the intelligent device navigation based on the distance parameters of the risk area.
2. The temporal bone surgical navigation system in a mixed reality environment according to claim 1, characterized in that, The 3D reconstruction module is used to acquire multimodal medical images, perform voxel-level deep fusion and 3D point cloud geometric reconstruction, and construct a virtual anatomical model of the temporal bone. Specifically, it is used for: Acquire multimodal medical images, including preoperative high-resolution CT, MRI, and intraoperative real-time ultrasound images; Feature points are marked on the multimodal medical images, and spatiotemporal registration is performed to obtain image spatial registration parameters. Based on image spatial registration parameters, voxel-level deep fusion is performed, and three-dimensional point cloud geometric reconstruction is carried out to construct a digital twin model of the temporal bone. Pre-defined physical properties of the temporal bone are calibrated on the digital twin model of the temporal bone to obtain a virtual anatomical model of the temporal bone.
3. The temporal bone surgical navigation system in a mixed reality environment according to claim 1, characterized in that, The risk analysis module is used to analyze the damage risk propagation path of the temporal bone virtual anatomical model to obtain a visualized risk anatomical model, specifically for: Semantic segmentation of key anatomical structures was performed on the virtual anatomical model of the temporal bone to extract the structures of the facial nerve, auditory nerve and internal carotid artery, and fitted into key temporal bone structural points. Geometric morphological analysis was performed on key temporal bone structural points to obtain their morphological characteristics. Spatial topological relationships are mined from the morphological features of structural points to obtain spatial adjacency relationships between structures; The structural safety distance threshold is calculated based on the spatial adjacency relationship between structures to obtain the structural safety distance threshold; Damage risk propagation path analysis is performed based on structural safety distance threshold to obtain the damage risk propagation network; Dynamic risk rendering of a virtual anatomical model of the temporal bone is performed based on a damage risk propagation network to construct a visualized risk anatomical model.
4. The temporal bone surgical navigation system in a mixed reality environment according to claim 1, characterized in that, The risk distance calculation module is used to detect the real-time pose parameters of surgical instruments, calculate the risk area distance of the visualized risk anatomy model, and extract the risk area distance parameters. Specifically, it is used for: Detect the real-time pose parameters of surgical instruments and extract the six degrees of freedom motion data of the instruments; Based on the six-degree-of-freedom motion data, a virtual space coordinate transformation is performed to obtain the instrument position mapping parameters; Virtual instrument registration and correction are performed on the visualized risk anatomy model based on the instrument position mapping parameters to obtain the instrument registration anatomy model; The distance parameters between the instrument tip and the injury risk area are calculated based on the instrument registration anatomical model, and the distance parameters of the risk area are extracted.
5. The temporal bone surgical navigation system in a mixed reality environment according to claim 1, characterized in that, The intelligent navigation optimization module is used to adjust the device movement trajectory parameters and optimize intelligent device navigation based on the risk area distance parameter, specifically for: The temporal threat level is calculated based on the risk area distance parameter to generate a continuous security threat index; The color intensity rendering of risk areas is defined based on the continuous security threat index to obtain color rendering parameters; Dynamic rendering modulation of the instrument registration anatomical model is performed based on color rendering parameters to obtain a risk intensity rendering model; Multimodal early warning is based on a risk intensity rendering model, and real-time operation behavior records are collected, including equipment movement trajectory, dwell time and operation force information. Based on the real-time operation behavior recording information, the movement trajectory parameters are adjusted to construct the optimal instrument trajectory parameters; Intelligent instrument navigation optimization based on optimal instrument trajectory parameters.
6. The temporal bone surgical navigation system in a mixed reality environment according to claim 5, characterized in that, The specific steps for adjusting the movement trajectory parameters based on the real-time operation behavior recording information and constructing the optimal instrument trajectory parameters are as follows: The real-time operation behavior recording information is subjected to time-series device motion pattern mining to obtain the time-series device motion pattern; Based on the temporal movement pattern of the device, navigation decisions for risk avoidance devices are made, and a trajectory for risk avoidance devices is generated. Based on the trajectory of the risk avoidance device, a temporal movement pre-simulation of the risk intensity rendering model is performed to obtain device movement pre-simulation data. Based on the pre-simulation data of instrument movement, tissue deformation is predicted to obtain the tissue deformation and nerve compression status; Based on tissue deformation and nerve compression, the movement trajectory parameters are adjusted to construct the optimal instrument trajectory parameters.
7. The temporal bone surgical navigation system in a mixed reality environment according to claim 5, characterized in that, The specific steps for optimizing intelligent medical device navigation based on optimal medical device trajectory parameters are as follows: Multi-dimensional quantitative analysis was performed on the visualized risk anatomy model and optimal instrument trajectory parameters to obtain the quantitative analysis results; Based on the quantitative analysis results, iterative training in mixed reality is performed to generate iterative training samples. Intelligent instrument navigation is optimized based on iterative training samples, and an intelligent instrument navigation model is constructed to perform temporal bone treatment instrument navigation operations.