Virtual Reality and Real-World Image Fusion Navigation Method and System for Temporal Bone Surgery
By using a virtual reality and real-world image fusion navigation method for temporal bone surgery, dynamic registration and mapping are performed using medical image data and optical tracking data to construct an augmented reality navigation view and provide real-time warnings. This solves the problems of navigation accuracy and safety in temporal bone surgery and improves the accuracy and safety of the surgery.
Patent Information
- Application Number
- CN202510968142.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2045-07-14
AI Technical Summary
Existing surgical navigation technologies cannot provide real-time feedback on dynamic changes caused by tissue displacement and instrument manipulation during temporal bone surgery, resulting in limited navigation accuracy. Furthermore, traditional methods lack real-time distance warning capabilities between surgical instruments and anatomical structures, increasing the risk of intraoperative injury.
By acquiring medical imaging data of the temporal bone region for 3D reconstruction, and combining optical tracking data and instrument spatial pose data for dynamic registration and mapping, an augmented reality navigation view is constructed. Structural distance analysis and dynamic warning parameter generation are performed to adjust the navigation strategy in real time.
It achieves high-precision matching between virtual models and actual anatomical structures, improves the accuracy and safety of surgical navigation, enhances the real-time nature and interactivity of intraoperative operations, reduces the probability of accidental damage to important nerves and blood vessels, and improves surgical efficiency.
Smart Images

Figure CN120788730B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of surgical navigation technology, and in particular to a virtual reality and real-world image fusion navigation method and system for temporal bone surgery. Background Technology
[0002] The application of virtual reality (VR) and augmented reality (AR) navigation systems in temporal bone surgery is gradually becoming an important means to improve surgical safety and accuracy. Existing surgical navigation technologies typically rely on registration of preoperative static images, but dynamic changes caused by tissue displacement, instrument manipulation, and other factors during surgery are difficult to reflect in real time, limiting navigation accuracy. Furthermore, traditional methods lack sufficient real-time distance warning capabilities between surgical instruments and anatomical structures, increasing the risk of intraoperative injury. Summary of the Invention
[0003] The main objective of this invention is to provide a virtual reality and real-world image fusion navigation method and system for temporal bone surgery, which can effectively solve the problem of registration lag in traditional static images and ensure that the virtual model and the actual anatomical structure always maintain a high-precision match.
[0004] To achieve the above objectives, the present invention provides a virtual reality and real-world image fusion navigation method for temporal bone surgery, comprising:
[0005] Medical imaging data of the temporal bone region is acquired and three-dimensional reconstruction is performed to obtain a dynamic virtual model of the temporal bone;
[0006] Acquire optical tracking data and instrument spatial pose data, perform dynamic registration and mapping on the dynamic virtual model of the temporal bone, and obtain real-time spatial fusion relationship;
[0007] An augmented reality navigation view is obtained by constructing a navigation view based on the fusion relationship between the instrument spatial pose data and the real-time space.
[0008] Structural distance analysis is performed on the augmented reality navigation view to obtain dynamic warning parameters;
[0009] The display is adjusted based on the augmented reality navigation view and the dynamic warning parameters to obtain a real-time navigation strategy.
[0010] Furthermore, the acquisition of medical imaging data of the temporal bone region for three-dimensional reconstruction to obtain a dynamic virtual model of the temporal bone includes:
[0011] The CT and MRI images of the temporal bone region are acquired and multimodal registration and alignment are performed to obtain the medical image data.
[0012] The medical image data is segmented into tissue structures to obtain an initial structure mask set;
[0013] The initial structure mask set is reconstructed into a three-dimensional surface mesh to obtain the target structure mesh model;
[0014] Spatial attributes are assigned and dynamic response parameters are bound to the target structural mesh model to obtain the dynamic virtual model of the temporal bone.
[0015] Furthermore, the acquisition of optical tracking data and instrument spatial pose data, and the dynamic registration and mapping of the temporal bone dynamic virtual model to obtain real-time spatial fusion relationships, include:
[0016] Optical data collected by a preset head marker array is acquired, and optical tracking and recognition are performed to obtain the optical tracking data;
[0017] The spatial pose data of the instrument is obtained by performing a free-degree pose analysis of the instrument using a preset surgical instrument tracking adapter.
[0018] The optical tracking data and the instrument spatial pose data are rigidly aligned in coordinate systems to obtain a registration reference dataset.
[0019] The real-time spatial fusion relationship is obtained by performing surface iterative registration between the registration benchmark dataset and the temporal bone dynamic virtual model.
[0020] Further, the step of performing surface iterative registration based on the registration benchmark dataset and the temporal bone dynamic virtual model to obtain the real-time spatial fusion relationship includes:
[0021] Based on the registration reference dataset, target points are located to obtain a set of target feature points;
[0022] Differential manifold matching is performed between the target feature point cloud and the dynamic virtual model of the temporal bone to obtain the initial registration matrix;
[0023] The vibration dominant frequency is separated from the spatial pose data of the instrument to obtain the main mode component parameters;
[0024] Based on the principal mode component parameters, the initial registration matrix is spatially registered to generate an anti-vibration registration matrix;
[0025] The tympanic cavity connectivity is verified by performing a vibration-resistant registration matrix to obtain the real-time spatial fusion relationship.
[0026] Furthermore, the step of constructing a navigation view based on the spatial pose data of the device and the real-time spatial fusion relationship to obtain an augmented reality navigation view includes:
[0027] Based on the spatial pose data of the instruments, the position of the surgical instruments is identified to obtain the position information of the surgical instruments.
[0028] The position information of the surgical instruments and the real-time spatial fusion relationship are transformed by spatial coordinates to obtain virtual space instrument mapping data;
[0029] The dynamic virtual model of the temporal bone is subjected to structure-aware rendering to obtain a multi-layered transparency volume drawing view.
[0030] The virtual space instrument mapping data and the multi-layer transparency volume drawing view are deeply fused to obtain a primary view frame;
[0031] The augmented reality navigation view is obtained by reconstructing the physical occlusion relationship and correcting optical distortion on the primary view frame.
[0032] Furthermore, the step of performing structural distance analysis on the augmented reality navigation view to obtain dynamic warning parameters includes:
[0033] Spatial trajectory extraction is performed on the augmented reality navigation view to obtain the motion vector information of the device;
[0034] Based on the motion vector information of the instrument and the dynamic virtual model of the temporal bone, a nearest neighbor structure search is performed to obtain the target spatial distribution set;
[0035] Dynamic distance field calculation is performed on the target spatial distribution set to obtain a multi-structure real-time distance mapping table;
[0036] Trajectory conflict prediction is performed based on the target spatial distribution set and the multi-structure real-time distance mapping table to obtain collision distribution prediction data.
[0037] The collision distribution prediction data and the multi-structure real-time distance mapping table are spatially correlated and integrated to obtain dynamic early warning parameters.
[0038] Furthermore, the step of adjusting the display based on the augmented reality navigation view and the dynamic warning parameters to obtain a real-time navigation strategy includes:
[0039] Spatial weight analysis of the temporal bone subregion was performed on the dynamic early warning parameters to obtain the target region regulatory factor set;
[0040] Based on the target region regulation factor set and the augmented reality navigation view, structural visual association optimization is performed to obtain a temporal bone specialized navigation view;
[0041] Based on the temporal bone-specific navigation view, drilling trajectory conflict prediction is performed to obtain a safe trajectory obstacle avoidance strategy;
[0042] Based on the aforementioned safe trajectory obstacle avoidance strategy, the temporal bone specialized navigation view is mapped to the operational field of view to obtain an operational guidance view;
[0043] Optical rendering fusion and adaptation are performed on the operation guidance view to obtain a real-time navigation strategy.
[0044] Furthermore, the step of predicting drilling trajectory conflicts based on the temporal bone-specific navigation view to obtain a safe trajectory obstacle avoidance strategy includes:
[0045] Phase constraints are constructed on the temporal bone specialized navigation view to obtain a third-order operation channel;
[0046] Collision detection is performed on the dynamic virtual model of the temporal bone based on the third-order operation channel to obtain collision interference volume information;
[0047] Based on the collision interference volume information, planar obstacle avoidance planning is performed to obtain a set of conflict-free trajectory options;
[0048] Perform path resolution and visual transformation on the conflict-free trajectory option set to obtain a visual cognition strategy;
[0049] The visual cognition strategy was verified by virtual drilling to obtain a safe trajectory obstacle avoidance strategy.
[0050] This invention also provides a virtual reality and real-world image fusion navigation system for temporal bone surgery, applied to the virtual reality and real-world image fusion navigation method for temporal bone surgery described in any one of the above, comprising:
[0051] The acquisition module is used to acquire medical image data of the temporal bone region for three-dimensional reconstruction to obtain a dynamic virtual model of the temporal bone.
[0052] The analysis module is used to acquire optical tracking data and instrument spatial pose data, and to perform dynamic registration and mapping on the dynamic virtual model of the temporal bone to obtain real-time spatial fusion relationship;
[0053] The association module is used to construct a navigation view based on the spatial pose data of the instrument and the real-time spatial fusion relationship, thereby obtaining an augmented reality navigation view.
[0054] The processing module is used to perform structural distance analysis on the augmented reality navigation view to obtain dynamic warning parameters;
[0055] The control module is used to adjust the display based on the augmented reality navigation view and the dynamic warning parameters to obtain a real-time navigation strategy.
[0056] The present invention provides a virtual reality and real image fusion navigation method and system for temporal bone surgery, which has the following beneficial effects:
[0057] By acquiring real-time optical tracking data and instrument spatial pose data, and combining this with a dynamic virtual model of the temporal bone for dynamic registration and mapping, the problem of registration lag in traditional static images can be effectively solved. This ensures that the virtual model and the actual anatomical structure maintain a high-precision match, thereby improving the accuracy of surgical navigation. An augmented reality navigation view is constructed based on the instrument spatial pose data and the real-time spatial fusion relationship, allowing for intuitive observation of the relative positions of surgical instruments and key anatomical structures. This enhances the real-time nature and interactivity of intraoperative operations and adapts to rapidly changing surgical field environments. By performing structural distance analysis on the augmented reality navigation view, the distance between instruments and key tissues is calculated in real time, and dynamic warning parameters are generated. This provides timely alerts to potential risks, reduces the probability of accidental injury to important nerves and blood vessels, and improves surgical safety. Real-time display adjustment based on the augmented reality navigation view and dynamic warning parameters provides optimal navigation strategies, assists in precise operations, improves surgical efficiency, and reduces operational errors caused by visual errors or delayed feedback. Attached Figure Description
[0058] Figure 1 This is a flowchart of a virtual reality and real image fusion navigation method for temporal bone surgery provided by the present invention;
[0059] Figure 2 This is a structural diagram of a virtual reality and real-world image fusion navigation system for temporal bone surgery provided by the present invention.
[0060] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0061] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0062] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments.
[0063] Reference Figure 1 As shown, this invention provides a virtual reality and real-world image fusion navigation method for temporal bone surgery, comprising:
[0064] Step S1: Obtain medical imaging data of the temporal bone region for three-dimensional reconstruction to obtain a dynamic virtual model of the temporal bone;
[0065] Step S2: Acquire optical tracking data and instrument spatial pose data, perform dynamic registration and mapping on the dynamic virtual model of the temporal bone, and obtain the real-time spatial fusion relationship;
[0066] Step S3: Construct a navigation view by fusing the spatial pose data of the instrument with the real-time space to obtain an augmented reality navigation view;
[0067] Step S4: Perform structural distance analysis on the augmented reality navigation view to obtain dynamic warning parameters;
[0068] Step S5: Adjust the display based on the augmented reality navigation view and dynamic warning parameters to obtain a real-time navigation strategy.
[0069] Based on the steps described above, the detailed process is as follows:
[0070] Step S1:
[0071] CT and MRI images of the patient's temporal bone region were acquired using medical imaging equipment. CT scans focused on capturing bony structures such as the ossicles and semicircular canals, with slice thickness controlled within 0.5 mm to ensure accuracy. MRI focused on imaging soft tissue structures such as the facial nerve and blood vessels. The acquired multimodal image data underwent registration and fusion processing to precisely align the bony information from CT with the soft tissue contrast from MRI in space. The fused dataset was then input into a pre-defined deep learning segmentation model, which automatically identified and extracted seven key anatomical structures within the temporal bone, including the facial nerve canal, cochlea, and vestibule. Morphological processing was then used to optimize the continuity of the segmentation results, particularly in fine areas such as the basal cochlear turn. After segmentation, surface reconstruction technology was used to generate a three-dimensional mesh model, with smoothing optimization at the facial nerve bend to ensure natural surface transitions.
[0072] Dynamic response parameters, including the elastic modulus of bone tissue and the strain threshold of neural structures, are embedded in the model to simulate the deformation effects of surgical instrument interactions in real time. The model supports intraoperative updates via cone-beam computed tomography (CBCT) to ensure synchronization with the patient's actual anatomical state.
[0073] Step S2:
[0074] By activating the optical tracking system, an infrared camera captures passive reflective markers on the patient's head fixation device at a rate of 120 frames per second, calculating their three-dimensional coordinates to establish a reference coordinate system. Tracking adapters mounted on instruments such as bone drills provide real-time feedback on the instrument's position and orientation, calculating the precise coordinates of the instrument's tip using a pre-calibrated geometric offset matrix. The tracking coordinate system is then spatially aligned with the instrument coordinate system, unifying them to the same world coordinate system.
[0075] In the exposed temporal bone area, spatial point cloud data was acquired using a calibration probe at anatomical landmarks such as the tympanic antrum entrance and the conus eminence, serving as key input for registration. A point cloud registration algorithm was used to match the intraoperatively acquired physical feature points with the virtual model, while compensating for pose drift caused by high-frequency vibrations of the bone drill. After registration, verification points were selected in the promontory region for error detection, and the topological connectivity of the scala tympani to scala vestibulae channel was verified. The final registration result was output when the spatial error was less than 0.3 mm and the channel was intact.
[0076] Step S3:
[0077] The spatial relationships obtained from registration are applied to the visualization system. Real-time instrument position data are converted to a virtual coordinate system, and motion smoothing is used to eliminate operational jitter, generating a stable display of the instrument's spatial pose. Subsequently, structure-aware rendering is performed on the virtual model. Bony structures are set to 20% transparency to maintain perspective, while nerves and blood vessels are rendered with 40% and 60% transparency respectively to enhance differentiation. Curvature-aware rendering is enabled in the cochlear spiral region to highlight complex geometric features. A depth buffer algorithm is used to analyze the spatial occlusion relationship between the instrument and anatomical structures; when the instrument is located behind the facial nerve, the occluded parts of the virtual model are automatically hidden.
[0078] The rendering results are adapted to different display devices: a virtual outline with 50% transparency is superimposed on the surgical microscope via an optical coupler, while a binocular stereoscopic view is generated in a VR headset. Finally, the inherent distortion of the optical equipment is corrected in real time, and the projection parameters are dynamically adjusted according to the interpupillary distance to ensure that the visual information accurately matches the actual surgical field in space.
[0079] Step S4:
[0080] Spatial relationship quantification analysis is performed based on a real-time rendered navigation view. The motion trajectory vectors of surgical instruments are extracted from the view data stream, including the displacement rate and directional changes of the instrument tip. Simultaneously, the spatial coordinates of anatomical structures in the dynamic virtual model of the temporal bone are invoked, with a focus on searching for adjacent structures in high-risk areas such as the tympanic segment of the facial nerve canal, the convexity of the lateral semicircular canal, and the vestibular window. A dynamic distance field calculation engine is constructed to measure the Euclidean distance from the instrument tip to each dangerous structure in real time, triggering a spatial marking mechanism when the instrument enters a preset warning range.
[0081] Based on the predicted path envelope within 60 milliseconds of the instrument's motion trajectory, and combined with the vibration characteristics of the bone drill to anticipate potential collision risks, a multi-dimensional early warning dataset containing distance values, collision probabilities, and time thresholds is generated. Finally, the distance information is bound to spatial coordinates, and after channel connectivity verification in the Gujiao area, structured early warning parameters are output, providing a quantitative basis for display and control.
[0082] Step S5:
[0083] The distance and risk level data in the dynamic warning parameters are analyzed, and visual control factors are generated by classifying temporal bone subregions such as the facial nerve area (threshold 0.5mm) and the bony labyrinth area (threshold 1.0mm). The transparency distribution of the navigation view is dynamically adjusted according to the factor weights: the transparency of bony areas is reduced to 15% to highlight the instrument position, and a red pulsating boundary (frequency 2Hz) is superimposed on the surface of the facial nerve canal to enhance the warning effect.
[0084] Simultaneously, path planning decisions are made. When warning parameters indicate high risk, an obstacle avoidance path (offset angle 8°-15°) bypassing the posterior vertical segment of the facial nerve is automatically generated and projected onto the surgical field as a green arrow. On the microscope eyepiece or head-mounted display, spatial distortion is optimized using a distortion correction matrix, and binocular parallax is calibrated based on interpupillary distance. Finally, visual guidance signals and path parameters are integrated, and a safe drilling speed scale is superimposed on the mastoid antrum entrance area, forming a comprehensive strategy output consistent with clinical operational intuition.
[0085] This invention provides a virtual reality and real-world image fusion navigation method for temporal bone surgery. By acquiring optical tracking data and instrument spatial pose data in real time and combining them with a dynamic virtual model of the temporal bone for dynamic registration and mapping, it effectively solves the problem of registration lag in traditional static images, ensuring that the virtual model and the actual anatomical structure maintain a high-precision match at all times, thereby improving the accuracy of surgical navigation. An augmented reality navigation view is constructed based on the instrument spatial pose data and the real-time spatial fusion relationship, allowing for intuitive observation of the relative positions of surgical instruments and key anatomical structures, enhancing the real-time nature and interactivity of intraoperative operations, and adapting to rapidly changing surgical field environments. By performing structural distance analysis on the augmented reality navigation view, the distance between instruments and key tissues is calculated in real time, and dynamic warning parameters are generated, which can promptly alert potential risks, reduce the probability of accidental injury to important nerves and blood vessels, and improve surgical safety. Real-time display adjustment based on the augmented reality navigation view and dynamic warning parameters can provide optimal navigation strategies, assist in precise operation, improve surgical efficiency, and reduce operational errors caused by visual errors or delayed feedback.
[0086] In one embodiment, medical imaging data of the temporal bone region is acquired and three-dimensionally reconstructed to obtain a dynamic virtual model of the temporal bone, including:
[0087] Temporal bone region imaging was performed using standardized scanning procedures with medical imaging equipment. CT image data was acquired via a spiral scanner, with slice thickness controlled to not exceed a preset threshold voltage value, conforming to clinical standards. MRI image data used specific sequence parameters, with slice interval settings ensuring spatial resolution requirements. Both types of data were input into a multimodal registration system to perform coordinate alignment: a similarity measurement method based on tissue grayscale distribution characteristics drove spatial transformation parameter optimization, achieving coordinate system unification between bony structures and soft tissue contrast information.
[0088] The registration process utilizes a rigid transformation model to adjust for displacement and rotational deviations between scanning coordinate systems, ultimately generating medical imaging data. This data combines the bone window visualization capabilities of CT images with the neurovascular resolution characteristics of MRI images, covering the complete anatomical range from the petrous part of the temporal bone to the mastoid region. Registration accuracy has been verified to meet sub-millimeter error standards, ensuring geometric fidelity and spatial consistency in subsequent processing. A verification mechanism is employed during image data transmission to guarantee data integrity.
[0089] Medical image data is input into a 3D segmentation processing module, which classifies it into multiple target structures based on anatomical features. The segmentation process employs a neural network architecture to extract pixel-level features, capturing global context through downsampling in the encoding stage and restoring spatial details through upsampling in the decoding stage. The segmentation results are binarized to generate mask sequences corresponding to each structure. In the post-processing stage, morphological operations are performed to correct segmentation defects: a cavity-filling algorithm is applied to thin-walled regions to ensure structural continuity; discrete noise regions are removed based on graph theory connectivity analysis. A manual interactive verification mechanism is implemented for key transition regions, such as lumen junctions, to ensure the topological integrity of tubular structures. The final output initial structural mask set retains substructural separation information of suture spaces and provides a discrete voxel representation basis for subsequent surface reconstruction. The mask data storage format supports 3D spatial coordinate mapping queries.
[0090] The initial structural mask set is input into the meshing processing unit. Surface reconstruction uses an isosurface extraction algorithm to convert binary voxels into triangular mesh surfaces, with the isosurface threshold setting balancing structural integrity and detail preservation requirements. The mesh generation process employs an adaptive subdivision strategy for regions with abrupt curvature changes, controlling the balance between mesh density and geometric deviation. The original mesh is iteratively optimized using a smoothing algorithm, constraining vertex displacement amplitudes to maintain the original structural morphology. Parametric surface reconstruction technology is applied to complex surface intersection regions to improve surface smoothness. The output target structural mesh model passes closure and self-intersection checks, with the number of triangular faces remaining within a controllable range. The mesh topology in narrow channel regions conforms to physiological and anatomical structural characteristics, and the turning angles and curvature distribution meet clinical identification standards. The mesh model file format includes vertex coordinates and facet index data structures.
[0091] The system includes a target structure mesh model with attribute mapping functionality. Spatial attribute assignment involves parsing raw image metadata, converting physical resolution parameters into vertex spatial coordinates, and assigning preset transparency levels and color coding schemes based on tissue type labels. Dynamic response parameters are bound to a mechanical property simulation component: vibration transfer functions are set for bony structural nodes, and strain threshold response mechanisms are defined; kinematic constraint models are established in movable joint areas; and deformation monitoring rules are deployed in thin-walled regions. The model integrates a real-time update interface supporting intraoperative image data input, driving mesh vertex position updates through a displacement field calculation engine. The constructed temporal bone dynamic virtual model integrates spatial attribute data and dynamic response logic, supporting real-time interactive simulation and spatial state feedback during surgical navigation. The model output format is compatible with mainstream 3D rendering pipeline architectures.
[0092] This embodiment utilizes multimodal image registration and alignment to ensure spatial consistency between bony structure and soft tissue image data, forming fused medical image data with both structural recognizability and geometric accuracy, effectively supporting the anatomical fidelity requirements of subsequent 3D reconstruction. The tissue segmentation process employs layered processing based on anatomical features, preserving the spatial topological characteristics of thin-walled regions through a neural network architecture. The initial structural mask set fully reflects the connectivity of key temporal bone channels, significantly improving the structural integrity of subsequent mesh modeling. The 3D surface mesh reconstruction combines isosurface extraction and parametric surface optimization, maintaining the morphological characteristics of complex surfaces while controlling geometric deviations. The generated target structural mesh model meets clinical standards for identifying narrow cavities.
[0093] In one embodiment, optical tracking data and instrument spatial pose data are acquired, and a dynamic registration and mapping are performed on a dynamic virtual model of the temporal bone to obtain a real-time spatial fusion relationship, including:
[0094] The infrared optical tracking system consists of three or more high-resolution cameras mounted in the surgical area, angled to cover the operating space of a pre-defined array of markers around the patient's head. This array comprises passive reflective spheres, with diameters matching clinical navigation accuracy standards and mounted on the headframe in a non-coplanar geometric configuration. The camera array synchronously acquires images of the reflected light spots, with a frame rate set to 120Hz to meet real-time requirements.
[0095] The image processing module extracts the center coordinates of the light spot and filters out background interference, applying spatial triangulation principles to calculate the three-dimensional spatial coordinates of each marker point. The raw coordinate data undergoes filtering and noise reduction to remove motion jitter and environmental vibration noise. The coordinate sequence is timestamped and encapsulated into a data stream, with a digital checksum mechanism ensuring transmission integrity. The output optical tracking data accuracy has been verified to be less than 0.3 mm, meeting the standards for temporal bone surgical navigation. The coordinate storage format includes spatial displacement vectors and rotation quaternion parameters.
[0096] The surgical instrument handle integrates an active tracking adapter, with an array of infrared luminescent markers embedded on its surface. The spatial arrangement of these markers establishes a geometric mapping relationship with the instrument's functional endpoints through a pre-calibration process, controlling the calibration error within 0.1 mm. The pose calculation engine receives the real-time coordinates of the adapter's markers and calculates the six-DOF pose parameters based on an orthogonal projection model.
[0097] Rotational components are expressed as quaternions to avoid singularity issues, while translational components output the absolute coordinates of the instrument tip in the tracking coordinate system. A kinematic compensation mechanism for instrument spatial pose data fusion is implemented: a built-in vibration spectrum analysis module for high-speed drill bits automatically corrects axial offsets caused by vibrations in the 12-18kHz range. The data output frequency is strictly synchronized with the optical tracking frame rate, with a time delay verified to be less than 2 milliseconds. The pose data stream is transmitted to the navigation system interface via an encrypted protocol, and a verification mechanism prevents data tampering. The format specification includes the pose matrix and instrument type identification code.
[0098] The module aligns the optical tracking data with the spatial pose data of the surgical instruments using the input coordinate system. The processing flow defines the head reference frame coordinate system as the spatial registration reference system, and uses a singular value decomposition algorithm to solve for the rigid transformation matrix from the tracking world coordinate system to the reference frame coordinate system. The instrument spatial pose data undergoes coordinate transformation based on the transformation matrix, generating six-DOF pose parameters in a unified coordinate system. The transformation process includes a residual verification step: detecting the spatial coordinate transformation residual while the instrument is stationary, with a threshold set at 0.2 mm. The registration reference dataset constructs an integrated 3D spatial point cloud and instrument pose mapping relationship, with the point cloud acquisition range covering the temporal bone mastoid cortex, cone eminence, and anatomical landmarks of the mastoid antrum entrance region. The dataset is labeled with a time-synchronized stamp, and the storage format includes coordinate transformation parameters and a spatial point topological relationship index. An encrypted digest verification is performed before data output to ensure the integrity of input in subsequent registration processes.
[0099] The registration benchmark dataset is loaded into the surface registration engine, which calls the triangular patch geometry data of the dynamic virtual model of the temporal bone. The iterative registration process is divided into a two-stage architecture: the initial coarse registration stage estimates the spatial transformation parameters through principal component analysis; the fine registration stage uses an iterative nearest-point algorithm to optimize the distance error between corresponding points. The registration process introduces an anatomical feature constraint mechanism: limiting the deviation of the horizontal segment length of the facial nerve canal to no more than 0.25 mm, and maintaining the continuity of curvature changes in the tympanic scutum region.
[0100] The optimized engine integrates a vibration compensation module to generate a dynamic pose correction field based on the instrument's acceleration spectrum characteristics. The registration output includes seven-dimensional transformation parameters (rotation quaternions + translation vectors) and confidence coefficients. A spatial verification process is performed in the Gujia area: five non-coplanar verification points are selected, and the root mean square error between the measured coordinates and the virtual model mapping points must be less than a threshold of 0.3 mm. Finally, the real-time spatial fusion relationship data stream is kept synchronized with the surgical operation via timestamps, and the output interface is compatible with augmented reality rendering pipeline protocols.
[0101] This embodiment utilizes the collaborative processing of optical tracking with a head marker array and device adapter pose analysis to ensure sub-millimeter spatial accuracy of optical tracking data and real-time device pose calculation, establishing a reliable spatial benchmark for multi-source data. Rigid coordinate system alignment unifies the head reference system and device coordinate system through rigid transformation, eliminating inherent system biases. The registration benchmark dataset comprehensively covers key anatomical landmarks of the temporal bone, providing highly discriminative spatial features for dynamic registration. The surface iterative registration process combines anatomical constraints and vibration dynamic compensation, effectively suppressing pose drift caused by high-frequency vibrations of the bone drill while ensuring the integrity of the facial nerve canal length and the thin-walled morphology of the tympanic cavity. Before outputting the real-time spatial fusion relationship, multi-point spatial verification of the promontory region is performed to confirm that the spatial mapping error is below the clinically safe threshold.
[0102] In one embodiment, surface iterative registration is performed based on a registration benchmark dataset and a dynamic virtual model of the temporal bone to obtain a real-time spatial fusion relationship, including:
[0103] The coordinate sequences in the registration benchmark dataset were analyzed to filter out points located in the exposed area of the mastoid process of the temporal bone. The mastoid process is the core operating area for temporal bone surgery, containing anatomical landmarks such as the entrance to the mastoid antrum, the conus eminence, and the posterior superior spur of the external auditory canal. Its bony structures exhibit significant grayscale differences in images, facilitating feature identification. By setting a curvature threshold, geometric abrupt change points were identified to ensure that the selected points are the tips or inflection points of bony structures, rather than soft tissue attachment areas.
[0104] To ensure the uniformity and coverage of the point cloud distribution, a spatial constraint strategy is adopted: the minimum distance between points must be greater than 0.5 mm to avoid computational redundancy caused by dense points; the point density in key areas (such as the periphery of the tympanic sulcus entrance) must be higher than in other areas to ensure the positioning accuracy of these areas. Morphological filtering is performed on candidate points, using a 3×3×3 cube kernel for closing operations to fill tiny holes and smooth the surface, removing isolated points caused by noise.
[0105] The resulting target feature point cloud contains 60-120 points, covering 6-8 key anatomical landmarks such as the entrance to the mastoid antrum, the apex of the triangular fossa, and the posterior superior spur of the external auditory canal. Each point is accompanied by three-dimensional coordinates and curvature attributes. These point clouds are time-stamped and aligned with the surgical operation timeline to ensure real-time processing of subsequent procedures. The point cloud accuracy has been verified to be less than 0.3 mm, meeting sub-millimeter navigation requirements.
[0106] Differential manifold matching aims to geometrically align the target feature point cloud with the surface of a dynamic virtual model of the temporal bone to find the optimal spatial transformation matrix. This process must balance global matching accuracy with preservation of local anatomical structure, avoiding deformation of critical structures.
[0107] A spatial index of the target feature point cloud and the virtual model is constructed. The matching algorithm adopts an improved Iterative Closest Point (ICP) algorithm: in the initial stage, principal component analysis (PCA) is used to estimate the rough rotation and translation parameters to roughly align the point cloud to the model surface; in the fine optimization stage, anatomical constraints are introduced to rigidly limit the length error of the tympanic segment of the facial nerve canal and the curvature gradient of the tympanic shield area to ensure the integrity of the key structural morphology.
[0108] The optimization process employs a constrained nonlinear least squares method, iteratively adjusting transformation parameters through gradient descent to minimize the weighted sum of distances between the point cloud and corresponding points in the model. Weight allocation is tilted towards critical regions (e.g., the facial nerve canal region has twice the weight of other regions), prioritizing alignment of important structures. Upon convergence of the iterations, an initial registration matrix is obtained, containing rotation quaternions and translation vectors, which maps the target point cloud to the virtual model coordinate system.
[0109] The validity of the matrix is verified through residual evaluation: the average distance of all matching points is calculated, and if it is less than 0.3 mm, the initial registration is considered successful; otherwise, the constraints are adjusted and the iteration is repeated. The final output initial registration matrix provides the initial transformation relationship for subsequent vibration-resistant registration, and its numerical stability is verified by orthogonality to ensure that the transformation is free of singularities.
[0110] High-frequency vibrations generated during the operation of surgical instruments (such as bone drills and suction devices) can cause pose data drift. It is necessary to separate the dominant vibration frequency from the pose sequence and extract the dominant mode component parameters to provide a basis for vibration compensation.
[0111] The instrument pose data is decomposed into a time series, and the acceleration signal is split into multiple intrinsic mode functions (IMFs) through empirical mode decomposition (EMD). Each IMF corresponds to a vibration component at a different frequency. By selecting a termination condition (the number of extreme points of two consecutive IMFs < a threshold), the dominant vibration modes (usually the first 3) are extracted, corresponding to the 8000-15000Hz frequency band (the main vibration range of the bone drill).
[0112] The frequency domain analysis module performs a Fast Fourier Transform (FFT) on the extracted IMF to identify the amplitude, frequency, and phase information of the principal mode components. Bone drills focus on extracting the first three amplitude attenuation coefficients (describing the characteristics of vibration energy decay over time); suction devices capture the phase shift of low-frequency drift (typically <100Hz).
[0113] The principal mode parameters are output in time-series format, including amplitude, frequency, phase, and corresponding timestamps, synchronized with the instrument's motion state. These parameters are subsequently used to construct a vibration compensation field, eliminating vibration-induced pose drift through inverse displacement correction, thus ensuring the stability of the registration process.
[0114] The principal mode shape parameters reflect the influence of surgical instrument vibration on pose during operation and need to be incorporated into the optimization process of the initial registration matrix to counteract pose drift caused by vibration. In practice, the principal mode shape parameters (including amplitude, frequency, phase, and timestamp) are first spatiotemporally aligned with the initial registration matrix to ensure that the time dimension of vibration compensation is consistent with the time axis of pose transformation.
[0115] The core of spatial registration is the construction of a vibration compensation field: based on the dynamic changes in amplitude in the principal mode parameters, a displacement correction vector opposite to the vibration direction is generated; the duration of the correction vector's action is adjusted in conjunction with frequency characteristics to synchronize the compensation effect with the vibration period; and the direction of the correction vector is calibrated using phase information to avoid reverse errors caused by phase shift. This process is achieved through transformation synthesis in Lie group space—the rotation component of the initial registration matrix is multiplied by the rotation correction amount of the vibration compensation using quaternions, while the translation component is vector-superimposed with the translation correction amount of the vibration, ultimately generating the vibration-resistant registration matrix.
[0116] To ensure the effectiveness of compensation, a dynamic error suppression mechanism is introduced during registration: when the instrument vibration amplitude exceeds a preset threshold (e.g., 0.5 mm), the compensation weight in the corresponding direction is automatically increased; if the vibration frequency deviates from the statistical mean of the principal mode parameters (e.g., ±10%), the duration of the compensation vector is adjusted. Before outputting the vibration-resistant registration matrix, local area verification is performed: three non-coplanar points around the tympanic inlet are selected, and their positional deviations in the virtual model and actual space are calculated. If the average deviation is less than 0.2 mm, the compensation is considered effective; otherwise, the extraction strategy for the principal mode parameters is adjusted retrospectively. The final generated vibration-resistant registration matrix integrates the spatial relationship of the initial registration with the dynamic correction of vibration compensation, providing a stable foundation for subsequent topology verification.
[0117] Tympanic cavity connectivity verification is a crucial step in confirming whether the vibration-resistant registration matrix accurately reflects the continuity of the internal anatomical structure of the temporal bone, directly affecting the reliability of the navigation system. In practice, the three-dimensional geometric data of the tympanic cavity passage (including the spatial distribution of the scala tympani, scala vestibulae, and mastoid air cells) in the dynamic virtual model of the temporal bone is first called, and then spatially superimposed with the virtual point cloud mapped by the vibration-resistant registration matrix.
[0118] The verification process focuses on the morphological integrity of the tympanic passage: by calculating the continuity indicators (such as curvature abrupt change rate and normal vector angle deviation) of the superimposed point cloud at the tympanic step-vestibulocochlear step interface, it is determined whether there are structural breaks or overlaps caused by registration errors. If the continuity indicators exceed the preset threshold (such as curvature abrupt change rate > 0.3 / mm), it is determined that the registration matrix has local distortion, and the principal mode parameters need to be re-applied for compensation and correction; if the indicators meet the requirements, the connectivity of the passage is further verified—five non-coplanar verification points are selected, and their geodesic distances in the virtual model and actual space are calculated. If the maximum deviation of all geodesic distances is less than 0.3mm, the spatial fusion relationship is determined to be valid.
[0119] The validated real-time spatial fusion relationship includes the final transformation parameters of the vibration-resistant registration matrix, the spatial coordinates of the validation points, and continuity indicators. The output embeds timestamps and instrument status codes to ensure real-time synchronization with the surgical procedure. This real-time spatial fusion relationship is directly used for rendering the augmented reality navigation view, providing accurate virtual-reality spatial mapping and supporting subsequent obstacle avoidance strategy generation and operational guidance.
[0120] This embodiment establishes a highly discriminative spatial benchmark for temporal bone surgical navigation through precise localization of target feature point clusters. Sub-millimeter-level localization accuracy of key anatomical landmarks (such as the tympanic antrum entrance and the tympanic cone) effectively reduces initial deviations in subsequent registration. The differential manifold matching process balances global geometric alignment with local anatomical constraints, preserving the morphological integrity of key parameters such as facial nerve canal length and tympanic scutum curvature, thus avoiding spatial deformation deviations between the virtual model and the actual structure. Vibration frequency separation and the generation of the anti-vibration registration matrix specifically suppress pose drift caused by high-frequency vibrations of the bone drill, while a dynamic compensation mechanism ensures the real-time performance and stability of spatial mapping during high-speed drilling operations.
[0121] In one embodiment, a navigation view is constructed by fusing the instrument's spatial pose data with real-time spatial data to obtain an augmented reality navigation view, including:
[0122] Spatial pose data of surgical instruments is acquired in real time by a pre-set tracking adapter, which is typically installed near the instrument handle or functional end. The adapter obtains the three-dimensional coordinates and rotational orientation of the instrument in the surgical field through infrared reflective markers or active light-emitting elements. The pose data includes translational components (X, Y, Z axis coordinates) and rotational components, which reflect the absolute position and orientation of the instrument in the tracking coordinate system.
[0123] The core of position recognition is converting these pose parameters into actual spatial positions within the surgical field. In practice, the translation vector is first extracted from the pose data; this vector directly represents the displacement of the instrument tip relative to the tracking origin. The rotation component is used to determine the instrument's pointing direction, which is converted to a direction cosine matrix using Euler angles to obtain the unit direction vector of the instrument tip. To ensure the accuracy of the position information, the raw data undergoes dual verification: first, cross-verification using multiple cameras in the optical tracking system eliminates single-view measurement errors; second, by combining the instrument's kinematic model, low-pass filtering is applied to the pose jitter caused by high-frequency vibrations, preserving the low-frequency steady-state position. For example, a bone drill generates vibrations of 12-18kHz during high-speed rotation; filtering removes high-frequency noise, retaining the low-frequency steady-state displacement of 0.1-10Hz. The final output surgical instrument position information includes the three-dimensional coordinates of the instrument tip (accuracy ≤0.3mm) and its pointing direction (error ≤2°), stored as a timestamp sequence, strictly synchronized with the surgical operation sequence, providing a reliable physical position reference for subsequent spatial mapping.
[0124] The real-time spatial fusion relationship defines the coordinate mapping rules between the virtual model and the real surgical field, including rotation matrices and translation vectors, which are generated by the registration process in step 4. The goal of spatial coordinate transformation is to convert the actual positions of surgical instruments (from the position information in step 1) into corresponding coordinates in the virtual model coordinate system, so as to achieve precise superposition of virtual and reality.
[0125] The transformation parameters (rotation matrix R and translation vector T) in the real-time spatial fusion relationship are invoked to transform the instrument position information from the tracking coordinate system to the virtual model coordinate system. Among them, the rotation matrix R is determined by the registration process to ensure that the spatial orientation of the virtual model is consistent with that of the real surgical field; the translation vector T compensates for the origin offset between the tracking coordinate system and the virtual model coordinate system.
[0126] The virtual coordinates are constrained and validated: if the projection of the instrument tip in the virtual model exceeds the anatomical boundary of the temporal bone (e.g., entering the intracranial space or mastoid air cells), a position correction mechanism is triggered. Correction strategies include adjusting the pitch angle of the rotation matrix or the anterior-posterior displacement of the translation vector to restrict the instrument position within the effective range of the temporal bone bony structures (e.g., between the mastoid cortex and the tympanic tectum). The final output virtual space instrument mapping data includes the coordinates (X, Y, Z) of the instrument tip in the virtual model and its direction vectors (U, V, W), with a correspondence error ≤0.5mm between the instrument tip and the virtual model surface mesh, providing a spatial alignment basis for subsequent rendering.
[0127] The dynamic virtual model of the temporal bone includes various anatomical elements such as bony structures (e.g., mastoid cortex, petrous apex) and soft tissues (e.g., facial nerve, blood vessels). The core of structure-aware rendering is to differentiate rendering parameters based on the anatomical characteristics and clinical importance of the tissues to highlight key structures.
[0128] The model is classified by material: bony structures (such as the mastoid cortex) are given a high-density material with a refractive index of 1.5 to simulate their dense characteristics; soft tissues (such as the facial nerve sheath) are given a low-density material with a refractive index of 1.3 to highlight the interface difference between them and bony structures; and nerves and blood vessels (such as nerve fibers in the facial nerve canal) are given a translucent material with a transparency of 0.4 to maintain visibility while avoiding obscuring deep structures.
[0129] The rendering process employs a layered rendering strategy: the bottom layer consists of bony structures, rendered using a depth-first algorithm to ensure the clarity of anatomical contours; the middle layer consists of soft tissues, whose translucency is simulated through ray tracing, and the refractive index difference of light at the soft tissue-bone structure interface enhances the visibility of the interface; the top layer consists of nerves and blood vessels, whose boundaries are strengthened using an edge enhancement algorithm, and the transparency is set to 0.6 to balance visibility and permeability.
[0130] To adapt to the dynamic changes in the surgical scene, rendering parameters support real-time adjustment: when the instrument approaches the facial nerve (e.g., distance < 2mm), the transparency of the facial nerve layer is automatically increased to 0.8, while the transparency of the surrounding bony structures is decreased to 0.3, guiding attention to key areas; when the instrument moves away from key structures, the default transparency settings are restored to maintain the integrity of the overall anatomical structure. The final multi-layered transparency volume rendering view includes depth buffer information and a material property table, providing a visual foundation for subsequent fusion with instrument mapping data.
[0131] The core of fusing virtual instrument mapping data with multi-layered transparency volume rendering views lies in establishing consistency between the virtual instruments and the 3D anatomical model in terms of spatial position and visual presentation. In practice, firstly, based on the coordinate information in the virtual instrument mapping data, a spatial correspondence is established between the 3D position of the instrument tip and the corresponding anatomical structure in the multi-layered transparency volume rendering view. For example, the coordinates of the instrument tip in the virtual model are mapped to the rendering layer corresponding to those coordinates in the multi-layered view.
[0132] During the fusion process, the occlusion relationship between instruments and models needs to be addressed: if the tip of an instrument is located behind an anatomical structure (such as the facial nerve canal) in the virtual model, the drawing order of the instrument in the multi-layer view is adjusted so that it is occluded by the structure; if the instrument is in front, its transparency is increased to ensure the visibility of the structure behind it. At the same time, the display attributes of the instrument are dynamically adjusted according to its functional type (such as bone drills and suction devices)—bone drills are highlighted at the edges when they are close to bony structures, and suction devices are less transparent when they are close to soft tissues to highlight the operating area.
[0133] The output primary view framework includes the spatial location of the instrument, its display attributes (transparency, brightness), and its occlusion relationship with the model, forming a preliminary visualization result that overlays virtual and reality, providing a foundation for subsequent optimization.
[0134] Optimizing the primary view frame requires addressing two core issues: accurate reproduction of physical occlusion and optimization of optical display.
[0135] Physical occlusion relationship reconstruction is achieved through depth information: First, the depth values of the instrument tip and each structural point in the model (i.e., the distance to the surgical field camera) are calculated. The drawing order is determined based on the depth values—structures with smaller depth values (closer to the camera) are drawn first, and structures with larger depth values (farthest from the camera) are drawn later, ensuring that occluded structures are not displayed incorrectly. For complex occlusion scenarios (such as the instrument partially occluding the facial nerve canal), the occlusion relationship is corrected by adjusting the drawing order or transparency to ensure that the occlusion logic is consistent with the actual surgical field.
[0136] Optical distortion correction addresses the inherent characteristics of display devices. After obtaining distortion parameters through device calibration, each pixel in the primary view frame is adjusted to ensure the displayed content conforms to the geometry of the actual surgical field. For example, barrel distortion, which may be caused by wide-angle lenses, is corrected by adjusting the pixel coordinate distribution to adjust the display distortion of instruments and models, ensuring the accuracy of spatial relationships.
[0137] The output augmented reality navigation view integrates the real logic of physical occlusion with the geometric optimization of optical display. The spatial relationship between instruments and models is completely consistent with the actual surgical field, providing accurate virtual-reality fusion navigation information.
[0138] This embodiment ensures accurate instrument position recognition and spatial coordinate transformation, guaranteeing the positional accuracy of surgical instruments in both virtual and real spaces and providing a reliable benchmark for navigation. Structure-aware rendering, through layered transparency settings, highlights key anatomical structures of the temporal bone, improving the efficiency of identifying deep structures. Deep fusion of virtual and real views and physical occlusion reconstruction eliminate display misalignment between virtual instruments and models, ensuring consistency between occlusion relationships and the actual surgical field. Optical distortion correction corrects the inherent deformation of the display device, ensuring the geometric realism of the navigation screen.
[0139] In one embodiment, structural distance analysis is performed on the augmented reality navigation view to obtain dynamic warning parameters, including:
[0140] Augmented reality navigation views integrate spatial information from virtual models and the real surgical field. The movement trajectory of instruments within the view intuitively reflects their real-time positional changes relative to the temporal bone anatomy. The core of spatial trajectory extraction is capturing the positional changes of the instrument tip through continuous frame analysis, generating vector information describing its direction and velocity. In practice, the view is first scanned frame by frame. Image feature matching or optical flow tracking techniques are used to identify the pixel coordinates of the instrument tip in each frame, and the corresponding timestamp is recorded. Subsequently, by comparing the coordinate changes of adjacent frames and combining the inter-frame time interval Δt, the instantaneous direction and velocity of the instrument are calculated, ultimately generating instrument motion vector information containing the position sequence and corresponding velocity directions (such as along the positive X-axis, negative Y-axis, etc.). The temporal resolution of this information is synchronized with the view refresh rate, providing a dynamic positional reference for subsequent search of neighboring structures.
[0141] The goal of nearest neighbor structure search is to identify key anatomical structures near the current position of the instrument, providing spatial context for subsequent distance calculations and collision prediction. In practice, a search range matching the instrument's size and anatomical density is first defined, centered on the instrument's current position in its motion vector. Then, the triangular mesh data of the temporal bone dynamic virtual model is traversed to filter out all structural points whose vertex coordinates fall within this search range. To improve efficiency, spatial indexing techniques are used to pre-organize the model vertices, limiting the search range to a local area around the instrument. The selected structural points must meet two core conditions: first, they must belong to core temporal bone anatomical structures (such as the facial canal, semicircular canals, and ossicles); second, they must pose a potential contact risk with the instrument in the virtual model (e.g., a distance less than a safety threshold). The final target spatial distribution set includes the vertex coordinates of these neighboring structures, their anatomical type (e.g., bony / soft tissue), and their initial distance from the instrument, providing input data for dynamic distance field calculations.
[0142] The core of dynamic distance field calculation is to quantify the spatial distance between the instrument and each neighboring structure in real time, forming a distance mapping relationship that changes over time. In practice, for each structural point in the target spatial distribution set, the spatial distance between them (e.g., the straight-line distance between the structural point coordinates and the instrument coordinates) is calculated by comparing its position coordinates with the instrument's coordinates in the current frame. To avoid redundant calculations, an incremental update strategy is adopted: if the displacement of the instrument's motion vector is small (e.g., the position difference from the previous frame is less than 0.1 mm), the distance value from the previous frame is reused and minor errors are corrected; if the displacement is large (e.g., exceeding 0.5 mm), the distances of all neighboring points are recalculated. The calculation results are stored according to anatomical structure type, forming a multi-structure real-time distance mapping table. The table includes the structure type (e.g., facial nerve canal, lateral semicircular canal), structural point coordinates, current distance value, and timestamp. This mapping table is updated in real time with the instrument's movement, providing dynamic distance data support for subsequent trajectory conflict prediction.
[0143] The core of trajectory conflict prediction is to predict the risk of collision in the future by analyzing the spatial relationship between the current motion state of the instrument and the adjacent anatomical structures. In practice, a short-term trajectory prediction model is first constructed based on the velocity direction and historical displacement patterns in the instrument's motion vector. For example, assuming the instrument moves at a constant current speed (ignoring the influence of high-frequency vibration), its future position can be represented as a linear superposition of the current position and the velocity vector.
[0144] Subsequently, the predicted trajectory is spatially overlaid with the locations of neighboring structures in the target spatial distribution set for analysis: for each neighboring structure point, the minimum spatial distance between each point on the predicted trajectory and the structure point is calculated; if the predicted distance at a certain moment is less than the collision safety threshold of the structure, the structure point is marked as a potential collision risk point, and the time interval and spatial location of the collision are recorded.
[0145] To improve prediction accuracy, a motion trend correction mechanism is introduced: if the velocity or direction of the machine's motion vector changes abruptly (such as when encountering resistance during drilling causing a speed decrease), the time step of the prediction model is adjusted to shorten the prediction cycle and improve real-time performance. The final collision distribution prediction data includes the potential collision structure type, predicted collision time (within 0.5-2 seconds), and collision risk level (high / medium / low), providing a risk distribution basis for the subsequent generation of warning parameters.
[0146] The generation of dynamic warning parameters requires combining the spatial risk of collision prediction with the dynamic changes in real-time distance to form specific warning information that can guide operations. In practice, each risk point in the collision distribution prediction data (such as a certain region of the facial nerve canal) is first associated with the current distance of the corresponding structural point in the multi-structure real-time distance mapping table. For example, if the current distance of the structural point corresponding to the risk point is 1.8mm (close to the safety threshold of 2mm), and the distance is predicted to decrease at a rate of 0.1mm / second within the next 0.3 seconds, then the dynamic warning parameter of the risk point can be defined as "the anterior region of the facial nerve canal, high collision risk within the next 0.3 seconds, current distance 1.8mm".
[0147] Differentiated markings are used for collision points of different risk levels: high-risk points (such as predicted collision time < 1 second and current distance < safety threshold) are marked with a red flashing boundary; medium-risk points (such as predicted collision time 1-2 seconds and current distance close to safety threshold) are marked with a yellow gradient boundary; low-risk points (such as predicted collision time > 2 seconds or current distance much greater than safety threshold) are marked with a green static boundary.
[0148] The output dynamic warning parameters include the spatial coordinates of the risk point, the predicted collision time, the current distance, and the risk level indicator. These parameters are overlaid on the surgical field in real time through the augmented reality navigation view, providing the surgeon with clear obstacle avoidance guidance.
[0149] This embodiment uses spatial trajectory extraction to track instrument movement in real time, providing a dynamic positional reference for subsequent analysis; neighboring structure search accurately locates key anatomical areas, avoiding unnecessary calculation interference; dynamic distance field calculation quantifies the spatial relationship between the instrument and surrounding structures, making risks perceptible; trajectory conflict prediction identifies potential collision risks in advance, buying time for operational adjustments; dynamic early warning parameters transform risk information into intuitive guidance, directly assisting the surgeon's decision-making. Overall, it achieves end-to-end coordination of instrument movement, anatomical structure, and risk warning, reducing the probability of accidental structural contact and improving the accuracy and safety of surgical procedures.
[0150] In one embodiment, a real-time navigation strategy is obtained by adjusting the display based on the augmented reality navigation view and dynamic warning parameters, including:
[0151] Dynamic early warning parameters include real-time distances between the instrument and various anatomical structures of the temporal bone, collision risk levels, and predicted collision times. Spatial weighting needs to be assigned based on the functional zoning characteristics of the temporal bone. Temporal bone subregions can be divided into high-sensitivity, medium-risk, and low-sensitivity areas according to anatomical function. Weighting analysis is guided by clinical risk: high-sensitivity areas, involving critical structures such as nerves and blood vessels, are weighted at 0.7-1.0; medium-risk areas, potentially affecting hearing or balance, are weighted at 0.4-0.6; and low-sensitivity areas, with high anatomical redundancy, are weighted at 0.1-0.3.
[0152] Risk levels (e.g., high / medium / low) and corresponding structural types are extracted from dynamic early warning parameters. A predefined anatomical partition mapping table is used to associate risk levels with sub-region weights. For example, a high-risk warning for the facial nerve canal corresponds to a weight of 1.0, while a medium-risk warning for the lateral semicircular canal corresponds to a weight of 0.5. Subsequently, local weight adjustments are made for risk points at different locations within the same sub-region—if a point is less than a safety threshold (e.g., 2mm) from a critical structure, its local weight is increased by 0.2; if the distance is greater than the threshold but less than the warning threshold (e.g., 5mm), the base weight is maintained. The final target region control factor set contains the weight values for each sub-region and sub-region, providing a basis for differentiated control in subsequent display optimization.
[0153] Augmented reality navigation views integrate spatial information from virtual models and the real surgical field. They need to be optimized through structural visual association to highlight key areas and reduce the surgeon's cognitive load. In practice, the first step is to spatially match the target region's modulatory factor set with the anatomical structures in the view—for example, high-weight structures in the facial nerve canal region correspond to specific colors (e.g., red) and transparency (e.g., 0.8) in the view; medium-weight structures in the lateral semicircular canal region correspond to yellow and transparency 0.5; and low-weight structures correspond to gray and transparency 0.2.
[0154] The optimization process employs a layered rendering strategy: the bottom layer represents the skeletal structure, with a depth buffering algorithm ensuring clear outlines; the middle layer consists of soft tissues and nerves / blood vessels, with display intensity adjusted according to regulatory factors—high-weighted soft tissues enhance edge contrast, while medium-weighted structures maintain normal display; the top layer contains risk warning signs, with flashing animations to enhance attention.
[0155] To adapt to the dynamic changes in the surgical scenario, optimized parameters support real-time adjustments: when instruments approach high-risk areas, the transparency of that area is automatically increased to 0.9 and a highlighted border is overlaid; when instruments move away from risk areas, default display parameters are restored. The resulting temporal bone-specific navigation view, through differentiated visual expression, intuitively presents key structures and risk information, helping surgeons quickly locate anatomical areas requiring attention.
[0156] The core of drilling trajectory conflict prediction is to perform spatial overlay analysis of the instrument's motion path and key structures in the specialized view to identify potential collision risks and generate avoidance commands. In practice, firstly, the spatial boundaries of high-weight structures in the specialized view are extracted and converted into prohibited areas in three-dimensional space. Then, the real-time motion trajectory of the instrument is acquired, and the minimum spatial distance between the trajectory line and the prohibited area is calculated.
[0157] If the minimum distance is less than the safety threshold (e.g., the safety distance for the facial nerve canal is 2mm), it is judged as a high-risk conflict, and a level 1 obstacle avoidance strategy is generated: prompting "The facial nerve canal is 2mm ahead, it is recommended to adjust the drilling angle to 15°"; if the minimum distance is between the safety threshold and the warning threshold (e.g., 5mm), it is judged as a medium-risk conflict, and a level 2 obstacle avoidance strategy is generated: prompting "The facial nerve canal is 5mm ahead, the current speed needs to be reduced to 1000rpm"; if the minimum distance is greater than the warning threshold, it is judged as low-risk, and no adjustment is required.
[0158] By introducing a motion trend correction mechanism, if the curvature of the instrument's trajectory suddenly increases (such as when drilling encounters resistance), the prediction time step is shortened (from 1 second to 0.5 seconds), and the conflict risk is recalculated. The final safe trajectory obstacle avoidance strategy includes parameters such as risk level, conflict location, and adjustment suggestions (angle / speed / pause), which are overlaid onto the surgical field in real time through augmented reality, providing the surgeon with clear operational guidance.
[0159] The safety trajectory obstacle avoidance strategy includes key information such as risk level, conflict location, and adjustment suggestions. This information needs to be spatially correlated with the anatomical structures in the temporal bone specialized navigation view to generate operational guidelines that the surgeon can directly refer to. In practice, the spatial coordinate range of the risk areas (such as the facial nerve canal and lateral semicircular canals) in the obstacle avoidance strategy is first extracted and matched with the corresponding structures in the specialized view—for example, the conflict location in the facial nerve canal area corresponds to the surface area of its 3D model in the view.
[0160] The mapping process employs spatial projection technology: the two-dimensional risk warnings in the strategy are converted into three-dimensional spatial annotations in the view, and their actual projection positions in the surgical field are determined by ray casting. For high-risk areas, a semi-transparent red warning box is overlaid in the view, with the edge of the box aligned with the anatomical outline of the risk area; for medium-risk areas, a yellow gradient box is overlaid, with the box extending to a 5mm buffer zone outside the risk area.
[0161] Adjustment suggestions are labeled as text labels near the corresponding risk areas, with the label content dynamically linked to specific parameters in the strategy. To avoid information overload, a layered display strategy is adopted: high-risk warnings are always displayed at the top, while medium- and low-risk warnings are dynamically displayed based on the instrument's movement—displayed when the instrument approaches the risk area and hidden when it moves away. The final operation guidance view deeply integrates strategy information with anatomical structures through spatial correlation, allowing operators to directly identify areas of concern and corresponding operational requirements through the view.
[0162] The operation guidance view needs to be integrated with the real-time surgical field image to ensure that the information is synchronized with the actual operation scenario, while also adapting to the display characteristics of different optical devices. In practice, the real-time video stream of the surgical field is first acquired, and the annotations (such as warning boxes and text labels) in the operation guidance view are aligned with the corresponding positions in the video stream using image registration technology—for example, the red warning box in the view needs to accurately cover the actual display area of the facial nerve canal in the surgical field.
[0163] The fusion process employs a transparency blending algorithm: overlaying the annotation layer of the guide view with the surgical field video layer ensures clear visibility of the annotation information without affecting the overall observation of the surgical field. Rendering parameters are dynamically adjusted for different optical devices: in microscope mode, contrast is enhanced to highlight the annotations; in head-mounted display mode, stereo parallax is optimized to match human visual perception. A dual-buffering rendering mechanism is used: the main buffer stores the fusion result of the current frame, while the secondary buffer pre-renders the guide view for the next frame, with frame rate synchronization achieved through a vertical synchronization signal. The final output real-time navigation strategy includes the fused surgical field image and dynamically annotated information.
[0164] This embodiment uses dynamic weight analysis to accurately locate highly sensitive areas of the temporal bone. Combined with structural visual association optimization, it presents key structures and risk information differentially in the augmented reality navigation view, significantly improving the surgeon's efficiency in recognizing anatomical points. Trajectory conflict prediction identifies potential collision risks in advance through spatial overlay analysis, generating a graded warning strategy to allow the surgeon time for adjustment and reduce the probability of accidentally touching key structures. Operational field of view mapping deeply integrates obstacle avoidance strategies with real-time surgical field images. Spatial projection technology ensures that the prompts accurately correspond to the actual anatomical locations, avoiding decision-making biases caused by information misalignment. Optical rendering adaptation optimizes display parameters for different devices, ensuring clear visibility of navigation information under complex lighting conditions. Overall, it achieves end-to-end collaboration from risk warning to operation guidance, improving the reliability of surgical navigation and the surgeon's confidence.
[0165] In one embodiment, drilling trajectory conflict prediction is performed based on a temporal bone-specific navigation view to obtain a safe trajectory obstacle avoidance strategy, including:
[0166] The temporal bone specialized navigation view, through structural visual association optimization, has clearly defined the spatial distribution and display weight of high-sensitivity, medium-risk, and low-sensitivity areas. The core of the stage constraint construction is to combine the anatomical exposure patterns and risk evolution characteristics of the surgical process, dividing the drilling operation into three stages with clear spatial boundaries, and defining the allowed operation channels for each stage.
[0167] Based on the natural anatomical boundaries of the temporal bone, the surgical field is divided into the initial exposure area (stage one), the deep access area (stage two), and the critical structure adjacent area (stage three). The initial exposure area corresponds to the superficial layer of the temporal bone, and the operating channel is defined as a three-dimensional columnar region from the entrance of the surgical field to the surface of the mastoid cortex, with a cross-sectional diameter twice the diameter of the drill bit, allowing instruments to move freely in and out.
[0168] The deep access zone corresponds to the area where the mastoid air cells and semicircular canals are located. The operating channel is contracted into a conical area extending deeper along the original path, with the cross-sectional diameter reduced by 30% compared to the initial stage, thus limiting the lateral swing of the instrument.
[0169] The critical structure adjacent area corresponds to highly sensitive areas such as the facial nerve canal and vestibular window. The operating channel is further narrowed to a linear area that only accommodates the drill tip, with a width not exceeding 1.2 times the drill diameter, to ensure that the instrument tip and the critical structure maintain a preset safe distance.
[0170] The boundary parameters of the manipulation channel (such as contraction ratio and linear width) are synchronized in real time with the dynamic virtual model of the temporal bone. If mastoid air cell fusion is detected, the angle of the cone region in the second stage is automatically adjusted to avoid abnormal air cells. Through real-time registration data identification, when the position of the facial nerve canal deviates from the anatomical atlas, the linear channel in the third stage is synchronously deviated by the same distance, always maintaining a preset safe distance from key structures. The final three-dimensional manipulation channel is stored in the form of a three-dimensional spatial grid. Each grid cell is labeled with its stage (1 / 2 / 3) and the maximum allowable drilling depth (e.g., first stage ≤2mm, second stage ≤5mm, third stage ≤1mm), providing spatial constraint rules for subsequent collision detection.
[0171] The goal of collision detection is to identify whether the spatial position of the drilling instrument in the current operation phase overlaps with high-risk structures in the dynamic virtual model of the temporal bone. The real-time position of the instrument tip (from pose data in the augmented reality navigation view) is mapped to the virtual model coordinate system by calling the 3D mesh data of the third-order operation channel, thus obtaining its permissible operating space range in the current phase (e.g., the linear channel of the third phase).
[0172] Centered on the instrument tip, a cylindrical detection area matching the drill bit diameter (diameter = drill bit diameter + 0.2 mm, to compensate for slight instrument oscillation) is constructed. This area is then spatially intersected with the triangular mesh of the dynamic virtual model of the temporal bone.
[0173] A layered detection strategy improves detection efficiency by detecting whether the instrument tip exceeds the boundary of the current stage's operating channel (such as the linear channel width limit in the third stage). If it does, it is directly marked as a high-risk collision; if it does not exceed the boundary, the overlap between the detection area and the high-risk structural mesh is further detected. The overlapping area is determined by comparing the coordinates of each point in the detection area with the coordinates of the virtual model mesh vertices and calculating the minimum Euclidean distance. When the minimum distance is less than the drill radius, it is determined to be a collision interference. The final generated collision interference volume information includes the three-dimensional coordinate range of the interference area, the interference type (such as light contact / severe penetration), and the corresponding anatomical structure name, providing a specific risk location description for subsequent obstacle avoidance planning.
[0174] Planar obstacle avoidance planning, based on collision interference volume information and constrained by a third-order operation channel, generates drilling path options that avoid all high-risk areas. The boundary surfaces of the collision interference volumes are extracted (fitted using triangular patches from a virtual model mesh) and transformed into a set of obstacle planes in 3D space. Subsequently, for each obstacle plane, its projected area within the operation channel at the current operation stage is calculated to determine the spatial range that needs to be avoided.
[0175] Starting from the current position of the instrument, an initial path is generated along the original drilling direction. If the path intersects with an obstacle plane, two branch paths are generated at the intersection point—a left offset path and a right offset path—with an offset distance of 1.5 times the drill diameter, ensuring a safe distance from the obstacle plane. This process is repeated for each branch path until the path is completely free from all collision interference volumes. The resulting conflict-free trajectory option set contains multiple alternative paths, each labeled with its start point, end point, maximum allowable depth, and minimum safe distance from critical structures. These paths all satisfy the spatial constraints of the third-order operating channel and have no spatial overlap with high-risk structures in the temporal bone dynamic virtual model.
[0176] Perform path resolution mapping and visual transformation on the set of conflict-free trajectory options to obtain the visual cognitive strategy.
[0177] The conflict-free trajectory option set contains multiple alternative paths that satisfy third-order operational channel constraints and do not spatially overlap with high-risk structures. These paths need to be converted into visual information that the surgeon can intuitively understand to support operational decisions. In practice, the three-dimensional coordinate data of each path is first mapped to the coordinate system of the augmented reality navigation view. Spatial registration technology is used to ensure that the path is consistent with the actual spatial position of the surgical field—for example, the starting point of the path corresponds to the current tip position of the instrument, and the ending point corresponds to the target depth.
[0178] To differentiate path priority and safety, each path is assigned distinct visual attributes: recommended paths (lowest overall risk, closest to the target depth) are highlighted in green, auxiliary paths are marked with yellow dashed lines, and hazardous area boundaries are marked with red gradient shading. The path width is dynamically adjusted based on the drill bit diameter—if a 2mm diameter drill bit is currently being used, the path width is set to 2.5mm (with a 0.5mm safety margin) to ensure it matches the actual operating dimensions.
[0179] Based on real-time pose data from the surgical field camera, the projected coordinates of the path on the retina are calculated to ensure that the path shape matches the actual spatial orientation. Simultaneously, key information labels are overlaid on the path: the starting point is labeled "Start Point (X1, Y1, Z1)," turning points are labeled "Turn Point (X2, Y2, Z3)," and the ending point is labeled "End Point (X3, Y3, Z3)," with depth information displayed next to each label (e.g., "Current Depth: 3mm, Maximum Allowable Depth: 5mm"). The resulting visual perception strategy is overlaid onto the surgical field using augmented reality, allowing the surgeon to directly observe the spatial distribution and safety limits of multiple optional paths, providing an intuitive basis for operational decisions.
[0180] Visual cognitive strategies need to be validated for their effectiveness in a virtual environment to ensure they accurately guide the surgeon to avoid risky areas during actual surgery. In practice, a virtual validation scenario consistent with the patient's temporal bone anatomy is first constructed: a dynamic virtual model of the temporal bone is imported, and the same initial conditions as during surgery are set (e.g., instrument type, initial position, drilling direction). Virtual drilling operations are then performed sequentially according to the path outlined in the visual cognitive strategy, recording the instrument position, speed, and spatial relationship with the model at each step.
[0181] The verification process focuses on detecting three types of risks: First, path deviation risk—if the actual position of the instrument deviates from the planned path position by more than 0.3mm (caused by optical tracking system errors or model deformation), the path is marked as "needs correction". Second, collision omission risk—using a virtual collision detection algorithm (based on spatial intersection between the model mesh and the instrument cylinder), it checks for potential collision areas not marked in the visual strategy (such as branches of the facial nerve canal or semicircular canal protrusions). Third, depth loss control risk—comparing the actual depth of virtual drilling with the maximum allowable depth set in the strategy, if the error exceeds 0.2mm (caused by drill wear or speed fluctuations), the path is marked as "depth unreliable".
[0182] For problematic paths discovered during verification, a replanning process is triggered: based on the offset or the location of the collision area, the offset direction of the path is adjusted (e.g., shifted 1mm to the left / right) or the depth is shortened (e.g., adjusted from 5mm to 4.5mm), generating a corrected path and verifying it again until all risk indicators meet the standards. The final safe trajectory obstacle avoidance strategy includes optimized path parameters (e.g., corrected coordinates, depth, and width) and a real-time update mechanism—when a slight change in the anatomical structure is detected, the boundary parameters of the path are automatically adjusted to ensure that a preset safe distance is always maintained from critical structures.
[0183] This embodiment divides the drilling operation into three spatial channels through phase constraints, clearly defining the safe operating range of different areas. Collision detection and obstacle avoidance planning identify interference volumes in real time and generate conflict-free paths to avoid potential collision risks in advance. Path mapping and visual transformation convert the abstract 3D path into intuitive annotations, improving the clarity of operation guidance. Virtual drilling verification verifies the reliability of the path by simulating a real-world scenario, ensuring the effectiveness of the planning.
[0184] Reference Figure 2 As shown, the present invention also provides a virtual reality and real-world image fusion navigation system for temporal bone surgery, applicable to any of the above-mentioned virtual reality and real-world image fusion navigation methods for temporal bone surgery, comprising:
[0185] The acquisition module is used to acquire medical imaging data of the temporal bone region for three-dimensional reconstruction to obtain a dynamic virtual model of the temporal bone.
[0186] The analysis module is used to acquire optical tracking data and instrument spatial pose data, perform dynamic registration and mapping on the dynamic virtual model of the temporal bone, and obtain real-time spatial fusion relationship.
[0187] The association module is used to construct a navigation view by integrating the spatial pose data of the instrument with the real-time spatial fusion relationship, thereby obtaining an augmented reality navigation view.
[0188] The processing module is used to perform structural distance analysis on the augmented reality navigation view to obtain dynamic warning parameters;
[0189] The control module is used to adjust the display based on the augmented reality navigation view and dynamic warning parameters to obtain a real-time navigation strategy.
[0190] This invention provides a virtual reality and real-world image fusion navigation system for temporal bone surgery. By acquiring optical tracking data and instrument spatial pose data in real time and combining them with a dynamic virtual model of the temporal bone for dynamic registration and mapping, it effectively solves the problem of registration lag in traditional static images, ensuring that the virtual model and the actual anatomical structure maintain a high-precision match at all times, thereby improving the accuracy of surgical navigation. An augmented reality navigation view is constructed based on the instrument spatial pose data and the real-time spatial fusion relationship, allowing for intuitive observation of the relative positions of surgical instruments and key anatomical structures, enhancing the real-time nature and interactivity of intraoperative operations, and adapting to rapidly changing surgical field environments. By performing structural distance analysis on the augmented reality navigation view, the distance between instruments and key tissues is calculated in real time, and dynamic warning parameters are generated to promptly alert potential risks, reducing the probability of accidental injury to important nerves and blood vessels, and improving surgical safety. Real-time display adjustment based on the augmented reality navigation view and dynamic warning parameters provides optimal navigation strategies, assists in precise operation, improves surgical efficiency, and reduces operational errors caused by visual errors or delayed feedback.
[0191] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the system and each module described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0192] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A virtual reality and real image fusion navigation method for temporal bone surgery, characterized in that, The method comprises the following steps: obtaining medical image data of the temporal bone region for three-dimensional reconstruction to obtain a dynamic virtual model of the temporal bone; obtaining optical tracking data and instrument spatial pose data, and performing dynamic registration mapping on the dynamic virtual model of the temporal bone to obtain real-time spatial fusion relationship; constructing a navigation view of the instrument spatial pose data and the real-time spatial fusion relationship to obtain an augmented reality navigation view; analyzing the structure distance of the augmented reality navigation view to obtain a dynamic early warning parameter; displaying and controlling the augmented reality navigation view and the dynamic early warning parameter to obtain a real-time navigation strategy; the method for obtaining medical image data of the temporal bone region for three-dimensional reconstruction to obtain a dynamic virtual model of the temporal bone comprises the following steps: aligning the CT image data and the MRI image data of the temporal bone region by multi-modal registration to obtain the medical image data; segmenting the medical image data to obtain an initial structure mask set; reconstructing a three-dimensional surface mesh of the initial structure mask set to obtain a target structure mesh model; assigning spatial attributes to the target structure mesh model and binding dynamic response parameters to obtain the dynamic virtual model of the temporal bone; the dynamic response parameters include the elastic modulus of bone tissue and the strain threshold of nerve structure, so that the deformation effect during the interaction of the surgical instrument can be simulated in real time.
2. The virtual reality and real image fusion navigation method for temporal bone surgery according to claim 1, wherein, the method for obtaining optical tracking data and instrument spatial pose data, and performing dynamic registration mapping on the dynamic virtual model of the temporal bone to obtain real-time spatial fusion relationship comprises the following steps: obtaining optical data collected by a preset head marker array, and performing optical tracking identification to obtain the optical tracking data; analyzing the instrument spatial pose data by a preset surgical instrument tracking adapter; aligning the coordinate systems of the optical tracking data and the instrument spatial pose data to obtain a registration reference data set; performing surface iterative registration on the registration reference data set and the dynamic virtual model of the temporal bone to obtain the real-time spatial fusion relationship.
3. The virtual reality and real image fusion navigation method for temporal bone surgery according to claim 2, wherein, the method for performing surface iterative registration on the registration reference data set and the dynamic virtual model of the temporal bone to obtain the real-time spatial fusion relationship comprises the following steps: locating target points according to the registration reference data set to obtain a target feature point cloud set; performing differential manifold matching on the target feature point cloud set and the dynamic virtual model of the temporal bone to obtain an initial registration matrix; separating the instrument spatial pose data into main vibration mode components to obtain main vibration mode component parameters; performing spatial registration on the initial registration matrix according to the main vibration mode component parameters to generate an anti-vibration registration matrix; verifying the anti-vibration registration matrix for tympanic cavity connectivity to obtain the real-time spatial fusion relationship.
4. The virtual reality and real image fusion navigation method for temporal bone surgery according to claim 1, wherein, the method for constructing a navigation view of the instrument spatial pose data and the real-time spatial fusion relationship to obtain an augmented reality navigation view comprises the following steps: identifying the position of the surgical instrument based on the instrument spatial pose data to obtain surgical instrument position information; performing spatial coordinate transformation on the surgical instrument position information and the real-time spatial fusion relationship to obtain virtual space instrument mapping data; The temporal bone dynamic virtual model is subjected to structure-aware rendering to obtain a multi-layer transparency volume rendering view; The virtual space instrument mapping data and the multi-layer transparency volume rendering view are subjected to deep fusion to obtain a primary view framework; The primary view framework is subjected to physical occlusion relationship reconstruction and optical distortion correction to obtain the augmented reality navigation view.
5. The virtual reality and real image fusion navigation method for temporal bone surgery according to claim 1, wherein, The augmented reality navigation view is subjected to structure distance analysis to obtain a dynamic early warning parameter, including: The augmented reality navigation view is subjected to space trajectory extraction to obtain instrument motion vector information; According to the instrument motion vector information and the temporal bone dynamic virtual model, a near-neighbor structure search is performed to obtain a target space distribution set; The target space distribution set is subjected to dynamic distance field calculation to obtain a multi-structure real-time distance mapping table; Based on the target space distribution set and the multi-structure real-time distance mapping table, trajectory conflict prediction is performed to obtain collision distribution prediction data; The collision distribution prediction data and the multi-structure real-time distance mapping table are subjected to spatial correlation integration to obtain a dynamic early warning parameter.
6. The virtual reality and real image fusion navigation method for temporal bone surgery according to claim 1, wherein, The augmented reality navigation view and the dynamic early warning parameter are subjected to display regulation to obtain a real-time navigation strategy, including: The temporal bone sub-region space weight analysis is performed on the dynamic early warning parameter to obtain a target area regulation factor set; Based on the target area regulation factor set and the augmented reality navigation view, structure visual correlation optimization is performed to obtain a temporal bone specialized navigation view; According to the temporal bone specialized navigation view, drilling trajectory conflict prediction is performed to obtain a safe trajectory obstacle avoidance strategy; According to the safe trajectory obstacle avoidance strategy, the temporal bone specialized navigation view is subjected to operation field of view mapping to obtain an operation guide view; The operation guide view is subjected to optical rendering fusion adaptation to obtain a real-time navigation strategy.
7. The virtual reality and real image fusion navigation method for temporal bone surgery according to claim 6, wherein, The drilling trajectory conflict prediction is performed according to the temporal bone specialized navigation view to obtain a safe trajectory obstacle avoidance strategy, including: The temporal bone specialized navigation view is subjected to stage constraint construction to obtain a three-order operation channel; Based on the three-order operation channel, collision detection is performed on the temporal bone dynamic virtual model to obtain collision interference volume information; According to the collision interference volume information, plane obstacle avoidance planning is performed to obtain a conflict-free trajectory option set; The conflict-free trajectory option set is subjected to resolution path mapping and visual conversion to obtain a visual cognitive strategy; The visual cognitive strategy is subjected to virtual drilling verification to obtain a safe trajectory obstacle avoidance strategy.
8. A virtual reality and real image fusion navigation system for temporal bone surgery, characterized by, The virtual reality and real image fusion navigation method for temporal bone surgery according to any one of claims 1-7, comprising: A collection module, the collection module is used for acquiring medical image data of the temporal bone region for three-dimensional reconstruction to obtain a temporal bone dynamic virtual model; An analysis module, the analysis module is used for acquiring optical tracking data and instrument space pose data, and performing dynamic registration mapping on the temporal bone dynamic virtual model to obtain a real-time space fusion relationship; An association module, the association module is used for constructing a navigation view on the instrument space pose data and the real-time space fusion relationship to obtain an augmented reality navigation view; The processing module is configured to perform structural distance analysis on the augmented reality navigation view to obtain a dynamic early warning parameter. The control module is configured to perform display regulation according to the augmented reality navigation view and the dynamic early warning parameter to obtain a real-time navigation strategy. The medical image data of the temporal bone region is acquired for three-dimensional reconstruction to obtain a dynamic virtual model of the temporal bone, including: The CT image data and the MRI image data of the temporal bone region are acquired for multi-modal registration and alignment to obtain the medical image data; The medical image data is subjected to tissue structure segmentation to obtain an initial structure mask set; The initial structure mask set is subjected to three-dimensional surface mesh reconstruction to obtain a target structure mesh model; The target structure mesh model is subjected to space attribute assignment and dynamic response parameter binding to obtain the dynamic virtual model of the temporal bone. The dynamic response parameter includes the elastic modulus of bone tissue and the strain threshold of neural structure, so that the deformation effect during interaction with a surgical instrument can be simulated in real time.
Citation Information
Patent Citations
Operation holographic navigation method and system based on mixed reality
CN115105207A
Spine three-dimensional visualization surgical navigation system based on mixed reality
CN116492052A