Neurosurgery visual operation navigation system based on multi-modal fusion

By performing three-dimensional registration and feature fusion of intracranial DTI and CT images, combined with intraoperative ultrasound image correction and vascular centerline tracking, the problems of image registration error and slow response of navigation interface in neurosurgery were solved, achieving higher precision and real-time navigation guidance and reducing the risk of accidental injury.

CN120938601AInactive Publication Date: 2025-11-14THE FIRST AFFILIATED HOSPITAL OF ARMY MEDICAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511056149.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies in neurosurgery suffer from problems such as large image registration errors, navigation image drift and superposition errors, slow response of the navigation interface, and inaccurate location of risk areas. In particular, when brain tissue is displaced and blood flow is pulsating during surgery, the determination of surgical instrument paths becomes unstable, increasing the risk of accidental damage to critical structures.

Method used

By aligning intracranial DTI and CT images in three dimensions, the direction of white matter fibers and the direction of bony edges of the temporal bone are extracted. A three-dimensional fusion feature map is generated by combining a multimodal feature fusion algorithm. The edge direction of intraoperative ultrasound images is corrected by a structural similarity algorithm. The trajectory of the vascular centerline is tracked, deformation correction and layer labeling are performed, and a dynamic collision detection algorithm is introduced for navigation interaction to ensure the real-time response and accuracy of the surgical instrument path.

Benefits of technology

It improves the accuracy and stability of image registration, reduces the interference of intraoperative image deviation on navigation results, enhances the immediacy and interpretability of navigation interaction, and is suitable for guiding complex surgical paths around high-risk structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120938601A_ABST
    Figure CN120938601A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-modal registration, in particular to a neurosurgery visual operation navigation system based on multi-modal fusion, which comprises a feature fusion module, an image correction module, a deformation correction module, a layer labeling module and a navigation interaction module. According to the method, a DTI image and a CT image are aligned through three-dimensional pairing, white matter fiber and temporal bone edge directions are extracted, feature integrity is enhanced, edges are corrected through a structural similarity algorithm, space mapping precision is improved, a blood vessel center line and an ultrasonic image edge point track are combined, deformation is corrected, and drift and artifacts in an operation are restrained; the structure continuity of the blood vessel and the fiber is guaranteed, the positions of the white matter fiber and the blood vessel are integrated, a consistent directional diagram layer is formed, the navigation response sensitivity is improved through three-dimensional distance evaluation, a self-matching adjustment mechanism is set, the feature fusion and correction precision is enhanced, image deviation interference is reduced, and the navigation real-time performance and the analytic performance are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimodal registration technology, and in particular to a neurosurgical visualization and surgical navigation system based on multimodal fusion. Background Technology

[0002] Multimodal registration technology encompasses multiple areas, including medical image processing, image registration, image fusion, and intraoperative navigation. Its core content involves the spatial alignment and fusion of medical image data from different modalities, such as magnetic resonance imaging (MRI), computed tomography (CT), and ultrasound images, to obtain more comprehensive tissue structure information. In delicate surgical fields such as neurosurgery, surgeons often rely on multiple image modalities to accurately identify lesion locations, tissue boundaries, and functional area distributions. Therefore, multimodal registration technology has become a crucial component of preoperative planning and intraoperative guidance. This technology also includes image preprocessing, feature extraction, registration model construction, transformation relationship optimization, and 3D visualization reconstruction. These methods improve registration accuracy and computational efficiency, enhancing the practicality and reliability of intraoperative navigation.

[0003] One such system is a multimodal fusion-based neurosurgical visualization navigation system. In neurosurgical procedures, preoperative magnetic resonance imaging (MRI) images and intraoperative ultrasound images are used as the primary data sources. A method based on image grayscale mutual information registration is employed to register images of different modalities, and then the registered data is used to construct an intraoperative three-dimensional visualization anatomical model. Simultaneously, the spatial positioning information of surgical instruments is used to map their real-time positions onto the visualization model, forming the surgical navigation interface. This patent also incorporates fusion rules to fuse tissue boundary features from intraoperative images with anatomical information from preoperative images, enhancing key anatomical landmarks in the navigation model. Furthermore, volume rendering technology is used to render the three-dimensional images in real time, enabling real-time visualization updates of the intraoperative navigation images.

[0004] While existing technologies have achieved the registration and fusion of MRI and ultrasound images and constructed a three-dimensional navigation view, limitations remain in feature selection and registration mechanisms. Conventional grayscale mutual information, as the primary matching criterion, cannot adequately identify tissue regions with significant structural differences, especially during surgery due to brain tissue displacement, blood flow pulsation, or obstruction by bony structures, which can easily lead to registration errors. Furthermore, the lack of refined utilization of the directional characteristics of blood vessels and nerve fibers results in significant deviations in the location of risk areas in the navigation image. Because there is no real-time tracking and correction of dynamic tissue drift during surgery, image drift superposition errors may occur during intraoperative navigation, leading to instability in surgical instrument path determination. In addition, the navigation interface lacks a dynamic response mechanism with risk structures, which can easily cause slow reactions during sudden operations, increasing the probability of accidental damage to critical structures. For example, in temporal lobe surgery, slight deformation of brain tissue due to traction occurs during surgery. Without real-time correction and interactive warnings, image navigation may not reflect the true spatial relationship between the actual instruments and risk points, resulting in reduced navigation accuracy and increasing the difficulty and risk of intraoperative judgment. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the existing technology and propose a neurosurgical visualization navigation system based on multimodal fusion.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a neurosurgical visualization navigation system based on multimodal fusion includes:

[0007] The feature fusion module acquires intracranial DTI images and CT images, aligns and pairs them in three dimensions, extracts the direction of white matter fibers and the direction of bony edges of the temporal bone, performs feature fusion through a multimodal feature fusion algorithm, generates a three-dimensional fused feature map, and transmits it to the image correction module.

[0008] The image correction module acquires intraoperative temporal lobe ultrasound images based on the three-dimensional fusion feature map and analyzes the edge direction. It compares and corrects the bony edge direction of the temporal bone and the edge direction of the temporal lobe through a structural similarity algorithm, generates image correction results, and transmits them to the deformation correction module.

[0009] The deformation correction module obtains the CT vessel centerline based on the image correction results and tracks the trajectory of edge points in the intraoperative brain ultrasound image. It calculates the displacement direction and velocity changes and selects a stable set of edge points. It then performs back projection and linear interpolation correction on the vessel centerline in combination with the white matter fiber direction, generates a deformation correction image, and transmits it to the layer annotation module.

[0010] The layer annotation module extracts the DTI fiber alignment direction and CT vessel centerline position based on the deformation correction image, and generates layers according to the direction and spatial relationship to generate a three-dimensional layer set.

[0011] The navigation interaction module, based on the three-dimensional layer set, obtains the spatial position of surgical instruments during the operation, performs three-dimensional distance judgment between the instrument tip and the risk point, calls the dynamic collision detection algorithm to set a dynamic threshold to trigger the red layer to flash, adjusts the layer precision according to the interaction delay, and outputs the navigation view interaction data results.

[0012] As a further aspect of the present invention, the three-dimensional fusion feature map specifically includes white matter fiber direction distribution data, temporal bone bony margin spatial data, and paired three-dimensional coordinate information. The image correction result includes a temporal lobe margin direction vector set and a structural similarity comparison coefficient. The deformation correction image specifically refers to the corrected vascular centerline trajectory, stable edge point spatial set, and displacement direction. The three-dimensional layer set includes directional relationship angles, vascular centerline spatial positioning, and layer spatial sequence encoding. The navigation view interactive data result includes device risk distance values, collision triggering state values, and layer rendering accuracy parameters.

[0013] As a further aspect of the present invention, the feature fusion module includes:

[0014] The image matching submodule acquires three-dimensional images of intracranial DTI and CT images, calls the three-dimensional coordinate registration function to perform rigid transformation on the CT image and align it with the DTI image, compares the coordinate distance between the center point of multiple voxels in the DTI image and the corresponding point in the CT image, and generates a three-dimensional voxel matching coordinate set.

[0015] The structure extraction submodule, based on the three-dimensional voxel matching coordinate set, locates the temporal bone edge region in the CT image and obtains the gradient direction of the region edge. At the same time, it extracts the white matter fiber direction of the corresponding voxel in the DTI image, calculates the angle between the directions, and generates a spatial direction corresponding angle distribution map.

[0016] The feature fusion submodule, based on the spatial direction corresponding angle distribution map, calculates the angular consistency between the CT edge direction and the DTI fiber direction through a multimodal feature fusion algorithm, constructs a three-dimensional orientation field structure map, reorganizes the fusion direction of each spatial location and transforms it into a tensor structure, and obtains a three-dimensional fusion feature map.

[0017] As a further aspect of the present invention, the image correction module includes:

[0018] The image acquisition submodule acquires ultrasound image frame data of the temporal lobe region during surgery, calls the three-dimensional coordinates in the three-dimensional fusion feature map to locate the region within the ultrasound image, extracts pixels of the temporal lobe region edge direction based on the grayscale edge distribution of the image, and performs linear interpolation calculation on each group of directions to generate the temporal lobe edge direction.

[0019] The structural comparison submodule obtains the direction of the bony edge of the temporal bone based on the direction of the temporal lobe edge, calculates the consistency of the two types of direction values ​​within the same spatial range through a structural similarity algorithm, records the similarity interval corresponding to the position, and obtains the direction similarity interval.

[0020] The orientation correction submodule performs orientation adjustment operations on the orientation data located in the three-dimensional fusion feature map point by point based on the orientation similarity interval. It performs orientation replacement and records the positions where there is a deviation between the temporal bone orientation and the temporal lobe orientation to obtain the image orientation correction result.

[0021] As a further aspect of the present invention, the deformation correction module includes:

[0022] The centerline extraction submodule, based on the image orientation correction results, locates the strong echo regions of blood vessel segments in the CT image, performs morphological filtering on the regions, collects blood vessel boundary pixels and arranges the point positions in sequence, extracts the centerline structure based on spatial connectivity, and obtains the blood vessel centerline length sequence.

[0023] The edge tracking submodule selects a set of edge points in a continuous time frame in the intraoperative brain ultrasound image based on the blood vessel centerline length sequence, compares the spatial positions of edge points in adjacent frames and calculates the displacement direction and velocity, and filters out abnormal points with positional changes through a stability judgment threshold to obtain a stable edge displacement range.

[0024] The trajectory correction submodule calls the stable edge displacement interval and the white matter fiber direction to sequentially perform reverse point projection operation on the spatial coordinates in the blood vessel centerline length sequence, and at the same time performs linear interpolation translation correction within the original coordinate distribution interval to generate a deformation correction image.

[0025] As a further aspect of the present invention, the stability judgment threshold is determined by calculating the rate of change of the spatial position of each edge point between adjacent frames, taking the average value of all rate values ​​as the basic reference, and combining it with a range of three times the standard deviation to construct a tolerance range for the movement amplitude of the edge point.

[0026] As a further aspect of the present invention, the layer annotation module includes:

[0027] The orientation extraction submodule, based on the deformation correction image, identifies the fiber orientation distribution area in the DTI image, filters points whose orientation value changes within the stability threshold, locates the blood vessel centerline in the CT image, calculates and maps the angle change between the orientation point and the blood vessel line segment, and generates the orientation relationship angle.

[0028] The layer construction submodule constructs a fiber orientation layer based on the directional relationship angle and a structural positioning layer based on the spatial path of the blood vessel centerline. Each layer is labeled according to its spatial sequence number in the 3D image, and the structural type and orientation sequence of the region in each layer are marked to obtain the 3D layer order.

[0029] The layer fusion submodule calls the three-dimensional layer sequence, merges the direction layer and structure layer according to their spatial correspondence, sorts all layers in spatial position and adjusts the coordinate pointing consistency, assigns a unique identification number to each layer and integrates them to establish a three-dimensional layer set.

[0030] As a further aspect of the present invention, the stability threshold is set based on the variance and average value of the orientation values ​​of the entire region of the DTI image. First, the average change value of all orientation points is statistically analyzed, and then the standard deviation is calculated. The threshold is set as the average change value plus one standard deviation, which is used to screen regions with stable orientation and remove points with drastic orientation fluctuations.

[0031] As a further aspect of the present invention, the navigation interaction module includes:

[0032] The location acquisition submodule, based on the three-dimensional layer set, collects the spatial coordinate information of the end of the surgical instrument during the operation, combines it with the risk areas marked in the layer, calculates the Euclidean distance between the end of the instrument and all risk points in each frame, and filters the effective point set with the best distance in the current frame to generate the instrument risk distance.

[0033] The collision judgment submodule, based on the device risk distance, calls the dynamic collision detection algorithm to calculate the product of the device movement speed and the layer refresh frequency as a dynamic safety threshold, compares the minimum distance value in each time slice with the threshold, marks the trigger state, and obtains the collision trigger state.

[0034] The layer response submodule, based on the collision trigger state, controls the red area at the corresponding spatial position in the 3D layer set to perform periodic flashing, compares the layer frame rate data with the interaction response delay record value, adjusts the rendering precision of the area in the layer, and establishes the navigation view interaction data result.

[0035] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0036] In this invention, by pairing and aligning intracranial DTI and CT images in three-dimensional space, and simultaneously extracting information on white matter fiber direction and temporal bone bony margin direction, an input basis for fusing structural and functional features is established, making the features used for registration more complete and recognizable. Based on this, the fused feature map is compared and corrected for structural similarity with the edge direction of intraoperative ultrasound images, establishing a more stable correspondence between edge features across different modalities and improving the accuracy of spatial mapping between images. Furthermore, by combining the vascular centerline and intraoperative ultrasound edge point trajectory information, image deformation is corrected through backprojection and linear interpolation, while a stable edge point set is selected to suppress intraoperative tissue drift and image artifacts, thereby ensuring the spatial continuity of vascular and fibrous structures. The white matter fiber direction and vascular spatial position are integrated to form a layer set with consistent direction, and a dynamic assessment mechanism for spatial position mapping and three-dimensional distance of risk areas is introduced to ensure that the real-time response to surgical instrument paths in the navigation view is both alerting and sensitive. By establishing an adaptive adjustment mechanism between layer accuracy and interaction delay, a dynamic balance is achieved between accurate alerts and navigation performance. The above processing logic enhances the fusion and correction accuracy between multimodal features, improves registration robustness, suppresses the interference of intraoperative image deviations on navigation results, and enhances the immediacy and interpretability of navigation interaction, making it suitable for guiding complex surgical paths around high-risk structures. Attached Figure Description

[0037] Figure 1 This is a system flowchart of the present invention;

[0038] Figure 2 This is a system block diagram of the present invention. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0040] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0041] Please see Figure 1A neurosurgical visualization and navigation system based on multimodal fusion includes:

[0042] The feature fusion module acquires intracranial DTI images and CT images, aligns and pairs them in three dimensions, extracts the direction of white matter fibers and the direction of bony edges of the temporal bone, performs feature fusion through a multimodal feature fusion algorithm, generates a three-dimensional fused feature map, and transmits it to the image correction module.

[0043] The image correction module acquires intraoperative temporal lobe ultrasound images based on three-dimensional fusion feature maps and analyzes the edge direction. It compares and corrects the direction of the bony edge of the temporal bone and the edge direction of the temporal lobe through a structural similarity algorithm, generates image correction results, and passes them to the deformation correction module.

[0044] The deformation correction module obtains the CT vessel centerline based on the image correction results and tracks the trajectory of edge points in the intraoperative brain ultrasound image. It calculates the displacement direction and velocity changes and selects a stable set of edge points. It then performs back projection and linear interpolation correction on the vessel centerline in combination with the white matter fiber direction, generates a deformation correction image, and passes it to the layer annotation module.

[0045] The layer annotation module extracts the DTI fiber alignment direction and CT vessel centerline position from the deformation-corrected image, generates layers according to the direction and spatial relationship, and generates a three-dimensional layer set.

[0046] The navigation interaction module, based on a set of 3D layers, obtains the spatial position of surgical instruments during the operation, performs 3D distance judgment between the instrument tip and risk points, calls a dynamic collision detection algorithm to set a dynamic threshold to trigger the red layer to flash, adjusts the layer precision according to the interaction delay, and outputs the navigation view interaction data results.

[0047] The three-dimensional fusion feature map specifically includes white matter fiber direction distribution data, temporal bone bony margin spatial data, and paired three-dimensional coordinate information. The image correction results include temporal lobe margin direction vector set and structural similarity comparison coefficient. The deformation correction image specifically refers to the corrected vascular centerline trajectory, stable edge point spatial set, and displacement direction. The three-dimensional layer set includes directional relationship angles, vascular centerline spatial positioning, and layer spatial sequence encoding. The navigation view interactive data results include device risk distance value, collision triggering state value, and layer rendering accuracy parameters.

[0048] Please see Figure 2 The feature fusion module includes:

[0049] The image matching submodule acquires three-dimensional images of intracranial DTI and CT images, calls the three-dimensional coordinate registration function to perform rigid transformation on the CT image and align it with the DTI image, compares the coordinate distance between the center point of multiple voxels in the DTI image and the corresponding point in the CT image, and generates a three-dimensional voxel matching coordinate set.

[0050] The DTI and CT image datasets for patient P001 were obtained via the DICOM interface. In the DTI images, the entrance point of the auditory nerve canal (coordinates (x1, y1, z1) = (32.5, 45.2, 18.7) mm) and the midpoint of the internal auditory canal (x2, y2, z2) = (29.8, 47.6, 20.3) were selected as reference points. In the CT images, the corresponding anatomical points were manually labeled with coordinates (34.1, 44.9, 19.2) and (31.2, 46.8, 21.0). The rotation vector R and translation vector T were calculated using a three-dimensional rigid body registration function. After optimization using an iterative nearest-point algorithm, R was obtained as follows: [[0.998, 0.012, -0.059], [-0.015, 0.999, 0.041], [0.058, -0.042, 0.997]], T = (1.6, -0.3, 0.5) mm. After transforming the CT image coordinate system, the registration error of the reference point is calculated to be 0.21 mm. Traverse the 512×512×200 voxel space in the DTI image and establish a spatial grid with a step size of 3 mm. Calculate the Euclidean distance between the center point of each grid and the corresponding voxel in the CT image. When the distance is ≤2 mm, record the matching coordinates to generate a three-dimensional coordinate set including 28,672 sets of matching coordinates, as shown in Table 1.

[0051] Table 1 Image Registration Parameters

[0052]

[0053]

[0054] The structure extraction submodule, based on a three-dimensional voxel matching coordinate set, locates the temporal bone edge region in CT images and obtains the gradient direction of the region edge. At the same time, it extracts the white matter fiber direction of the corresponding voxel in DTI images, calculates the angle between the directions, and generates a spatial direction corresponding angle distribution map.

[0055] In the matching coordinate set, the coordinates (x, y, z) = (35, 50, 22) mm of the petrous part of the temporal bone are selected, where g represents the CT gradient vector. The spatial gradient is calculated using the 3×3×3 neighborhood voxel values ​​of the CT image. The difference ΔH between adjacent voxel CT values ​​in the x-direction is taken at the coordinates (35, 50, 22) mm of the petrous part of the temporal bone. x =160-145=15HU, y-direction difference ΔH y =155-150=5HU, z-direction difference ΔH z=158-152=6HU, combined with the volume element spacing Δd=1.5mm, we get g=(15 / 1.5, 5 / 1.5, 6 / 1.5)=(10.0, 3.3, 4.0), f represents the DTI fiber vector: using voxels with a DTI image FA value of 0.8, the principal feature vector is (0.707, 0.707, 0), after normalization we get ||f||=1.0, N g Representing the gradient normalization factor: the maximum gradient magnitude of the entire CT image, max(||g||), is 150 HU / mm (from a database of 50 samples), N f Fiber normalization factor: The fiber modulus corresponding to the maximum FA value in DTI images is 1.5 (FA value range 0.2-0.8 corresponds to modulus length 0.5-1.5). λ represents the balance coefficient: Through comparison of pre- and post-operative data from 20 cases, the registration error is minimized when λ = 0.3, calculated as λ = 0.5 × (1-FA) (0.3 when FA = 0.8). FA is an index for quantifying the orientational consistency of white matter fibers in diffusion tensor imaging (DTI). ∈ represents the smoothing constant: set to 1 × 10⁻⁶. ;5 To prevent the denominator from being zero, substitute into the formula:

[0056] Substitute into the formula to calculate Numerator = 9.43 + 0.3 × 0.67 = 9.63, Denominator = 11.2 × 1.0 + 0.00001 = 11.2, θ = 9.63 / 11.2 = 0.86 radians ≈ 49.3°. The calculated θ = 49.3° is within the normal range of the temporal bone-auditory nerve angle (45°-55°), indicating that the fiber orientation and the spatial relationship with the bony structure are normal.

[0057] The feature fusion submodule, based on the spatial direction corresponding angle distribution map, calculates the angular consistency between the CT edge direction and the DTI fiber direction through a multimodal feature fusion algorithm, constructs a three-dimensional orientation field structure map, reorganizes the fusion direction of each spatial location and transforms it into a tensor structure, and obtains a three-dimensional fusion feature map;

[0058] Taking the angle values ​​θ1 = 0.75 rad at coordinates (38, 52, 25) and θ2 = 0.35 rad at coordinates (40, 55, 28), an angle consistency weight w = 1 / (1+10θ) is constructed, yielding w1 = 0.12 and w2 = 0.22. For the CT gradient direction g1 = (9.8, 2.5, 3.2) and the DTI fiber direction f2 = (0.866, 0.5, 0), fusion is performed: v = 0.12g1 + 0.22f2 = (1.18, 0.30, 0.38) + (0.19, 0.11, 0) = (1.37, 0.41, 0.38). After normalization, v represents the fusion direction vector (0.92, 0.28, 0.26). A three-dimensional tensor is constructed: The tensor eigenvalue λ1 = 0.95, the tensor eigenvalue λ 2,3 =0.05 indicates significant consistency in the principal direction, generating a 512×512×200 tensor field.

[0059] Please see Figure 2 The image correction module includes:

[0060] The image acquisition submodule acquires ultrasound image frame data of the temporal lobe region during surgery, calls the three-dimensional coordinates in the three-dimensional fusion feature map to locate the region within the ultrasound image, extracts pixels of the temporal lobe region edge direction based on the grayscale edge distribution of the image, and performs linear interpolation calculation on each group of directions to generate the temporal lobe edge direction.

[0061] An intraoperative ultrasound image of the temporal lobe region of patient P002 was acquired using an intraoperative ultrasound probe. The image resolution was 0.2 mm. The origin of the ultrasound image coordinate system (x0, y0, z0) = (120, 85, 40) was located in the 3D fusion feature map. A region of interest (ROI) of 128×128 pixels in the temporal lobe edge region of the ultrasound image was selected, and the gray-level gradient G in the x-direction was calculated for each pixel (i, j). x = (I(i+1,j)-I(i-1,j)) / 2Δx, where the voxel spacing of the ultrasound image is Δx, Δx = 0.2mm, and I(i,j) represents the gray value of the pixel in the i-th row and j-th column of the ultrasound image. When I(65,70) = 150, I(66,70) = 145, and I(64,70) = 148, G x = (145-148) / (2×0.2) = -7.5, similarly calculate the gradient G in the y-direction. y = (I(i,j+1)-I(i,j-1)) / 2Δy, for the gradient vector (G x G y Cubic spline interpolation was performed to generate a direction field with an accuracy of 0.1 mm, including the interpolated direction vector φ = (0.866, 0.5, 0) at coordinates (122.3, 86.7, 41.5) mm. After normalization, ||φ|| = 1.0, as shown in Table 2.

[0062] Table 2. Calculation of Ultrasound Image Gradient

[0063] coordinate x gradient y gradient (122.3,86.7) -7.5 4.2 (122.4,86.8) -6.8 5.1

[0064] The structural comparison submodule obtains the direction of the bony edge of the temporal bone based on the direction of the temporal lobe edge, calculates the consistency of the two types of direction values ​​within the same spatial range through a structural similarity algorithm, records the similarity interval corresponding to the position, and obtains the direction similarity interval.

[0065] Select the ultrasonic direction φ at spatial location j=1 (1) = (0.866, 0.5, 0) and ψ in the CT temporal bone direction (1) = (0.707, 0.707, 0), M represents the number of spatial locations: 1000 overlapping voxels are taken from the ultrasound image and CT registration area, and after grid sampling, they actually participate in the calculation M = 872, 2.φ (j) Representing the ultrasound direction vector: φ, calculated from the gradient of the ultrasound image at coordinate j=25 in the example. x =0.5, φ y =0.866, φ z =0, modulus verification 3.ψ (j) Represents the CT direction vector: the CT gradient direction ψ at the corresponding coordinates. x =0.707, ψ y =0.707, ψ z =0, modulus 4.δ j Representative offset adjustment amount: Calculated based on the axial resolution of 0.3 mm for an ultrasonic probe with a frequency of 5 MHz. j =

[0066] 0.3 / (2×5)=0.03mm, adjusted to 0.05 after statistical analysis of 20 samples, 5.ε j Smoothing constant: set to 1×10 ;6 To prevent the denominator from being zero, substitute the values ​​into the formula to calculate:

[0067] The numerator is calculated as follows: (0.5 × 0.707 + 0.866 × 0.707 + 0 × 0) - 0.05 = 0.92. The overall κ is calculated sim =0.55 falls within the moderate similarity range of 0.4-0.6, indicating a correctable deviation between the intraoperative ultrasound and preoperative CT orientations. This result triggered orientation replacement for 32% of the low similarity points <0.5, reducing the mean curvature of the fusion orientation field from 0.12 to 0.07.

[0068] The orientation correction submodule performs orientation adjustment operations on the orientation data located in the 3D fusion feature map point by point based on the orientation similarity interval. It performs orientation replacement and records the positions where there is a deviation between the orientation of the temporal bone and the orientation of the temporal lobe to obtain the image orientation correction result.

[0069] The ultrasonic direction was detected at coordinates (125.6, 88.2, 43.1) mm.

[0070] φ (25) = (0.5, 0.866, 0) and ψ in the CT direction (25) = (0.707, 0.707, 0) Similarity 0.52, below the threshold of 0.6, calculate the orientation deviation angle arccos(0.5×0.707+0.866×0.707)=30°, when the deviation is >25°, perform orientation replacement, update the position orientation in the fused feature map to (0.707, 0.707, 0), record the number of replacement operations as 153 times / thousand points, generate the correction orientation field, and reduce the maximum local curvature from 0.15 to 0.08.

[0071] Please see Figure 2 The deformation correction module includes:

[0072] The centerline extraction submodule locates the strong echo regions of blood vessel segments in CT images based on image orientation correction results, performs morphological filtering on the regions, collects blood vessel boundary pixels and arranges the point positions in sequence, extracts the centerline structure based on spatial connectivity, and obtains the blood vessel centerline length sequence.

[0073] Based on the corrected orientation field data, the petrous temporal bone vascular region (CT value > 300 HU) was selected in the CT image. Morphological operations were performed using 3×3 circular structuring elements to remove noise and extract the vascular boundaries. For the pixel at coordinates (x1, y1, z1) = (35.2, 50.1, 22.3) mm, the standard deviation of the CT value of the 8-neighborhood was calculated as σ = 15.6 HU (threshold > 10 HU was considered a boundary). Contour points were extracted every 0.5 mm along the z-axis. Delaunay triangulation was used to establish the spatial topology, and morphological filtering was performed on the region. At slice z = 22.5 mm, the boundary point sequence P1(35.0, 50.3), P2(35.5, 50.0), P... 12 (36.2, 49.8) The center points of multiple slices are calculated by skeletonization algorithm, including the center point C1 = (35.6, 50.1) mm of slice z = 22.5 mm. The center points of adjacent slices are connected to form a center line. The total length L = 18.7 mm is measured (the distance between adjacent points ≤ 0.3 mm). The coordinate sequence of 56 center points is generated, as shown in Table 3.

[0074] Table 3. Coordinates of the Central Line of Vascular Vessels

[0075]

[0076]

[0077] The edge tracking submodule selects a set of edge points in a continuous time frame of intraoperative brain ultrasound images based on the blood vessel centerline length sequence. It compares the spatial positions of edge points in adjacent frames and calculates the displacement direction and velocity. It filters out abnormal points with positional changes through a stability judgment threshold to obtain a stable edge displacement range.

[0078] In intraoperative brain ultrasound images, edge point sets 100-105 were selected from a temporally continuous frame. Edge point set E was extracted from frame 100. 100 ={(122.3, 86.7),(122.5, 86.9)}, analyze the point set E corresponding to frame 101. 101 ={(122.4, 86.8), (122.6, 87.0)}, and simultaneously calculate the displacement vector Δd = (0.1, 0.1) mm for point (122.3, 86.7), velocity v = 0.14 mm / frame (4.2 mm / s for a frame rate of 30 fps), and set the velocity threshold v. max =5 mm / s (determined based on the maximum blood flow velocity of 10 samples). Points with a velocity of 5.3 mm / s at (123.0, 87.2) were excluded. Points with a displacement direction ≤15° of the average direction of the vessel centerline (including point (122.5, 86.9) with a 12° angle) were retained. 28 stable displacement points were obtained, with velocity ranges...

[0079] [2.1,4.8]mm / s.

[0080] The trajectory correction submodule calls the stable edge displacement interval and white matter fiber direction to sequentially perform reverse point projection operation on the spatial coordinates in the blood vessel centerline length sequence, and at the same time performs linear interpolation translation correction within the original coordinate distribution interval to generate a deformation correction image.

[0081] Extract centerline point C 25 (36.1, 50.3, 23.0) mm, based on the stable edge displacement range, corresponding to the edge point of ultrasonic frame 102. Calculate the back projection vector v = (-0.2, 0.1) mm (making an angle of 18° with the white matter fiber direction (0.707, 0.707, 0)), and insert correction points C within the original coordinate interval of 0.3 mm. 25(36.1-0.2×0.3,50.3+0.1×0.3)=(36.04,50.33)mm. Linear interpolation was performed on 5 consecutive center points. After correction, the radius of curvature of the center line increased from 2.1mm to 3.5mm, generating a 512×512 pixel corrected image. The measurement error of the blood vessel diameter decreased from ±0.15mm to ±0.08mm.

[0082] Please see Figure 2 The layer annotation module includes:

[0083] The orientation extraction submodule, based on the deformation-corrected image, identifies the fiber orientation distribution area in the DTI image, filters points whose orientation value changes are within the stability threshold, locates the blood vessel centerline in the CT image, calculates and maps the angle change between the orientation point and the blood vessel line segment, and generates the orientation relationship angle.

[0084] Based on the deformation-corrected image DTI-0032, the fiber orientation vector f1 = (0.707, 0.707, 0) is extracted at coordinates (x, y, z) = (36.5, 51.2, 23.8) mm. The standard deviation of the orientation variation over 10 consecutive frames is σ = 0.08 rad (threshold set at 0.1 rad), which is considered a stable point. Simultaneously, the vessel centerline point C is also acquired. 15 The tangent direction at (36.3, 51.0, 23.8) mm is v1 = (0.866, 0.5, 0). The included angle θ = arccos(0.707 × 0.866 + 0.707 × 0.5) = 15° is calculated and mapped to a 512 × 512 angle, generating 3,856 valid angle data points, as shown in Table 4.

[0085] Table 4 Distribution of Fiber-Vessel Angle

[0086] Angle range Number of data points 0-10 892 10-20 1,564

[0087] The layer construction submodule constructs fiber orientation layers based on directional relationships and structural positioning layers based on the spatial path of the blood vessel centerline. Each layer is labeled according to its spatial sequence number in the 3D image, and the structural type and orientation sequence of the region within each layer are marked to obtain the 3D layer order.

[0088] Within a three-dimensional spatial grid (120:0.5:130, 85:0.5:90, 40:0.5:45) mm, spatial numbers L001-L020 (layer thickness 0.5 mm) are assigned to the fiber orientation layers. Layer L005 (z = 42.5 mm) is marked as a temporal lobe fiber bundle, and the orientation sequence [0.707, 0.707, 0] × 56 sets are recorded. The structural positioning layer S003 corresponds to the vascular segment V02 (centerline point 15-32). The type code CT-VASCULAR is set, and the orientation vector (0.866, 0.5, 0) and the type label vestibular aqueduct are stored at the coordinates (125.6, 88.2, 43.1) mm of layer L007.

[0089] The layer fusion submodule calls the 3D layer order, merges the orientation layer and structure layer according to the spatial correspondence, sorts all layers in spatial position and adjusts the coordinate pointing consistency, assigns a unique identification number to each layer and integrates them to create a 3D layer set.

[0090] Align the fibrous layer L012 (z = 44.0 mm) with the structural layer S008 (z = 44.0 mm), where z represents the coordinate in the sagittal plane. Check that the direction vector (0.707, 0.707, 0) at coordinates (127.3, 89.1, 44.0) mm has a pointing deviation ≤ 5° from the vascular tangent (0.866, 0.5, 0). Merge them according to their spatial correspondence and adjust the rotation vector.

[0091] The coordinates of R = [[0.999, 0.031, -0.012], [-0.031, 0.999, 0.012], [0.012, -0.012, 0.999]] are unified, and IDF-012-008 is assigned to the fusion layer. The integrated dataset includes 120 fusion layers and is stored as a 512×512×200 tensor field with a voxel resolution of 0.2×0.2×0.5.

[0092] Please see Figure 2 The navigation interaction module includes:

[0093] The location acquisition submodule, based on a set of 3D layers, collects the spatial coordinate information of the end of the surgical instrument during the operation. Combined with the risk areas marked in the layers, it calculates the Euclidean distance between the end of the instrument and all risk points in each frame, and selects the effective set of points with the best distance in the current frame to generate the instrument risk distance.

[0094] Based on the 3D layer set F-012-008, the spatial coordinates (x, y, y) of the surgical instrument tip at time stamp t = 15.3s were obtained. t y t , z t= (127.1, 89.0, 44.0) mm, traverse the 12 risk points R1(126.8, 88.9, 43.9) marked in the layer to R 12 (127.5, 89.2, 44.1) mm, x t y t and z t Calculate the Euclidean distance using the spatial coordinates at time t = 15.3s. Including the distance at R3 (127.0, 89.1, 44.0). Filter valid point sets with a distance of <2mm

[0095] {R3(0.14), R5(1.78), R8(1.95)}, generate the minimum risk distance of 0.14mm for the current frame, as shown in Table 5.

[0096] Table 5. Distance between Medical Devices and Risk Points

[0097] Risk point number distance R3 0.14 R5 1.78

[0098] The collision judgment submodule calculates the product of the device's movement speed and the layer refresh frequency as a dynamic safety threshold based on the device's risk distance, and compares the minimum distance value in each time slice with the threshold, marks the trigger state, and obtains the collision trigger state.

[0099] Set the machine's movement speed v = 3 mm / s (calculated from a displacement of 9 mm between t = 15.0 and 15.3 s), the layer refresh rate f = 30 Hz, and calculate the dynamic safety threshold T. d =v×(1 / f)=3×0.033=0.1mm. When t=15.3s, the minimum distance is 0.14mm>0.1mm, and the flag status is 0 (safe). When t=15.4s, the minimum distance is detected as 0.08mm<0.1mm, and the trigger status is 1 (warning). Record the number of triggers for 3 consecutive frames.

[0100] The layer response submodule, based on the collision trigger state, controls the red area at the corresponding spatial location in the 3D layer set to perform periodic blinking, compares the layer frame rate data with the interaction response delay record value, adjusts the rendering precision of the area in the layer, and establishes the navigation view interaction data results.

[0101] When state 1 is detected, the risk point R3 (127.0, 89.1, 44.0) mm is located, the area flashing frequency is set to 5 Hz (period 200 ms), the current rendering delay of the detection system is τ = 45 ms (threshold 50 ms), the rendering accuracy is maintained at 512×512, and when the delay > 50 ms, it is downgraded to 256×256. The transparency of the 12 risk areas in layer F-012-008 is updated from 0.8 to 0.9, and a navigation view dataset including 1536 state marker points is generated.

[0102] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A neurosurgical visualization and navigation system based on multimodal fusion, characterized in that, The system includes: The feature fusion module acquires intracranial DTI images and CT images, aligns and pairs them in three dimensions, extracts the direction of white matter fibers and the direction of bony edges of the temporal bone, performs feature fusion through a multimodal feature fusion algorithm, generates a three-dimensional fused feature map, and transmits it to the image correction module. The image correction module acquires intraoperative temporal lobe ultrasound images based on the three-dimensional fusion feature map and analyzes the edge direction. It compares and corrects the bony edge direction of the temporal bone and the edge direction of the temporal lobe through a structural similarity algorithm, generates image correction results, and transmits them to the deformation correction module. The deformation correction module obtains the CT vessel centerline based on the image correction results and tracks the trajectory of edge points in the intraoperative brain ultrasound image. It calculates the displacement direction and velocity changes and selects a stable set of edge points. It then performs back projection and linear interpolation correction on the vessel centerline in combination with the white matter fiber direction, generates a deformation correction image, and transmits it to the layer annotation module. The layer annotation module extracts the DTI fiber alignment direction and CT vessel centerline position based on the deformation-corrected image, generates layers according to the direction and spatial relationship, and generates a three-dimensional layer set.

2. The neurosurgical visualization and navigation system based on multimodal fusion according to claim 1, characterized in that, The three-dimensional fusion feature map specifically includes white matter fiber direction distribution data, temporal bone bony margin spatial data, and paired three-dimensional coordinate information. The image correction result includes temporal lobe margin direction vector set and structural similarity comparison coefficient. The deformation correction image specifically refers to the corrected vascular centerline trajectory, stable edge point spatial set, and displacement direction. The three-dimensional layer set includes directional relationship angles, vascular centerline spatial positioning, and layer spatial sequence encoding.

3. The neurosurgical visualization and navigation system based on multimodal fusion according to claim 1, characterized in that, The feature fusion module includes: The image matching submodule acquires three-dimensional images of intracranial DTI and CT images, calls the three-dimensional coordinate registration function to perform rigid transformation on the CT image and align it with the DTI image, compares the coordinate distance between the center point of multiple voxels in the DTI image and the corresponding point in the CT image, and generates a three-dimensional voxel matching coordinate set. The structure extraction submodule, based on the three-dimensional voxel matching coordinate set, locates the temporal bone edge region in the CT image and obtains the gradient direction of the region edge. At the same time, it extracts the white matter fiber direction of the corresponding voxel in the DTI image, calculates the angle between the directions, and generates a spatial direction corresponding angle distribution map. The feature fusion submodule, based on the spatial direction corresponding angle distribution map, calculates the angular consistency between the CT edge direction and the DTI fiber direction through a multimodal feature fusion algorithm, constructs a three-dimensional orientation field structure map, reorganizes the fusion direction of each spatial location and transforms it into a tensor structure, and obtains a three-dimensional fusion feature map.

4. The neurosurgical visualization and navigation system based on multimodal fusion according to claim 1, characterized in that, The image correction module includes: The image acquisition submodule acquires ultrasound image frame data of the temporal lobe region during surgery, calls the three-dimensional coordinates in the three-dimensional fusion feature map to locate the region within the ultrasound image, extracts pixels of the temporal lobe region edge direction based on the grayscale edge distribution of the image, and performs linear interpolation calculation on each group of directions to generate the temporal lobe edge direction. The structural comparison submodule obtains the direction of the bony edge of the temporal bone based on the direction of the temporal lobe edge, calculates the consistency of the two types of direction values ​​within the same spatial range through a structural similarity algorithm, records the similarity interval corresponding to the position, and obtains the direction similarity interval. The orientation correction submodule performs orientation adjustment operations on the orientation data located in the three-dimensional fusion feature map point by point based on the orientation similarity interval. It performs orientation replacement and records the positions where there is a deviation between the temporal bone orientation and the temporal lobe orientation to obtain the image orientation correction result.

5. The neurosurgical visualization and navigation system based on multimodal fusion according to claim 1, characterized in that, The deformation correction module includes: The centerline extraction submodule, based on the image orientation correction results, locates the strong echo regions of blood vessel segments in the CT image, performs morphological filtering on the regions, collects blood vessel boundary pixels and arranges the point positions in sequence, extracts the centerline structure based on spatial connectivity, and obtains the blood vessel centerline length sequence. The edge tracking submodule selects a set of edge points in a continuous time frame in the intraoperative brain ultrasound image based on the blood vessel centerline length sequence, compares the spatial positions of edge points in adjacent frames and calculates the displacement direction and velocity, and filters out abnormal points with positional changes through a stability judgment threshold to obtain a stable edge displacement range. The trajectory correction submodule calls the stable edge displacement interval and the white matter fiber direction to sequentially perform reverse point projection operation on the spatial coordinates in the blood vessel centerline length sequence, and at the same time performs linear interpolation translation correction within the original coordinate distribution interval to generate a deformation correction image.

6. The neurosurgical visualization and navigation system based on multimodal fusion according to claim 5, characterized in that, The stability judgment threshold is determined by calculating the rate of change of the spatial position of each edge point between adjacent frames, taking the average value of all rate values ​​as the basic reference, and combining it with a range of three times the standard deviation to construct the tolerance range of the edge point movement.

7. The neurosurgical visualization and navigation system based on multimodal fusion according to claim 1, characterized in that, The layer annotation module includes: The orientation extraction submodule, based on the deformation correction image, identifies the fiber orientation distribution area in the DTI image, filters points whose orientation value changes within the stability threshold, locates the blood vessel centerline in the CT image, calculates and maps the angle change between the orientation point and the blood vessel line segment, and generates the orientation relationship angle. The layer construction submodule constructs a fiber orientation layer based on the directional relationship angle and a structural positioning layer based on the spatial path of the blood vessel centerline. Each layer is labeled according to its spatial sequence number in the 3D image, and the structural type and orientation sequence of the region in each layer are marked to obtain the 3D layer order. The layer fusion submodule calls the three-dimensional layer sequence, merges the direction layer and structure layer according to their spatial correspondence, sorts all layers in spatial position and adjusts the coordinate pointing consistency, assigns a unique identification number to each layer and integrates them to establish a three-dimensional layer set.

8. The neurosurgical visualization and navigation system based on multimodal fusion according to claim 7, characterized in that, The stability threshold is set based on the variance and average value of the orientation values ​​of the entire DTI image. First, the average change value of all orientation points is calculated, and then the standard deviation is calculated. The threshold is the average change value plus one standard deviation, which is used to screen out regions with stable orientation and remove points with drastic orientation fluctuations.

9. The neurosurgical visualization and navigation system based on multimodal fusion according to claim 1, characterized in that, The system also includes: The navigation interaction module, based on the three-dimensional layer set, obtains the spatial position of surgical instruments during the operation, performs three-dimensional distance judgment between the instrument tip and the risk point, calls the dynamic collision detection algorithm to set a dynamic threshold to trigger the red layer to flash, adjusts the layer precision according to the interaction delay, and outputs the navigation view interaction data results. The interactive data results of the navigation view include the instrument risk distance value, collision trigger status value, and layer rendering accuracy parameters.

10. The neurosurgical visualization and navigation system based on multimodal fusion according to claim 9, characterized in that, The navigation interaction module includes: The location acquisition submodule, based on the three-dimensional layer set, collects the spatial coordinate information of the end of the surgical instrument during the operation, combines it with the risk areas marked in the layer, calculates the Euclidean distance between the end of the instrument and all risk points in each frame, and filters the effective point set with the best distance in the current frame to generate the instrument risk distance. The collision judgment submodule, based on the device risk distance, calls the dynamic collision detection algorithm to calculate the product of the device movement speed and the layer refresh frequency as a dynamic safety threshold, compares the minimum distance value in each time slice with the threshold, marks the trigger state, and obtains the collision trigger state. The layer response submodule, based on the collision trigger state, controls the red area at the corresponding spatial position in the 3D layer set to perform periodic flashing, compares the layer frame rate data with the interaction response delay record value, adjusts the rendering precision of the area in the layer, and establishes the navigation view interaction data result.

Citation Information

Cited By

  • Ultrasonic fusion multi-modal image guided surgical navigation system of microsurgical robot

    CN121177016A