A needle application automatic navigation method based on multi-modal real-time registration
Patent Information
- Application Number
- CN202511607835.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2045-11-05
AI Technical Summary
误差积累和延迟校正导致针具位置映射出现瞬时偏差,影响施针精度和安全性;
[0021] (1) It can significantly improve the real-time registration accuracy and robustness of the automatic navigation method for needle application in highly dynamic and complex environments, realize the rapid identification and immediate compensation of dynamic errors, break through the problem of correction lag caused by response delay in traditional registration error correction, effectively reduce the impact of sudden events such as rapid tissue deformation and patient micro-movement on navigation accuracy, and ensure the real-time reliability of registration results in key operation stages.
Smart Images

Figure CN121694872B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image fusion and surgical navigation technology, and in particular to an automatic navigation method for needle administration based on multimodal real-time registration. Background Technology
[0002] Currently, intelligent medical image registration and automated navigation systems for minimally invasive surgery are widely used in clinical diagnosis and treatment. Especially in scenarios with significant tissue deformation and patient positional variations, higher demands are placed on the accuracy of real-time image registration and needle navigation. Mainstream technical solutions typically rely on deep learning-driven multimodal registration networks to spatially align preoperative high-resolution images (such as CT and MRI) with intraoperative real-time images (such as ultrasound and endoscopy), achieving precise registration of anatomical structures through non-rigid deformation field modeling. These systems generally employ a posterior error feedback mechanism, detecting post-hoc errors in the registration results and adjusting deformation parameters accordingly to correct the registration. This feedback mechanism performs well in static or low-dynamic environments and has become the mainstream method for intelligent registration and navigation in the industry.
[0003] However, existing real-time medical image registration and navigation systems face the following significant technical shortcomings when applied to minimally invasive interventions, living organ manipulation, and precise needle placement:
[0004] (1) When a patient experiences violent breathing, changes position, or instantaneous contact with the surgical instrument, the tissue structure undergoes rapid deformation. Traditional error identification mechanisms often rely on the accumulation of several frames of results before error detection and correction, resulting in a significant lag in the entire correction response chain. Error accumulation and delayed correction lead to instantaneous deviations in the needle position mapping, affecting the accuracy and safety of needle application.
[0005] (2) Registration networks often use global models or coarse pyramid feature sampling, which makes it difficult to achieve high-resolution incremental response to dynamic mismatch regions; once local high-precision correction is needed, full-frame calculation is usually repeated, which significantly increases the computational load and inference latency.
[0006] (3) Existing technologies have a weak mechanism for distinguishing types of dynamic changes in organizations. They cannot quickly and adaptively switch compensation strategies based on real-time confidence changes for periodic disturbances (such as breathing) and sudden non-steady-state events (such as tool touch), resulting in unstable registration effects in dynamic environments.
[0007] (4) The error correction mechanism is triggered late, and adjustments are often made only after the registration mismatch has significantly affected the navigation path or surgical operation, resulting in a significant serial bottleneck of "error accumulation to detection response". Summary of the Invention
[0008] In order to solve the above-mentioned technical problems, the present invention provides an automatic navigation method for needle application based on multimodal real-time registration.
[0009] The technical solution of this invention is implemented as follows: an automatic navigation method for needle application based on multimodal real-time registration, comprising:
[0010] S1: Acquire high-resolution preoperative images and real-time intraoperative images, which are used as the reference modality and floating modality inputs for the master registration network, respectively;
[0011] S2: Normalize and register the preoperative and intraoperative images, including isotropic resampling, intensity standardization, and coarse registration alignment based on anatomical structure.
[0012] S3: Input the preprocessed preoperative and intraoperative images into the encoder of the deep learning master registration network to extract multi-scale feature maps for subsequent semantic consistency evaluation and deformation field estimation.
[0013] S4: Construct a lightweight two-branch semantic consistency evaluation sub-network, which receives the preoperative and intraoperative feature maps output by the main registration network encoder, respectively, and generates a spatially distributed dynamic semantic matching confidence heatmap by cross-level local mutual information measurement and key point alignment scoring.
[0014] S5: Input the dynamic semantic matching confidence heatmap into the deformation field optimization module as a feedforward signal to drive the fine-grained correction path selection and weight allocation of the local deformation field;
[0015] S6: Determine whether the confidence value of each spatial location in the confidence heatmap is lower than the preset threshold. If it is lower than the threshold, activate the high-resolution local registration path of the corresponding area and call the feature pyramid layer with a higher sampling rate to perform incremental deformation parameter re-estimation.
[0016] S7: Based on time sliding window, classify and identify confidence fluctuation patterns between consecutive frames, distinguish between periodic disturbances and non-steady deformation types, and adaptively switch different compensation strategies according to the identification results;
[0017] S8: For periodic disturbances, a smoothing filter is used to suppress the impact of high-frequency noise; for non-steady-state deformations, an emergency local re-registration process is triggered to achieve a low-latency response to dynamic changes in the tissue.
[0018] S9: Combine the trend of confidence gradient change to determine the direction of deformation propagation, dynamically lock the set of deformation control points that need to be updated first, and avoid the delay caused by global recalculation;
[0019] S10: The optimized local deformation field is fused with the global deformation field output by the master registration network to generate the final real-time registration result, which is used for coordinate mapping and path updating of needle application navigation.
[0020] The present invention provides an automatic navigation method for needle application based on multimodal real-time registration, which has the following beneficial effects:
[0021] (1) It can significantly improve the real-time registration accuracy and robustness of the automatic navigation method for needle application in highly dynamic and complex environments, realize the rapid identification and immediate compensation of dynamic errors, break through the problem of correction lag caused by response delay in traditional registration error correction, effectively reduce the impact of sudden events such as rapid tissue deformation and patient micro-movement on navigation accuracy, and ensure the real-time reliability of registration results in key operation stages.
[0022] (2) It significantly improves the feedforward perception and adaptive adjustment capabilities for local mismatch risks, transforming the error correction process from a reactive response to an active intervention as soon as potential risks emerge, shortening the time window from error occurrence to correction execution, and optimizing the timeliness and continuity of the registration link;
[0023] (3) For different types of disturbances, adaptive switching compensation strategies are used, and algorithms and computing resources are invested in a targeted manner to achieve dual protection of steady-state suppression of periodic disturbances and local fine re-registration of sudden deformations, thereby expanding the scope of application for heterogeneous deformation scenarios;
[0024] (4) By integrating multi-resolution features, spatial confidence and temporal pattern analysis, the local resolution and global consistency control of the registration of key anatomical regions are improved, effectively supporting the real-time update of high-precision navigation paths. Attached Figure Description
[0025] Figure 1 This is a flowchart of an automatic navigation method for needle application based on multimodal real-time registration according to the present invention. Detailed Implementation
[0026] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0027] The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the invention. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0028] like Figure 1 As shown, this invention provides an automatic navigation method for needle application based on multimodal real-time registration, specifically including:
[0029] S1: Acquire high-resolution preoperative images and real-time intraoperative images, which are used as the reference modality and floating modality inputs for the master registration network, respectively;
[0030] S2: Normalize and register the preoperative and intraoperative images, including isotropic resampling, intensity standardization, and coarse registration alignment based on anatomical structure.
[0031] S3: Input the preprocessed preoperative and intraoperative images into the encoder of the deep learning master registration network to extract multi-scale feature maps for subsequent semantic consistency evaluation and deformation field estimation.
[0032] S4: Construct a lightweight two-branch semantic consistency evaluation sub-network, which receives the preoperative and intraoperative feature maps output by the main registration network encoder, respectively, and generates a spatially distributed dynamic semantic matching confidence heatmap by cross-level local mutual information measurement and key point alignment scoring.
[0033] S5: Input the dynamic semantic matching confidence heatmap into the deformation field optimization module as a feedforward signal to drive the fine-grained correction path selection and weight allocation of the local deformation field;
[0034] S6: Determine whether the confidence value of each spatial location in the confidence heatmap is lower than the preset threshold. If it is lower than the threshold, activate the high-resolution local registration path of the corresponding area and call the feature pyramid layer with a higher sampling rate to perform incremental deformation parameter re-estimation.
[0035] S7: Based on time sliding window, classify and identify confidence fluctuation patterns between consecutive frames, distinguish between periodic disturbances and non-steady deformation types, and adaptively switch different compensation strategies according to the identification results;
[0036] S8: For periodic disturbances, a smoothing filter is used to suppress the impact of high-frequency noise; for non-steady-state deformations, an emergency local re-registration process is triggered to achieve a low-latency response to dynamic changes in the tissue.
[0037] S9: Combine the trend of confidence gradient change to determine the direction of deformation propagation, dynamically lock the set of deformation control points that need to be updated first, and avoid the delay caused by global recalculation;
[0038] S10: The optimized local deformation field is fused with the global deformation field output by the master registration network to generate the final real-time registration result, which is used for coordinate mapping and path updating of needle application navigation.
[0039] Step S1: Acquire preoperative high-resolution images and intraoperative real-time images, which will be used as the reference modality and floating modality inputs for the master registration network, respectively. Specifically, this includes:
[0040] S1.1: Acquire high-resolution preoperative image data based on medical imaging equipment, wherein the image data includes one or more of CT, MRI or PET modalities, to obtain an accurate three-dimensional representation of the patient's anatomical structure as a reference modal input for the master registration network;
[0041] The input conditions for medical imaging equipment are a high-resolution image acquisition system with CT, MRI or PET modalities, and the ability to output raw digital image data files and synchronous imaging parameter metadata.
[0042] A CT three-dimensional volume data acquisition algorithm (parameters: spiral scan slice thickness 0.5mm, matrix size 512×512) was used to achieve high-resolution volumetric scanning of the patient's entire bony and soft tissue structures.
[0043] Furthermore, by using an MRI multi-sequence acquisition algorithm (parameters: T1 weighting, T2 weighting and proton density weighting, slice thickness 1mm, matrix size 256×256, field of view size 240mm), accurate imaging of soft tissue and vascular structures is achieved, and a multimodal tomographic image dataset is obtained.
[0044] Furthermore, a PET metabolic imaging acquisition algorithm (parameters: 370 MBq of radioactive tracer 18F-FDG, 20 minutes of scanning time, and 128×128 reconstruction matrix) was used to achieve full-area detection of metabolically active regions and generate a metabolic distribution map.
[0045] By using modal fusion processing, the original datasets of CT, MRI and PET are unified in spatial coordinates and standardized in imaging parameters to generate a set of multimodal three-dimensional volumetric images in a single coordinate system.
[0046] A tissue boundary detection algorithm based on gradient magnitude enhancement (parameter: Sobel operator threshold set to 0.15) is used to enhance the boundaries of the fused 3D image set in order to improve the contrast of anatomical structures in subsequent registration.
[0047] Through the above acquisition and fusion processing methods, the multimodal raw images from the previous step are transformed into reference modal three-dimensional structured data, thereby achieving high-precision reference support for the input end of the master registration network.
[0048] For example, in a clinical application, a CT spiral scan mode was selected, with a slice thickness of 0.5 mm and a matrix resolution of 512×512, acquiring full-field images from the patient's head to neck, with an original data size of 1.2 GB. Simultaneously, MRI multi-sequence acquisition was performed, with a T1-weighted sequence matrix of 256×256, a field of view of 240 mm, a repetition time (TR) of 2000 ms, and an echo time (TE) of 90 ms. PET scanning used 18F-FDG as a tracer at a dose of 370 MBq, a scan time of 20 minutes, and a matrix size of 128×128. The three modalities were aligned in the same spatial coordinate system through affine transformation, and a gradient amplitude enhancement algorithm was used to strengthen key tissue boundaries, resulting in a fused three-dimensional volumetric dataset of 2.8 GB. After data consistency checks, this dataset was input into the reference modality channel of the master registration network, verifying its significant improvement in the stability of anatomical structure localization and the accuracy of deformation field estimation in dynamic real-time registration tasks.
[0049] S1.2: Real-time intraoperative images are acquired through an intraoperative image acquisition system. The images include one or more of ultrasound, X-ray, endoscopy, or optical coherence tomography (OCT) modalities to reflect the patient's current anatomical state during the operation and serve as floating modal inputs to the master registration network.
[0050] In the automatic navigation system for needle administration, the input of real-time intraoperative images needs to be acquired with high fidelity through the intraoperative image acquisition system to ensure the processing accuracy of the floating mode in the master registration network.
[0051] A multimodal medical image acquisition method (parameters: multi-probe ultrasound imaging, digital flat panel X-ray imaging, high-definition endoscopic imaging, high-resolution optical coherence tomography) is used to achieve full-domain acquisition of the patient's current anatomical state and output the raw image data stream.
[0052] Furthermore, through the image synchronization control method (parameters: the clock signal embedded in the acquisition system and the trigger signal of the surgical console), the acquisition time is accurately calibrated, and intraoperative image timestamp data is generated to ensure the time consistency assessment between preoperative and intraoperative images;
[0053] Furthermore, through a device-specific geometric calibration algorithm (parameters: probe spatial orientation data, focal length parameters, imaging lens distortion coefficients), geometric distortion correction of intraoperative images is achieved, and a spatially normalized image matrix after distortion correction is output.
[0054] Furthermore, through real-time noise suppression and enhancement algorithms (parameters: adaptive filter kernel size, histogram equalization intensity coefficient), the signal-to-noise ratio of intraoperative images is optimized and structural details are enhanced, and signal-enhanced image data for registration is obtained;
[0055] Furthermore, by using modal identification and metadata encoding methods, the processed intraoperative images are attached with modal type labels (ultrasound, X-ray, endoscopy, or OCT) and spatial positioning metadata to form a structured floating modal image input set, which serves as the floating modal input interface data for the master registration network.
[0056] Through the above multi-stage algorithm processing, the original intraoperative acquisition results are transformed into spatially normalized and signal-enhanced floating modal image data, realizing a high-precision computable expression of the patient's real-time anatomical state.
[0057] For example, in a laparoscopic minimally invasive surgery scenario, the intraoperative image acquisition system is configured in a trimodal joint acquisition mode, including a 10MHz ultrasound probe, a 1920×1080 resolution endoscopic camera, and an OCT module with a center wavelength of 1300nm and a sampling depth of 3mm. All devices achieve frame synchronization through a unified trigger controller with an internal clock accuracy of ±0.5ms. Ultrasound image distortion correction uses a pinhole lens model, with lens distortion parameters k1 of -0.23 and k2 of 0.14, resulting in a corrected image matrix spatial resolution of 0.5mm / pixel. Endoscopic images undergo adaptive median filtering (window size 5×5) combined with histogram equalization (enhancement coefficient 1.2) for noise suppression and contrast enhancement, significantly improving the visualization of vascular boundaries. OCT data is spatially resampled using linear interpolation, adjusting the slice interval to 5μm to achieve spatial specification uniformity for the trimodal data. The final output set contains three modal synchronization frames, each with modal labels and surgical instrument pose data. This achieves high-precision representation of floating modalities and improves the robustness of subsequent registration processing during the input stage of the main registration network.
[0058] S1.3: Modal consistency correction processing is performed on the preoperative high-resolution images and intraoperative real-time images respectively, including spatial resolution alignment, image slice thickness unification and imaging coordinate system normalization, in order to eliminate the geometric scale differences between modes and generate a modally unified image input set.
[0059] For the preoperative high-resolution images and intraoperative real-time images obtained by steps S1.1 and S1.2, a spatial resolution alignment algorithm (parameter: target resolution is set according to the voxel size of the high-resolution images) is used to unify the pixel spacing of different modal images in the anisotropic direction, and generate a unified resolution image matrix adapted to the subsequent depth registration network.
[0060] Furthermore, by using a slice thickness normalization method (parameter: the target thickness value is based on the slice thickness of the preoperative image), the slice thickness of the intraoperative image is adjusted, and three-dimensional image volume data with consistent slice thickness is obtained, so as to reduce the geometric mismatch between modalities in longitudinal sampling.
[0061] Furthermore, an imaging coordinate system normalization algorithm is adopted (parameters: origin relocation based on anatomical landmarks, coordinate axis direction according to DICOM standard) to achieve the unification of different modal images at the origin and axial direction of the coordinate system, and to generate a standardized coordinate mapping matrix;
[0062] Furthermore, based on the output of the coordinate mapping matrix, affine transformation reprojection processing is performed to achieve pixel geometric consistency of the image under a unified spatial reference frame, and to generate an image dataset after modal consistency correction.
[0063] Modal consistency correction processing transforms the original preoperative high-resolution images and intraoperative real-time images from the previous step into an image input set with consistent spatial resolution, slice thickness, and coordinate system, thereby achieving geometric scale unification between different modalities and providing highly consistent spatial data conditions for subsequent temporal registration and depth feature extraction.
[0064] For example, in an acupuncture procedure, an automatic navigation system for needles acquires preoperative MRI images with a voxel resolution of 0.6mm × 0.6mm × 1.2mm and intraoperative ultrasound images with a spatial resolution of 0.8mm × 0.8mm × 5mm. A spatial resolution alignment algorithm is used to adjust the ultrasound images to 0.6mm × 0.6mm × 5mm, followed by slice thickness normalization to adjust the thickness to 1.2mm, resulting in three-dimensional volumetric data consistent with the MRI in slice thickness. An imaging coordinate system normalization algorithm is used to reposition the MRI origin to a human reference point (frontal-nasal point) and reposition the ultrasound image coordinate system to the same point, ensuring consistency in the z-axis direction, generating a unified coordinate mapping matrix. After affine transformation and reprojection, a modal consistency-corrected image set is output, with a unified resolution of 0.6mm × 0.6mm × 1.2mm, a consistent coordinate system origin, and standardized axial directions. In intraoperative data from 30 patients, the system verified that after adopting this modal consistency correction process, the feature matching stability of the registration network front end was significantly improved, the low confidence region was significantly reduced, and the number of dynamic error correction activations was greatly improved, significantly increasing the real-time registration accuracy of the system in high dynamic scenarios.
[0065] S1.4: Based on the image acquisition timestamp and surgical operation synchronization signal, perform time registration and alignment of preoperative and intraoperative images to ensure that the reference modality and floating modality images are synchronized in the time dimension, forming a temporally consistent image pair;
[0066] S1.5: The spatially and temporally aligned preoperative high-resolution images and intraoperative real-time images are input into the reference modality input channel and floating modality input channel of the master registration network, respectively, to generate initial image input pairs for subsequent feature extraction and semantic consistency evaluation.
[0067] Step S2: Normalization and registration preprocessing are performed on the preoperative and intraoperative images, including isotropic resampling, intensity normalization, and coarse registration alignment based on anatomical structure. Specifically, this includes:
[0068] S2.1: Perform isotropic resampling on preoperative high-resolution images and intraoperative real-time images respectively to unify spatial resolution and eliminate the influence of slice thickness differences on subsequent registration, thereby obtaining standardized image data in isotropic voxel space.
[0069] S2.2: Based on the image data after isotropic resampling, the histogram matching algorithm is used to perform intensity standardization processing on the preoperative and intraoperative images to eliminate the intensity inconsistency caused by differences in equipment imaging parameters and changes in tissue contrast, and obtain the image output after intensity normalization.
[0070] S2.3: Use a rigid registration algorithm based on anatomical landmarks to perform coarse registration and alignment operations on intensity-normalized preoperative and intraoperative images to establish an initial spatial mapping relationship and obtain a preliminary alignment global transformation matrix;
[0071] S2.4: Based on the global transformation matrix of the initial alignment, the preoperative images are mapped to the spatial coordinate system of the intraoperative images to generate a preoperatively aligned preoperative-intraoperative image pairing group, which serves as the initial input for subsequent fine registration;
[0072] S2.5: Perform image boundary cropping and background suppression on the preoperative-intraoperative image pairing group after initial alignment to remove interference from irrelevant areas and enhance the contrast of the region of interest, generating a preprocessed image dataset for input to the deep learning master registration network.
[0073] For the preoperative-intraoperative image pairing group after initial alignment, a boundary clipping algorithm based on segmentation mask (parameters: mask generation model = U-Net architecture, threshold setting based on Otsu automatic threshold segmentation of tissue grayscale distribution) is used to achieve spatial localization and removal of non-interest areas;
[0074] Furthermore, by using morphological opening and closing operations (parameters: structuring element = ellipse, size = 5×5 pixels), the clipping boundary is smoothed, and a clean and burr-free region of interest mask is obtained.
[0075] Furthermore, a background suppression algorithm (parameters: local contrast enhancement coefficient = 1.5, histogram cropping ratio = 0.02) is used to achieve intensity attenuation and noise suppression in the area outside the mask, and to generate image data after background suppression.
[0076] Furthermore, by using an adaptive histogram equalization method (parameters: block size = 8×8 pixels, cropping constraint coefficient = 2.0), local contrast enhancement is achieved within the region of interest, and a contrast-optimized image matrix is generated.
[0077] By normalizing the image data (parameter: intensity normalization range = 0 to 1), the enhanced image results from the previous step are transformed into numerical input data with uniform dimensions, enabling the deep learning master registration network to directly receive preprocessed image datasets.
[0078] For example, in the intraoperative registration scenario of needle placement navigation, a U-Net segmentation model was used to generate a region of interest mask with an accuracy of 256×256 pixels for the initially aligned preoperative MRI and intraoperative ultrasound image pairing group. An Otsu threshold segmenter was set to automatically segment at gray level 120, obtaining a binary mask that accurately covers the target organ. Morphological opening and closing operations used elliptical structuring elements with a radius of 3 pixels to smooth the mask boundaries, reducing edge pixel errors to within 0.5 pixels. During background suppression, a local contrast enhancement coefficient of 1.8 was set, attenuating the background intensity outside the mask to 20% of its original value, significantly reducing noise levels. Adaptive histogram equalization was performed with a block size of 16×16 pixels and a clipping constraint coefficient of 2.5, significantly improving the visibility of vascular textures within the organ. The processed image matrix is normalized to the intensity range of [0,1] and directly input into the encoder channel of the main registration network. The feature extraction output of the network on this dataset shows a reduction in low-confidence regions, which greatly improves the stability and accuracy of subsequent registration.
[0079] Step S3: The preprocessed preoperative and intraoperative images are input into the encoder of the deep learning master registration network to extract multi-scale feature maps for subsequent semantic consistency evaluation and deformation field estimation. Specifically, this includes:
[0080] S3.1: Perform multi-channel encoding on the normalized preoperative high-resolution images and intraoperative real-time images to construct a unified feature space representation;
[0081] S3.2: Construct a deep encoder structure based on a convolutional neural network, and perform multi-level feature extraction on the preoperative and intraoperative images to obtain a high-dimensional semantic feature map that is abstracted layer by layer;
[0082] For preoperative high-resolution images and intraoperative real-time image data after multi-channel encoding, a multi-level encoder structure constructed by a deep convolutional neural network (parameter settings include convolutional kernel size of 3×3, stride of 1, and same filling mode) is used to capture local texture features and abstract spatial structure patterns of the images.
[0083] Furthermore, by applying batch normalization (with parameters including a momentum coefficient of 0.9 and an ε value of 1e-5) after each convolutional layer, the feature distribution is stabilized, and a normalized feature mapping matrix is obtained, providing numerical scale constraints to prevent gradient vanishing or gradient explosion.
[0084] Furthermore, by introducing the nonlinear activation function ReLU (Rectified Linear Unit), the nonlinear representation capability of the encoded feature space is enhanced, and a composite feature response map containing high-frequency structure and low-frequency trend is generated;
[0085] Furthermore, multi-layer stacked convolution and max pooling operations (with parameters including a pooling window size of 2×2 and a stride of 2) are employed to achieve layer-by-layer downsampling of feature maps and expansion of the receptive field, thereby obtaining a high-dimensional feature tensor containing multi-scale information, providing an input basis for cross-level feature fusion.
[0086] Furthermore, by introducing a global average pooling mechanism into the high-level feature abstraction, a global semantic feature vector is generated to provide a holistic anatomical context reference in subsequent semantic consistency evaluation;
[0087] Through the above multi-layer convolution, normalization, non-linear activation and pooling chain processing, the multi-channel image input of the previous step is transformed into a high-dimensional feature map containing multi-scale and cross-modal structural semantics, realizing deep semantic encoding of preoperative and intraoperative images.
[0088] For example, in the acupuncture automatic navigation scenario, by configuring the first layer of the convolutional encoder with 64 convolutional kernels, a kernel size of 3×3, and a stride of 1, the low-level feature map extracted when processing normalized MRI reference images and real-time ultrasound images has a resolution of 128×128×64. After introducing batch normalization (momentum 0.9, ε=1e-5) and ReLU activation in the second layer, the feature resolution remains unchanged and the numerical distribution tends to stabilize. The third layer uses 128 3×3 convolutional kernels and max pooling (window 2×2, stride 2) to reduce the feature resolution to 64×64×128, and the receptive field range is significantly expanded. In the fourth layer, convolution and pooling are stacked to obtain a high-level semantic feature map of 32×32×256. The top layer obtains a global semantic vector of 1×1×256 through global average pooling, which is used for overall reliability judgment in the semantic consistency evaluation sub-network. With this parameter configuration, the encoder can significantly improve the discriminability of cross-modal features while ensuring computational efficiency. The multi-scale feature map can effectively lock the dynamic mismatch region in the subsequent deformation field estimation, thereby significantly improving the response speed and accuracy of real-time registration.
[0089] S3.3: Introduce feature pyramid modules in each level of the encoder to perform scale-invariant feature enhancement on the local anatomical structures of the preoperative and intraoperative images to improve the robustness of multi-scale registration.
[0090] S3.4: The encoded features of preoperative and intraoperative images are initially semantically aligned through a cross-modal feature alignment mechanism to generate a preliminary matching feature mapping relationship, which serves as the basis for subsequent dynamic semantic confidence assessment.
[0091] For the multi-scale feature maps of preoperative and intraoperative images after multi-channel encoding and feature pyramid enhancement, a cross-modal feature embedding alignment algorithm (parameters: channel attention weight matrix, feature similarity threshold (set to 0.85)) is used to establish the preliminary position and semantic correspondence of different modal features in a unified embedding space.
[0092] Furthermore, a local correlation mapping algorithm (parameters: window convolution kernel size k=5×5, stride s=2) is used to calculate the cross-modal correlation between the two sets of feature maps in the local neighborhood and obtain the matching mapping matrix of the high correlation region, which is used to guide the alignment accuracy.
[0093] Furthermore, a feature registration method based on maximizing mutual information is adopted (parameters: mutual information iteration calculation steps n=10, convergence threshold ε=0.001) to achieve semantic alignment of global distributed features between preoperative and intraoperative feature maps, and to generate a preliminary matching relationship matrix for the entire scene as a basic alignment index.
[0094] The initial matching relationship matrix is optimized by using the matching residual minimization algorithm (parameter: mean square error threshold δ=0.01), which suppresses abnormal matching points and strengthens the weight of highly stable corresponding points. The optimized mapping relationship is input into the subsequent dynamic semantic confidence evaluation sub-network to achieve a stable and consistent foundation for cross-modal features.
[0095] By using a cross-modal feature alignment mechanism, the multi-scale feature representation of the previous step is transformed into a matching mapping matrix containing anatomical consistency information, thereby achieving preliminary synchronization of multimodal data at the semantic level.
[0096] For example, in the dynamic registration process of an acupuncture navigation system, preoperative MRI images were isotropically resampled to 1 mm³ voxels, and intraoperative ultrasound images were resampled to the same resolution; both were processed by an encoder to construct a 256-channel multi-scale feature map. In the cross-modal embedding alignment algorithm, the channel attention weights were normalized to between 0 and 1 using Softmax, and the feature similarity threshold was set to 0.85. The local correlation mapping window size was 5×5 with a step size of 2, realizing cross-modal correlation search in the neighborhood of each pixel. In the mutual information alignment stage, the mutual information calculation iteration steps were 10, and the convergence threshold was set to 0.001, resulting in a preliminary global matching matrix with a significantly improved average mutual information value. Using a matching residual minimization algorithm, the mean square error was controlled within 0.01. The final output matching matrix was verified to reduce the alignment error of anatomical structures near key acupuncture points, significantly improving registration stability and laying a reliable foundation for subsequent confidence map generation.
[0097] S3.5: Output multi-scale feature maps to subsequent modules. The multi-scale feature maps include low-level texture information and high-level structural semantic information, which are used to support the collaborative work of the semantic consistency evaluation sub-network and the deformation field estimation module.
[0098] Step S4: Construct a lightweight dual-branch semantic consistency evaluation sub-network, which receives preoperative and intraoperative feature maps output by the main registration network encoder, respectively, and generates a spatially distributed dynamic semantic matching confidence heatmap through cross-level local mutual information measurement and keypoint alignment scoring. Specifically, this includes:
[0099] S4.1: Based on the preoperative multi-scale feature maps and intraoperative multi-scale feature maps output by the master registration network encoder, a dual-branch feature input channel is constructed, which is respectively connected to the preoperative branch and intraoperative branch of the semantic consistency evaluation sub-network to process different modal image features in parallel.
[0100] Based on the preoperative and intraoperative multi-scale feature maps output by the master registration network encoder, a dual-branch input channel construction method (parameters: preoperative feature branch, intraoperative feature branch, multi-scale feature hierarchy index) is adopted to achieve parallel input and isolated processing of image features of different modalities.
[0101] Furthermore, through the feature channel mapping algorithm (parameters: kernel size 3×3, stride 1, normalization method is batch normalization BN), the channel consistency adjustment of the preoperative branch and the intraoperative feature branch at their respective input stages is realized, and a structure-aligned multi-scale feature tensor is obtained.
[0102] Furthermore, a feature scale synchronization mechanism (parameters: feature pyramid level synchronization, interpolation method is bilinear interpolation) is adopted to achieve spatial resolution matching of corresponding level features of preoperative branches and intraoperative branches, and to generate a multi-scale feature alignment set required for registration calculation.
[0103] Furthermore, through a cross-branch data caching strategy (parameters: cache size 256MB, indexed circular queue), temporary storage and inter-frame synchronization control of features of different modal branches are achieved, and an input matrix for subsequent cross-level local mutual information calculation is generated;
[0104] By constructing parallel input channels and performing feature alignment processing as described above, the encoder output from the previous step is effectively mapped to the preoperative and intraoperative branches of the semantic consistency evaluation subnetwork, thereby achieving structural robustness and processing path independence of multimodal features.
[0105] For example, in an automated acupuncture navigation application, preoperative high-resolution MRI images are encoded to obtain 5-layer multi-scale features, and intraoperative real-time ultrasound images are also encoded to obtain 5-layer multi-scale features. The feature channel mapping algorithm unifies the number of channels in both sets of features to 64 dimensions, and the batch normalization parameter is set to ε=0.001. The feature scale synchronization mechanism interpolates the MRI's 3rd layer features from the original 128×128 pixels to 64×64 pixels, consistent with the ultrasound's 3rd layer features, with interpolation weights linearly distributed according to spatial location. The cross-branch data caching strategy simultaneously retains the 1st to 5th layers of MRI and ultrasound features in a 256MB cache using a circular queue index, ensuring complete timestamp matching. For the 4th layer feature pairs, the cache hit rate reaches over 95%, ensuring the temporal synchronization and spatial consistency of the mutual information calculation input. Finally, the two branch input channels stably output the multi-scale feature alignment set required for registration calculation, significantly improving the computational efficiency and stability of the local mutual information matrix in the subsequent S4.2 step.
[0106] S4.2: Perform cross-level local mutual information metric calculation on the feature maps of each level input from the preoperative and intraoperative branches to evaluate the semantic similarity between different anatomical regions in preoperative and intraoperative images, and generate a multi-level local mutual information matrix as a preliminary basis for semantic matching metric.
[0107] S4.3: Based on a predefined anatomical keypoint detector, keypoints are extracted and aligned from preoperative and intraoperative feature maps. Geometric transformation residuals and semantic consistency scores between keypoint pairs are calculated to quantify the matching stability of key structures during the registration process.
[0108] Based on the preoperative and intraoperative multi-scale feature maps output by the master registration network encoder, a predefined anatomical key point detector (parameters: custom key point type set, detection confidence threshold (set to 0.85), scale range) is used to achieve the function of extracting the location of anatomical key structures for two different modal features.
[0109] Furthermore, through scale normalization and coordinate system mapping algorithms (parameters: normalization factor, coordinate system transformation matrix), a unified representation of the coordinates of key points before and during surgery is achieved, and a set of key point coordinates that can be compared across modalities is obtained;
[0110] Furthermore, a keypoint matching algorithm based on nearest neighbor search (parameters: Euclidean distance as the distance metric and a matching threshold of 2.5 mm) is used to pair keypoints before and during the operation, and a set of keypoint pairs is generated for subsequent geometric transformation residual calculation.
[0111] Furthermore, the keypoint pair set is optimized and registered using a rigid transformation fitting algorithm (parameters: least squares fitting (limited to 50 times), iterative convergence criterion (termination when the convergence residual amplitude is less than 0.01 mm)), and the geometric transformation residual values between each keypoint pair are calculated. ,in Indexed for keypoint pairs;
[0112] Furthermore, a semantic consistency scoring function (parameters: mutual information weight α, structural similarity weight β) is used to calculate the cross-modal feature vector similarity of keypoint pairs, as shown in the following formula:
[0113]
[0114] in For the first Semantic consistency score of key point pairs For mutual information, For structural similarity, and These represent the feature vectors of the corresponding key points before and during the operation, respectively;
[0115] By combining geometric transformation residuals and semantic consistency scores, the key point matching results of the previous step are transformed into quantified structural matching stability data, thereby enabling reliability assessment of key structures during real-time registration.
[0116] For example, in a minimally invasive interventional surgery scenario, the preoperative MRI images, after being encoded, yielded multi-scale feature maps with resolutions of 128×128, 64×64, and 32×32. Intraoperative ultrasound images were also encoded to obtain multi-scale feature maps. The anatomical keypoint detector was configured to simultaneously detect three types of keypoints: bony prominences, vascular bifurcation points, and organ edges. The detection confidence threshold was set to 0.85, and the scale range covered all feature map levels. The coordinates of the detected keypoints were then transformed using a coordinate system transformation matrix. Mapping was performed to achieve cross-modal positional unification. During the matching phase, a nearest neighbor search with a Euclidean distance threshold of 2.5 mm was used, resulting in 48 keypoint pairings. In the rigid transformation fitting process, the number of least squares iterations was limited to 50, and the convergence criterion was a residual decrease of less than 0.01 mm. The average value of the geometric residuals was calculated. mm, maximum value mm. The semantic consistency score uses weighted coefficients of α=0.6 and β=0.4 to calculate the mutual information and structural similarity of each key point pair. The final average score is significantly improved, indicating that the key structure matching stability of this method is significantly improved in this scenario. As the initial input for S4.4 fusion processing, it can effectively improve the recognition accuracy of low confidence regions.
[0117] S4.4: The local mutual information measurement results and key point alignment scores are weighted and fused to generate a fused semantic consistency score map, which serves as the initial input feature for dynamic semantic matching confidence.
[0118] S4.5: Based on the fused semantic consistency score map, a spatial attention mechanism is performed through a lightweight convolutional neural network to enhance the processing, highlight the confidence response of high error risk areas, and output a spatially distributed dynamic semantic matching confidence heatmap;
[0119] Based on the fused semantic consistency scoring map, a lightweight convolutional neural network (parameters: kernel size 3×3, stride 1, number of channels adapted to the feature dimension of the fused scoring map) is used to model the local correlation of the initial spatial features;
[0120] Furthermore, by employing a multi-scale spatial attention mechanism (parameters: window scale includes dual-scale configurations of 5×5 and 9×9, and attention weight calculation uses Softmax normalization), the response to high error risk regions is enhanced, and a feature activation map modulated by attention weights is obtained.
[0121] Furthermore, a channel-recalibrated SE (Squeeze-and-Excitation) module (parameters: compression rate 1 / 16, activation function is a combination of ReLU and Sigmoid) is adopted to achieve adaptive adjustment of the weights of the fused scoring map on different semantic channels and generate the channel-weighted spatial feature mapping result;
[0122] Furthermore, by using spatial weight backpropagation constraints (parameter: regularization coefficient λ=0.001, constraint objective is to maximize the gradient magnitude of high-risk regions), the network's ability to discriminate low-confidence regions is improved during training, and a confidence enhancement matrix is generated.
[0123] By using convolutional fusion and spatial upsampling (parameters: upsampling factor of 2, interpolation method of bilinear interpolation), the result of the previous step is transformed into a spatially distributed dynamic semantic matching confidence heatmap, thereby realizing the instantaneous salience of potential mismatch regions during the registration process.
[0124] For example, in an automated navigation scenario for needle administration, for a fused semantic consistency score map with an input dimension of 256×256, the convolutional neural network is set to a 3-layer structure. The first layer has a 3×3 kernel size and 32 channels, the second layer has a 3×3 kernel size and 64 channels, and the third layer has a 1×1 kernel size and 1 channel, used to generate an initial single-channel confidence feature map. In the spatial attention mechanism, dual-scale windows extract the mean and variance of local regions as attention inputs, and the normalized attention weight matrix is calculated using Softmax. In channel recalibration, each channel of the fused scoring feature map is averaged in the global space to obtain a channel description vector. This vector is then passed through a two-layer fully connected network (compression ratio 1 / 16) to output channel weights, which are compressed to the [0,1] interval using the Sigmoid function. Spatial weight backpropagation constraints are implemented by introducing a gradient magnitude regularization term for high-risk regions into the loss function, enhancing the response to mismatch-prone regions during training. Finally, after bilinear interpolation upsampling to the original input resolution of 256×256, the generated dynamic semantic matching confidence heatmap, in the validation set targeting deformation field disturbances caused by simulated patient breathing, significantly improves the detection sensitivity of low-confidence regions and effectively reduces the registration residual values in these regions, achieving rapid feedforward compensation for dynamic errors.
[0125] Step S5: The dynamic semantic matching confidence heatmap is input to the deformation field optimization module as a feedforward signal to drive the fine-grained correction path selection and weight allocation of the local deformation field. Specifically, this includes:
[0126] S5.1: Perform spatial resolution adaptation processing on the dynamic semantic matching confidence heatmap to match the input dimension requirements of the deformation field optimization module, and obtain the adapted spatial confidence distribution feature map.
[0127] S5.2: Based on the adapted spatial confidence distribution feature map, calculate the local confidence gradient at each spatial location to obtain the confidence gradient vector field, which is used to characterize the dynamic propagation trend of potential mismatch regions;
[0128] S5.3: Perform directional clustering analysis on the confidence gradient vector field to identify the spatial propagation direction with the fastest decrease in confidence, and generate a deformation propagation direction indicator map to guide the priority update strategy of local deformation control points;
[0129] S5.4: Based on the deformation propagation direction indicator map and the preset deformation control point priority rules, dynamically lock the local control point set that needs to be updated first, and generate a control point update priority sequence;
[0130] In the processing chain of the deformation field optimization module, the input data are the deformation propagation direction indication diagram generated by S5.3 and the set of deformation control point priority rules preset by the system;
[0131] A directional feature matching algorithm (parameters: deformation propagation direction indicator map, control point spatial coordinate set) is used to establish the correspondence between the propagation direction and the spatial location of the control points, so as to identify the candidate set of control points on high propagation risk paths;
[0132] Furthermore, by using a priority weight calculation method (parameters: historical stability score of control points, local value of semantic matching confidence), the comprehensive priority index of control points is generated, and the weight distribution data of each control point in the candidate set is obtained.
[0133] Furthermore, a threshold-based sorting algorithm (parameters: comprehensive priority index, dynamic threshold value) is adopted to filter the weight distribution data, retain the control point records with priority indices higher than the threshold value, and arrange them in descending order of index value to generate a sorted list of control points.
[0134] Furthermore, a spatial clustering constraint method (parameters: control point list, anatomical structure adjacency matrix) is adopted to perform spatial connectivity analysis on the sorted list, grouping locally adjacent control points with similar priorities to ensure the consistency of local updates and the rationality of the anatomy, and forming a sequence of grouped control points;
[0135] The control point sequence generation module transforms the above grouped sequence into a control point update priority sequence, enabling precise driving of the subsequent local deformation parameter re-estimation process, significantly reducing the computational load in irrelevant regions and improving response speed.
[0136] For example, in an automated needle navigation scenario, assume that three main propagation paths have been identified in the deformation propagation direction indicator map, with a total of 120 control point coordinates K along these paths. Each control point has a local semantic matching confidence value and a historical stability score. The direction feature matching algorithm maps K to obtain a candidate set C of control points located on high-risk paths, ultimately selecting 48 points. Priority weights. The calculation uses a weighted summation formula:
[0137]
[0138] in This is the weighting coefficient for historical stability scoring. To score the stability of the control points, The confidence level weighting coefficient is... Let be the local confidence level. =0.6, =0.4, then the P value of each control point is calculated, and the dynamic threshold is set to 0.75, retaining only control points with P ≥ 0.75. The sorting results form 20 high-priority control point records. After spatial clustering constraint analysis, these 20 points are divided into 5 groups, each group corresponding to a dissected adjacent region. The final generated control point update priority sequence is input into the local deformation parameter estimation module. In the subsequent deformation field fusion, the registration error of the region where these control points are located is significantly reduced, the overall system response speed is improved, and high spatial calibration accuracy is maintained.
[0139] S5.5: Based on the control point update priority sequence, combined with the high sampling rate layer in the multi-resolution feature pyramid, incremental local deformation parameter estimation is performed to generate a local optimized deformation field;
[0140] S5.6: Based on the spatial consistency constraints between the local optimized deformation field and the global deformation field, a weighted fusion strategy is executed to generate a fused high-precision real-time deformation field output for coordinate mapping updates in needle navigation;
[0141] Based on the local optimized deformation field output by S5.5 and the global deformation field of the master registration network as input data, a spatial consistency constraint calculation method (parameters: three-dimensional coordinate system transformation matrix, anatomical structure topological constraint weight coefficients) is adopted to achieve geometric alignment of the two sets of deformation fields under a unified spatial reference.
[0142] Furthermore, by using a weighted fusion strategy algorithm (parameters: local region dynamic semantic matching confidence value, global smoothing weight coefficient), the deformation vector fusion weights at different spatial locations are adaptively allocated, and the fusion weight distribution matrix is obtained.
[0143] Furthermore, a pixel-wise weighted vector overlay method (parameters: fusion weight distribution matrix, local deformation vector field, global deformation vector field) is adopted to construct the fusion vector field and generate high-precision real-time deformation field data;
[0144] Furthermore, a nonlinear smoothing filtering algorithm (parameters: size of the three-dimensional Gaussian kernel, number of iterations) is applied to suppress local discontinuities in the fused vector field, resulting in a physically reasonable and continuous final deformation field output.
[0145] Furthermore, by using a needle navigation coordinate mapping update method (parameters: final deformation field, preoperative image coordinate model, intraoperative image coordinate model), the spatial positioning path of the needle is corrected in real time, and the update instruction data required by the needle navigation module is generated.
[0146] By using weighted fusion and spatial consistency alignment, the results of the local and global deformation fields from the previous step are transformed into a high-precision, low-latency fused deformation field, which significantly improves the registration accuracy and dynamic response performance during needle application.
[0147] For example, in a minimally invasive spinal surgery navigation scenario, the local optimized deformation field resolution is 0.5 mm voxels, the global deformation field resolution is 1 mm voxels, the spatial consistency constraint weight is set to 0.8, and the corresponding anatomical topology constraint model includes the spinal vertebral body and surrounding soft tissues. The fusion weight distribution matrix is calculated using the following formula:
[0148]
[0149] in, Let be the fusion weight at position i. This represents the dynamic semantic matching confidence value for this location. The nonlinear smoothing filtering algorithm uses the Gaussian kernel standard deviation. mm, number of iterations Secondly, it ensures the continuity of the vertebral body edges and soft tissue transition zone in the fusion deformation field without any breaks. In this scenario, when the final deformation field is used for needle path calculation, the vertebral body positioning deviation is reduced by more than 1.0 mm compared to the basic method, the navigation path update delay is reduced to less than 30 ms, and the navigation system can still maintain stable needle application accuracy and path following ability even with slight respiratory displacement of the patient.
[0150] Step S6: Determine whether the confidence value of each spatial location in the confidence heatmap is lower than a preset threshold. If it is lower than the threshold, activate the high-resolution local registration pathway of the corresponding region and call the feature pyramid layer with a higher sampling rate to perform incremental deformation parameter re-estimation. Specifically, this includes:
[0151] S6.1: Spatial discretization is performed on the dynamic semantic matching confidence heatmap, and confidence threshold comparison is performed in units of pixels or voxels to identify potential mismatch areas with confidence scores below the preset threshold.
[0152] S6.2: Based on the identified low-confidence regions, generate a local registration activation mask map and mark the spatial coordinate set that needs to be prioritized for high-resolution registration processing, as the input control signal for the local registration path;
[0153] Using the low-confidence region coordinate set output by S6.1 as input data, a spatial bit mask generation algorithm (parameters: region coordinate set, image spatial resolution) is used to visualize and encode potential mismatch regions, and form a preliminary mask matrix for navigation control of subsequent high-resolution processing paths.
[0154] Furthermore, through the regional connectivity analysis algorithm (parameters: 8-neighborhood or 26-neighborhood connection rule, area threshold), the connected domain segmentation of the initial mask matrix is realized, and a set of connected subdomains that meet the minimum spatial scale requirements is obtained, filtering out isolated or noisy point regions;
[0155] Furthermore, a morphological combination of region dilation and erosion is used (parameters: structuring element shape = sphere, radius = 1~2 pixels / voxel) to achieve smoothing of the boundaries of connected subdomains and restoration of structural integrity, and to generate a boundary-optimized binary mask matrix.
[0156] Furthermore, by using a coordinate space index remapping algorithm (parameter: preoperative-intraoperative unified coordinate system transformation matrix), the boundary-optimized binary mask matrix is converted into a global spatial coordinate set, ensuring that the coordinate set has consistency in the multimodal registration network;
[0157] Furthermore, a regional priority scalar assignment method (parameter: dynamic semantic matching confidence value back mapping coefficient) is adopted to assign a numerical processing priority label to each element in the global spatial coordinate set, thereby generating a local registration activation mask map with priority weights;
[0158] By combining spatial location mask generation, connected component analysis, morphological optimization, coordinate remapping and priority assignment, the low-confidence region identification results from the previous step are transformed into structured and weighted spatial control signals, enabling precise triggering and resource allocation optimization of local registration paths.
[0159] For example, in the real-time registration module of the needle application navigation system, the input data is a set of low-confidence region coordinates obtained after threshold segmentation of the 3D dynamic semantic matching confidence heatmap. This coordinate set contains approximately 1500 voxels with a spatial resolution of 0.5 mm³. The spatial mask generation algorithm constructs a preliminary mask matrix with dimensions of 512×512×256 based on this coordinate set. Region connectivity analysis employs a 26-neighborhood connection rule, with an area threshold set to 100 voxels. Noise regions smaller than this threshold are eliminated, retaining only 7 connected subdomains. Morphological processing selects spherical structural elements with a radius of 2 voxels for dilation and erosion operations to ensure smooth and continuous boundaries between subdomains. Coordinate remapping uses a global transformation matrix established from preoperative to intraoperative images to map the mask matrix to a unified coordinate system, ensuring coordinate consistency across different modalities. The priority assignment method uses a confidence value back-mapping coefficient k=10 to map the original confidence value of each coordinate point to a priority weight, thereby forming a local registration activation mask with the same dimension as the original image. This mask is then used as an input control signal in S6.3, enabling the system to instantly call the high-resolution feature pyramid layer for target region feature re-extraction. Verification shows that in scenes with rapid tissue deformation, the local registration response time is significantly shortened while maintaining stable deformation correction accuracy.
[0160] S6.3: Based on the local registration activation mask, call the feature pyramid level with a higher sampling rate in the main registration network to re-extract local features from the low confidence region to obtain more refined spatial structure information;
[0161] S6.4: Based on the re-extracted high-resolution local features, incremental deformation parameter estimation is performed, and a non-rigid registration algorithm is used to refine the deformation field of the local region to improve the local registration accuracy.
[0162] S6.5: The local deformation parameters estimated by incremental estimation are fused with the global deformation field output by the master registration network to generate the fused local optimized deformation field, which serves as the input data for the subsequent deformation field optimization module.
[0163] The deformation parameters estimated by the non-rigid registration algorithm are applied to the high-resolution local features obtained by re-extraction and used as local input data for fusion processing;
[0164] A spatial consistency-constrained weighted fusion method (constraint parameters: neighborhood structure similarity, local confidence weight) is adopted to achieve spatial alignment between local and global deformation field data.
[0165] Furthermore, by using a weighted vector superposition algorithm (weight allocation based on dynamic semantic matching confidence matrix), pixel-by-pixel fusion of local deformation vectors and global deformation vectors is achieved, resulting in a high-precision fused deformation vector field.
[0166] Furthermore, a nonlinear smoothing filtering method is adopted (filter kernel parameters: three-dimensional Gaussian kernel, standard deviation is adaptively adjusted according to the variance of the fused vector) to suppress the local discontinuities of the fused vector field and generate a locally optimized deformation field that is continuous in three-dimensional space and has a reasonable geometric structure.
[0167] Furthermore, the accuracy of the local optimized deformation field is evaluated by a verification algorithm based on the consistency check of multi-resolution feature pyramids, and a deformation field quality score index is generated.
[0168] Through the above fusion and smoothing process, the local deformation parameters of the previous step are transformed into a locally optimized deformation field with structural continuity and high spatial accuracy, so as to realize a high-confidence data stream input to the subsequent deformation field optimization module.
[0169] For example, in a minimally invasive spinal surgery navigation scenario, the sampling interval of the high-resolution feature pyramid layer was set to 0.25 mm. A local registration activation mask marked low-confidence voxel blocks around the third lumbar vertebra, totaling 32,768 voxel points. The non-rigid registration algorithm employed a B-spline deformation model based on maximizing mutual information, with a control point spacing of 8 mm and a mutual information calculation window of 5×5×5 voxels. In the spatial consistency constraint weighted fusion method, the neighborhood structure similarity threshold was set to 0.85, and the local confidence weight ranged from 0.6 to 1.0. The fusion vector field was calculated using the following formula. :
[0170]
[0171] in For local confidence weights, For local deformation parameter vector field, This represents the global deformation parameter vector field. Nonlinear smoothing filtering employs a three-dimensional Gaussian kernel. The filter standard deviation is adaptively calculated through the variance of the fused vector field. The formula for calculating the standard deviation is as follows:
[0172]
[0173] in The standard deviation of the filter. To fuse the vector magnitude at position i of the vector field, To achieve the mean of the fused vector field magnitudes, This represents the total number of vector field positions. After performing the above processing, the spatial continuity of the locally optimized deformation field in the L3 vertebral region is significantly improved, and the registration error of anatomical key points is reduced to within 0.12 mm, achieving high-precision coordinate mapping updates of the needle during navigation.
[0174] Step S7: Classify and identify confidence fluctuation patterns between consecutive frames based on a time sliding window, distinguish between periodic disturbances and non-steady-state deformation types, and adaptively switch different compensation strategies according to the identification results. Specifically, this includes:
[0175] S7.1: Perform time-sliding window sampling on the dynamic semantic matching confidence heatmap within continuous time frames to construct a time evolution model of the confidence sequence for subsequent fluctuation pattern classification analysis;
[0176] For the fused local optimized deformation field and the corresponding dynamic semantic matching confidence heatmap output by step S6, a time sliding window sampling method (parameters: sliding window length L, step size Δt) is used to achieve serialized capture of the confidence distribution of continuous time frames;
[0177] Furthermore, by using a sliding window traversal strategy (parameter: window start frame index k), the spatial confidence matrix sequence within each time window is stably extracted, and a set of confidence matrices arranged in chronological order is obtained.
[0178] Furthermore, a spatial compression mapping method (parameter: number of principal components p) is adopted to map the two-dimensional / three-dimensional confidence matrix of each frame into a one-dimensional feature vector, thereby eliminating the time modeling computation burden caused by excessive spatial dimension and obtaining a time-series-based confidence feature stream;
[0179] Furthermore, a normalization method (parameters: mean μ and value range) is adopted to standardize the amplitude of the time-series confidence feature stream, ensuring that the confidence values of different frames and different spatial regions are comparable under a unified dimension, and generating a normalized confidence time series;
[0180] Furthermore, through a sequence caching mechanism, the normalized confidence time series is stored in a sliding window buffer in real time, providing a continuous and ordered temporal evolution data structure for subsequent frequency domain conversion and pattern classification;
[0181] By using a time-sliding window sampling algorithm and a normalization process, the dynamic confidence matrix result from the previous step is transformed into a confidence time evolution model that can be used for fluctuation pattern analysis, thereby achieving a structured representation of the registration stability state over a continuous time period.
[0182] For example, in a needle application navigation system, a time-sliding window sampling is performed on the dynamic semantic matching confidence heatmap. The window length L is set to 20 frames, the step size Δt is set to 5 frames, and the starting frame index k of the window is sequentially increased. Each heatmap frame is 128×128 pixels in size. The dimensionality is reduced to a one-dimensional feature vector of p=5 dimensions through two-dimensional principal component analysis, forming a time window sequence matrix of length 20. Amplitude normalization is performed on this matrix, with the mean μ set to 0.5 and the value range [0,1]. The normalization formula is as follows:
[0183]
[0184] in, This is the normalized confidence score. The original confidence level value. For the normalized mean, and These represent the maximum and minimum confidence values within the window, respectively. The normalized time series is input into the sliding window buffer to form a 20×5 ordered matrix, which can provide stable temporal evolution characteristics before the occurrence of periodic disturbances caused by respiration and sudden deformations caused by device contact, providing low-latency and high-accuracy input data for the subsequent frequency domain conversion step in S7.2;
[0185] S7.2: The confidence sequence is frequency domain transformed based on Fast Fourier Transform (FFT) to extract spectral energy distribution features in order to identify whether there are periodic components;
[0186] S7.3: Perform peak detection algorithm on the energy spectrum after frequency domain conversion to identify the main frequency band and its harmonic components, and determine whether there are periodic disturbances caused by breathing or heartbeat;
[0187] S7.4: If a periodic dominant frequency component is detected, an autoregressive model (AR model) is used to model the confidence sequence and predict the confidence change trend of the next time frame in order to identify high-frequency noise interference.
[0188] S7.5: Perform wavelet packet decomposition algorithm on non-periodic confidence fluctuations to extract multi-scale detail coefficients and identify the occurrence time and spatial distribution characteristics of sudden non-steady-state deformation events;
[0189] S7.6: Based on the confidence fluctuation classification results, dynamically select the compensation strategy: for periodic disturbances, activate the smoothing filter to suppress high-frequency noise; for non-steady-state deformations, trigger the emergency local re-registration process.
[0190] Step S8: For periodic disturbances, a smoothing filter is used to suppress the influence of high-frequency noise; for non-steady-state deformations, an emergency local re-registration process is triggered to achieve a low-latency response to dynamic changes in the tissue. Specifically, this includes:
[0191] S8.1: Based on the confidence fluctuation sequence within the time sliding window, calculate the confidence change amplitude and periodicity characteristics between the current frame and historical frames to identify the type of disturbance;
[0192] S8.2: For regions identified as periodic disturbances, a moving average filtering algorithm based on adaptive window length is used to denoise the confidence heatmap in order to suppress the interference of high-frequency noise on deformation estimation.
[0193] S8.3: Based on the filtered confidence heatmap, calculate the confidence variance and gradient magnitude of the local region to assess the perturbation intensity and determine the priority of deformation field correction;
[0194] S8.4: For regions identified as non-steady-state deformation, based on the confidence decrease rate and spatial distribution characteristics, an emergency re-registration process based on the local feature pyramid is triggered, and the high-resolution feature layer is called to re-estimate the local deformation parameters.
[0195] S8.5: In the emergency local re-registration process, a non-rigid registration algorithm based on maximizing mutual information is used to perform fine deformation field correction on the local area to generate high-precision local registration increment information;
[0196] S8.6: The local registration increment information is fused with the global deformation field to generate updated deformation field parameters, which serve as the basis for subsequent needle coordinate mapping and navigation path adjustment.
[0197] Step S9: Determine the deformation propagation direction by combining the confidence gradient change trend, and dynamically lock the set of deformation control points that need to be updated first to avoid delays caused by global recalculation. Specifically, this includes:
[0198] S9.1: Calculate the gradient direction and magnitude of the confidence heatmap for dynamic semantic matching to obtain the local trend features of confidence changes in the spatial domain, providing a basic input for subsequent deformation propagation direction judgment;
[0199] S9.2: Based on gradient direction features, perform deformation propagation direction classification processing, use the direction consistency clustering algorithm to divide the gradient vector field into regions, identify the main direction of deformation propagation dominated by confidence decrease, and thus determine the potential propagation path of tissue deformation;
[0200] S9.3: Based on the identified deformation propagation direction and combined with the anatomical structure topological constraint model, a deformation propagation priority mask map is generated. This mask map marks the key control point regions on the deformation propagation path with high priority, which serve as candidate regions for subsequent deformation field updates.
[0201] S9.4: Perform local deformation sensitivity assessment on the deformation control points in the candidate region, calculate the sensitivity of each control point to the final registration error based on the deformation gradient tensor output by the master registration network, and generate a priority update sorting table for control points;
[0202] For the candidate deformation control point regions marked in the deformation propagation priority mask map, a local deformation sensitivity evaluation algorithm (parameters: deformation gradient tensor, local confidence weight, spatial neighborhood radius) is used to calculate the relative contribution of a single control point to the global registration error in the current registration state.
[0203] Furthermore, the sensitivity value of each control point is calculated using the weighted absolute gradient norm method (with weights based on local semantic matching confidence) through the three-dimensional deformation gradient tensor output by the master registration network, and the initial sensitivity distribution matrix is obtained.
[0204] Furthermore, the normalized impact response analysis method (parameters: time sliding window length, disturbance amplitude factor) is used to dynamically adjust the control point sensitivity matrix and establish the sensitivity curve of the control point under time-series changes, so as to screen out unstable high-sensitivity anomalies caused by short-term fluctuations.
[0205] Furthermore, based on the above sensitivity curve, a sorting and selection algorithm (parameters: sensitivity threshold, priority decrease coefficient) is used to sort the control points in descending order to generate a priority update sorting table for control points, wherein the sorting is based on a weighted comprehensive score of sensitivity value and confidence decrease rate.
[0206] The priority list generated by sensitivity assessment and ranking transforms the candidate region results from the previous step into an executable control point update plan, achieving the expected technical effect of avoiding global recalculation.
[0207] For example, in the real-time registration process of automatic needle application navigation, suppose the candidate region contains 50 deformation control points, and the range of each component of the deformation gradient tensor is... The local confidence weights are between 0.3 and 0.9, ranging from 2.0 to 2.0. For each control point i, the weighted absolute gradient norm is calculated. :
[0208]
[0209] in, The current local semantic matching confidence weight for the control point. , , These are the components of the deformation gradient tensor in three spatial directions. In this embodiment, the time window length is set to 5 frames, and the perturbation amplitude factor is set to 0.2. Instantaneous outliers caused by surgical instrument vibration are removed through shock response analysis. The top 15 control points in the final sort list are entered into the priority update plan, and local non-rigid registration parameter optimization is performed, which significantly improves the registration accuracy of local areas and reduces computational latency, maintaining the stability of the navigation path under high-speed tissue deformation conditions.
[0210] S9.5: Based on the control point priority update sorting table, dynamically lock the set of deformation control points that need to be updated first, and only perform local deformation parameter optimization on high-priority control points, avoiding repeated calculation of the deformation field of the whole scene, thereby significantly reducing the computational load and response latency of deformation field update.
[0211] Step S10: The optimized local deformation field is fused with the global deformation field output by the master registration network to generate the final real-time registration result, which is used for coordinate mapping and path updating of needle application navigation. Specifically, this includes:
[0212] S10.1: Spatial alignment is performed between the optimized local deformation field and the global deformation field output by the master registration network to ensure that the two have geometric consistency in the three-dimensional spatial coordinate system, thereby providing structured input for subsequent fusion operations;
[0213] S10.2: Based on the weighted fusion strategy, pixel-level weighted calculation of deformation vectors is performed on the aligned local deformation field and global deformation field. The weight coefficients are dynamically adjusted according to the semantic matching confidence heatmap of the local region to improve the registration accuracy in high uncertainty regions.
[0214] S10.3: Perform nonlinear smoothing filtering on the weighted fused deformation vector field to suppress the local deformation discontinuities introduced during the fusion process, and obtain a smooth and physically reasonable comprehensive deformation field as the basis data for real-time registration results;
[0215] S10.4: Based on the fused integrated deformation field, perform non-rigid transformation on the preoperative images to generate dynamic reference images that are spatially aligned with the real-time intraoperative images, which are used for anatomical structure mapping and positioning correction in needle application navigation.
[0216] S10.5: The dynamic reference image is compared with the current intraoperative image in real time, and the three-dimensional coordinate mapping relationship of the needle application path is updated based on the comparison result to generate a real-time path update command for the navigation control module, so as to realize closed-loop compensation of dynamic error.
[0217] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
[0218] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and rules of the present invention should be included within the scope of protection of the present invention.
Claims
1. A multi-modal real-time registration based automatic needle navigation method for needle insertion, characterized in that, Includes the following steps: S1: Acquire high-resolution preoperative images and real-time intraoperative images, which are used as the reference modality and floating modality inputs for the master registration network, respectively; S2: Normalize and register the preoperative high-resolution image and the intraoperative real-time image to generate preoperative and intraoperative images after normalization. S3: Input the preprocessed preoperative images and intraoperative images into the encoder of the deep learning master registration network to extract multi-scale feature maps respectively; S4: Construct a lightweight two-branch semantic consistency evaluation sub-network, which receives the preoperative and intraoperative feature maps output by the main registration network encoder, respectively, and generates a spatially distributed dynamic semantic matching confidence heatmap by cross-level local mutual information measurement and key point alignment scoring. S5: Input the dynamic semantic matching confidence heatmap into the deformation field optimization module; S6: Determine whether the confidence value of each spatial location in the confidence heatmap is lower than the preset threshold. If it is lower than the threshold, activate the high-resolution local registration path of the corresponding area and call the feature pyramid layer with a higher sampling rate to perform incremental deformation parameter re-estimation. S7: Based on the time sliding window, classify and identify the confidence fluctuation patterns between consecutive frames of the confidence heatmap, distinguish between periodic disturbances and non-steady-state deformation types, and adaptively switch different compensation strategies according to the identification results. S8: For periodic disturbances, a smoothing filter is used to suppress the impact of high-frequency noise; for non-steady-state deformations, an emergency local re-registration process is triggered. S9: Determine the direction of deformation propagation by combining the trend of confidence gradient change at each spatial location of the confidence heatmap, and dynamically lock the set of deformation control points that need to be updated first. S10: The optimized local deformation field is fused with the global deformation field output by the master registration network to generate the final real-time registration result.
2. The automatic needle navigation method based on multi-modal real-time registration of needle insertion according to claim 1, characterized in that, Step S1 specifically includes: Acquire high-resolution preoperative imaging data using medical imaging equipment; Intraoperative real-time images are acquired using an intraoperative imaging system; Modality consistency correction is performed on the preoperative high-resolution images and the intraoperative real-time images respectively to generate a modally unified image input set; Based on the image acquisition timestamp and the surgical operation synchronization signal, the time registration and alignment of preoperative and intraoperative images are performed to form a time-consistent image pair. The time-consistent image pairs are input to the reference mode input channel and the floating mode input channel of the master registration network to generate the initial image input pair.
3. The automatic needle navigation method based on multi-modal real-time registration of a needle tool according to claim 2, characterized in that, The preoperative high-resolution images include one or more of CT, MRI, or PET modalities, and the intraoperative real-time images include one or more of ultrasound, X-ray, endoscopy, or optical coherence tomography modalities.
4. The automatic navigation method for needle application based on multimodal real-time registration according to claim 1, characterized in that, Step S2 specifically includes: Isotropic resampling was performed on preoperative high-resolution images and intraoperative real-time images to obtain standardized image data in isotropic voxel space. Based on the standardized image data, the histogram matching algorithm is used to perform intensity standardization processing on the preoperative and intraoperative images to obtain the image output after intensity normalization. A rigid registration algorithm based on anatomical landmarks was used to perform coarse registration and alignment operations on intensity-normalized preoperative and intraoperative images to establish an initial spatial mapping relationship and obtain a preliminary alignment global transformation matrix. Based on the preliminary alignment global transformation matrix, the preoperative images are mapped to the spatial coordinate system of the intraoperative images to generate a preliminary aligned preoperative-intraoperative image pairing group. Image boundary cropping and background suppression are performed on the pre-aligned preoperative-intraoperative image pairing group to generate a preprocessed image dataset.
5. The automatic navigation method for needle application based on multimodal real-time registration according to claim 1, characterized in that, Step S3 specifically includes: Normalized preoperative high-resolution images and intraoperative real-time images are subjected to multi-channel encoding to construct a unified feature space representation. A deep encoder structure is constructed based on a convolutional neural network. Multi-level feature extraction is performed on the normalized preoperative high-resolution image and intraoperative real-time image to obtain a high-dimensional semantic feature map that is abstracted layer by layer. Feature pyramid modules are introduced into each level of the encoder to perform scale-invariant feature enhancement on local anatomical structures in preoperative and intraoperative images. A cross-modal feature alignment mechanism is used to perform preliminary semantic alignment of the encoded features of preoperative and intraoperative images, generating preliminary matching feature mapping relationships.
6. The automatic navigation method for needle application based on multimodal real-time registration according to claim 5, characterized in that, The cross-modal feature alignment mechanism includes normalizing the channel attention weight matrix and setting the feature similarity threshold to 0.
85. It also involves maximizing mutual information with an iteration step of 10, a convergence threshold of 0.001, and using a matching residual minimization algorithm to control the mean square error within 0.
01.
7. The automatic navigation method for needle application based on multimodal real-time registration according to claim 1, characterized in that, Step S4 specifically includes: Based on the preoperative multi-scale feature maps and intraoperative multi-scale feature maps output by the master registration network encoder, a dual-branch feature input channel is constructed, which is respectively connected to the preoperative branch and intraoperative branch of the semantic consistency evaluation sub-network. Perform cross-level local mutual information metric calculation on the feature maps of each level input from the preoperative and intraoperative branches to generate a multi-level local mutual information matrix; Based on a predefined anatomical keypoint detector, keypoints are extracted and aligned from preoperative and intraoperative feature maps. Geometric transformation residuals and semantic consistency scores between keypoint pairs are calculated to generate keypoint alignment scores. The multi-level local mutual information matrix and the key point alignment score are weighted and fused to generate a fused semantic consistency score map. Based on the fused semantic consistency score map, a spatial attention mechanism is used to enhance the processing through a lightweight convolutional neural network, outputting a spatially distributed dynamic semantic matching confidence heatmap.
8. The automatic navigation method for needle application based on multimodal real-time registration according to claim 7, characterized in that, In step S4, the anatomical key point detector function supports custom key point type sets, sets the detection confidence threshold to 0.85, uses an Euclidean distance threshold of 2.5 mm for key point pairing, and supports least squares iteration limit of 50 times for rigid transformation fitting, terminating when the convergence residual amplitude is less than 0.01 mm.
9. The automatic navigation method for needle application based on multimodal real-time registration according to claim 1, characterized in that, Step S5 includes: Spatial resolution adaptation processing is performed on the dynamic semantic matching confidence heatmap to match the input dimension requirements of the deformation field optimization module, thereby obtaining the adapted spatial confidence distribution feature map.
Citation Information
Patent Citations
Cross-module non-rigid body registration method and system based on deep learning, and medium
CN115690178A
Multi-mode large model assisted laparoscope soft tissue registration surgical navigation method and system
CN119941813A