Preoperative and intraoperative three-dimensional spinal registration method based on multi-scale reliable occlusion fusion
By constructing a three-dimensional spinal registration method based on multi-scale reliable occlusion fusion, the problem of insufficient handling of occlusion and blurred regions in existing technologies is solved, achieving high-precision and robust intraoperative three-dimensional registration and meeting the high-precision positioning requirements of spinal surgery navigation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING UNIV OF TECH
- Filing Date
- 2026-05-12
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies fail to effectively handle occlusion, blurring, and distortion areas in 3D registration during spinal surgery, leading to registration drift, segment confusion, and positioning deviation. Furthermore, they lack multi-scale cross-modal feature interaction and reliable occlusion suppression mechanisms, making it difficult to meet the high precision and robustness requirements of complex surgical scenarios.
A three-dimensional spine registration method based on multi-scale reliable occlusion fusion is constructed. Through a dual-view two-dimensional and three-dimensional image reconstruction and registration integrated network framework, a dual-branch encoder is used to extract multi-scale features, a multi-layer cross-modal fusion module is used to realize feature interaction, and a reliable occlusion guidance map is generated through an occlusion/reliable estimation module to suppress low-quality regional information. Spatial alignment is achieved by combining a deformation field prediction layer.
It improves the robustness and accuracy of intraoperative three-dimensional spinal registration, and can maintain high-precision spatial positioning information in complex scenarios, meeting the stable and accurate requirements of spinal surgery navigation.
Smart Images

Figure CN122492774A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing and surgical navigation technology, specifically to a preoperative and intraoperative three-dimensional spinal registration method based on multi-scale reliable occlusion fusion. Background Technology
[0002] Preoperative and intraoperative three-dimensional registration is a core technology of spinal surgical navigation systems. By establishing a spatial correspondence between preoperative CT and intraoperative images, it provides precise positioning for procedures such as pedicle screw placement, spinal correction, and decompression fusion. Intraoperative dual-view X-ray imaging has become the mainstream image source for real-time intraoperative positioning due to its convenient acquisition and low radiation dose. However, due to changes in body position, limited projection angles, overlapping bony structures, interference from metal instruments, and differences in exposure conditions, dual-view images often exhibit uneven clarity, local occlusion, missing edges, and structural distortion. Directly using two-dimensional and three-dimensional registration is susceptible to the compression of projection information. While reconstructing intraoperative three-dimensional data before performing three-dimensional registration can reduce information loss, it is difficult to handle artifacts, structural gaps, and unreliable areas in the reconstruction results, leading to registration drift, segment confusion, and positioning deviations. This makes it difficult to meet the high-precision and robust registration requirements in complex intraoperative scenarios.
[0003] In recent years, 3D reconstruction and registration of the spine based on biplane X-rays has become a research hotspot, with numerous academic achievements driving technological iteration. The TransVert network uses orthogonal X-ray images to infer the 3D pose of the spine, improving the accuracy of vertebral morphology reconstruction; the ProVLNet network introduces a projection geometry-aware mechanism to achieve 3D vertebral localization and keypoint regression in biplane X-rays; the EAR network employs an edge-aware strategy to enhance the reconstruction effect of sparse vertebral views; and the RadGS-Reg framework combines 3D radiometric Gaussian reconstruction with 3D registration, first reconstructing biplane X-rays into a 3D representation before aligning it with CT. While these methods improve reconstruction quality and registration efficiency, they all assume equal contribution of features from each region and do not explicitly model the credibility and occlusion degree of occluded, blurred, and distorted regions. This fails to suppress error propagation caused by low-quality views and reconstruction defects, resulting in a significant decrease in registration stability when biplane quality is uneven or local structural degradation occurs.
[0004] Chinese patent CN113538533A discloses a spinal registration method, device, equipment, and computer storage medium. It achieves MR and CT image registration through feature extraction and transformation units, improving cross-modal alignment accuracy. However, it only addresses static preoperative multimodal data and does not cover intraoperative X-ray and 3D reconstruction processes, failing to handle intraoperative occlusion and dynamic quality differences. Chinese patent CN116019554B proposes an accelerated spinal surgical navigation registration method, improving registration speed through optimized feature sampling and matching strategies. However, it relies on traditional manual features and rigid transformations, without employing deep learning reconstruction and adaptive fusion. It lacks robustness against intraoperative reconstruction artifacts and local structural defects, and does not construct reliable weights and occlusion suppression mechanisms. Chinese patent CN117017487B discloses a spinal registration device and method, achieving intraoperative registration through pose parameter optimization. It still uses a traditional iterative matching approach, lacking bi-branch coding, multi-scale fusion, and reliable occlusion guidance, making it difficult to handle complex intraoperative degradation scenarios.
[0005] Existing literature and patent methods generally suffer from four major defects: they do not explicitly model occlusion and reliable regions in dual-view reconstruction; they adopt equal-weight or fixed fusion strategies, resulting in direct propagation of errors in unreliable regions; they lack multi-scale cross-modal feature interaction mechanisms, and the structural correspondence between modalities is unclear; and they do not integrate reliable perception, occlusion suppression, and adaptive fusion in the reconstruction and registration process. Summary of the Invention
[0006] To address the aforementioned technical issues, this application discloses a preoperative and intraoperative three-dimensional spinal registration method based on multi-scale reliable occlusion fusion, specifically including:
[0007] S1. Construct a network framework integrating dual-view 2D and 3D image reconstruction and registration;
[0008] S2. Input the intraoperative anteroposterior X-ray image and lateral X-ray image into the three-dimensional reconstruction network to generate intraoperative reconstructed three-dimensional data;
[0009] S3. The preoperative CT body data is used as a fixed image and the intraoperative reconstructed 3D data is used as a moving image. The data is input into a 3D registration network based on reliable occlusion. Multi-scale features are extracted through a dual-branch encoder, and the multi-layer cross-modal fusion module MCF is used to realize the interaction of scale-by-scale features.
[0010] S4. The encoded and fused features are decoded step by step. Through the multi-scale aggregation module MSA, the occlusion / credible estimation module OB / OC, and the credible guided fusion module COGF, enhanced features for deformation field prediction are obtained.
[0011] S5. After channel stitching of the enhanced features and moving image features, input them into the deformation field prediction layer to obtain the three-dimensional deformation field. Then, the spatial transformation module STN performs spatial transformation on the intraoperative reconstructed three-dimensional data to obtain the registered three-dimensional data aligned with the preoperative CT body data.
[0012] Preferably, the integrated network framework constructed in S1 specifically includes a dual-view X-ray input terminal, a three-dimensional reconstruction terminal, a preoperative CT input terminal, and a three-dimensional registration terminal based on reliable occlusion;
[0013] The dual-view X-ray input terminal is used to input the anteroposterior X-ray image AP and lateral X-ray image LA acquired during the operation of the patient. The three-dimensional reconstruction terminal is used to restore the three-dimensional structure during the operation based on the dual-view two-dimensional image. The preoperative CT input terminal is used to input the patient's preoperative CT volume data. The three-dimensional registration terminal is used to perform spatial alignment between the preoperative CT volume data and the intraoperative reconstructed three-dimensional data, and finally outputs the registered three-dimensional result.
[0014] Preferably, the three-dimensional registration network in S3 uses a dual-branch encoder with the same structure but different parameters to extract multi-scale features from the fixed image and the moving image respectively.
[0015] The fixed image branch coding features and the moving image branch coding features of the i-th layer are respectively represented as follows:
[0016]
[0017]
[0018] in, Encode features for fixed image branches in layer i. For the encoding operation of the i-th layer of the fixed image branch, Preoperative CT scan data; Encode features for the i-th layer moving image branch. For the encoding operation of the i-th layer of the moving image branch, To reconstruct three-dimensional data during the operation.
[0019] Preferably, in step S3, after the feature extraction of each layer of coding is completed, a multi-layer cross-modal fusion module (MCF) is used to perform scale-wise interaction and fusion of the same-scale dual-branch features to obtain the fused features of the i-th layer:
[0020]
[0021] in, For the cross-modal fusion feature of the i-th layer, For the i-th layer multi-modal fusion operation, For the i-th layer, the image encoding features are fixed. Encode features for the i-th layer of the moving image;
[0022] The multi-layer cross-modal fusion module MCF first maps the two features to a unified feature space, and then performs channel splicing and convolution fusion to enhance the structural correspondence between modalities.
[0023] Preferably, in S4, the decoder outputs multi-scale decoding features. Bottleneck characteristics ,Will and Input the Multiscale Aggregation Module (MSA);
[0024] The module first performs resampling and channel alignment on features at different scales, then performs channel concatenation and convolution fusion to output multi-scale aggregated features:
[0025]
[0026] in, This is a multi-scale aggregation feature. For multi-scale aggregation operations, Features for mid-level and deep decoding This is a bottleneck characteristic.
[0027] Preferably, in S4, multi-scale aggregation features are used. The input occlusion / reliability estimation module OB / OC generates an occlusion weight map through two parallel branches. With Trusted Weights And construct a reliable occlusion guidance graph:
[0028]
[0029] in, For a reliable occlusion guide map, For a reliable weighted graph, To occlude the weight map, This is an element-wise multiplication operation;
[0030] By using guided maps, high-confidence, low-occlusion areas are highlighted, while invalid information from occluded and artifact areas is suppressed.
[0031] Preferably, in S4, a reliable occlusion guidance map is included. Multi-scale aggregation features Element-wise multiplication followed by convolution mapping yields occlusion / credibility-aware features:
[0032]
[0033] in, For occlusion / reliable perception features, For convolution mapping operations, This is a multi-scale aggregation feature. For a reliable occlusion guide map, This is an element-wise multiplication operation.
[0034] Preferably, in S4, the signal-guided fusion module COGF receives obstruction / trusted sensing features. shallow decoding features First, map the two feature paths to a unified fusion space to obtain... Then based on adaptive guiding weights Perform complementary weighted fusion:
[0035]
[0036]
[0037]
[0038] right Perform residual augmentation to obtain the final enhanced features used for deformation prediction. The formula is: ,in, Features after mapping For adaptive guiding weights, As a temporary fusion feature, This represents the residual enhancement mapping function.
[0039] Preferably, the enhanced features will be implemented in S5. With moving image features Channel splicing is performed, and the input deformation field prediction layer is used to obtain the three-dimensional deformation field through three-dimensional convolutional regression:
[0040]
[0041] in, For a three-dimensional deformation field, For the deformation field prediction function, This is for channel splicing operations. To ultimately enhance features, To reconstruct three-dimensional data during the operation.
[0042] According to the method of claim 9, the three-dimensional deformation field in step S5 is characterized in that... With intraoperative reconstruction of three-dimensional data Input the spatial transformation module STN to complete spatial alignment and obtain the registered 3D data:
[0043]
[0044] The network is trained using a joint loss function of image similarity and deformation smoothness.
[0045]
[0046] in, For the registered 3D data, For spatial transformation operations, For the total loss, For similarity loss, For deformation smoothing loss, This is the loss weighting coefficient.
[0047] Compared with the prior art, the technical solution of this application has the following technical effects:
[0048] This invention achieves end-to-end processing from dual-view 2D images to 3D structures by constructing an integrated reconstruction and registration network framework, providing complete workflow support for the spatial alignment of preoperative and intraoperative spinal surgery data. This framework deeply integrates intraoperative dual-view image reconstruction and 3D registration processes, ensuring consistent information transmission from 2D projection to 3D spatial transformation. It avoids information loss and error accumulation issues in step-by-step processing, providing a stable foundation for subsequent deformation field prediction and registration result output.
[0049] This invention employs a dual-branch encoder to independently extract multi-scale features from preoperative and intraoperative 3D data, and achieves interaction and fusion of features at the same scale through a multi-layer cross-modal fusion module. This fusion process first maps features from different modalities to a unified feature space, and then enhances the structural correspondence between modalities through channel stitching and convolutional fusion. This simultaneously preserves key structural information from both preoperative CT and intraoperative reconstructed 3D data, providing a stable and reliable cross-modal feature foundation for subsequent decoding stages and improving the completeness and consistency of feature representation.
[0050] This invention integrates decoding features from different levels through a multi-scale aggregation module, achieving effective utilization of multi-scale information. Simultaneously, it introduces an occlusion and reliability estimation module to generate a reliable occlusion guidance map, highlighting high-reliability regions and suppressing invalid information in low-quality regions. Building upon this, a reliable guidance fusion module complementarily fuses occlusion-based reliable perception features with shallow decoding features, generating enhanced features for deformation field prediction. This strengthens adaptability to reconstruction artifacts, structural defects, and local occlusion, improving the stability and reliability of deformation field prediction.
[0051] This invention combines enhanced features with moving image features, achieving adaptive alignment of intraoperative 3D data through a deformation field prediction layer and a spatial transformation module. Simultaneously, a joint loss function is used to train and optimize the network, ensuring the smoothness of the deformation field and the consistency of the registration results. This scheme effectively improves the robustness of preoperative and intraoperative 3D spinal registration through reliable perception and occlusion suppression mechanisms, providing stable and accurate spatial positioning information for spinal surgery navigation, and meeting the high-precision registration requirements in complex intraoperative scenarios.
[0052] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the preferred embodiments of this application are described in detail below with reference to the accompanying drawings.
[0053] The above and other objects, advantages and features of this application will become more apparent to those skilled in the art from the following detailed description of specific embodiments in conjunction with the accompanying drawings. Attached Figure Description
[0054] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In all drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.
[0055] Based on the description of the figures and their corresponding technical content in the document, the titles of the figures are as follows:
[0056] Figure 1 A schematic diagram of the overall process of preoperative and intraoperative 3D spinal registration method based on multi-scale reliable occlusion fusion;
[0057] Figure 2 A schematic diagram of the integrated network framework structure for dual-view 2D image to 3D reconstruction and 3D registration;
[0058] Figure 3 , three Schematic diagram of the overall structure of the dual-branch encoder, multi-layer cross-modal fusion and decoder of the registration network;
[0059] Figure 4 A schematic diagram of the internal operation process and key component structure of the Multi-Scale Aggregation Module (MSA);
[0060] Figure 5A schematic diagram of the structure of the occlusion confidence estimation module OB / OC and the confidence guidance fusion module COGF. Detailed Implementation
[0061] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. In the following description, specific details such as specific configurations and components are provided merely to help fully understand the embodiments of this application. Therefore, those skilled in the art should understand that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. In addition, for clarity and brevity, descriptions of known functions and structures are omitted in the embodiments.
[0062] It should be understood that the phrase "an embodiment" or "this embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "an embodiment" or "this embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.
[0063] Furthermore, reference numerals and / or letters may be repeated in different examples within this application. Such repetition is for the purpose of simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or settings discussed.
[0064] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, and A and B exist simultaneously. The term " / and" describes another type of relationship between related objects, indicating that two relationships can exist. For example, A / and B can mean: A exists alone, and A and B exist alone. In addition, the character " / " in this article generally indicates that the related objects before and after it are in an "or" relationship.
[0065] In this article, the term "at least one" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, "at least one of A and B" can mean: A exists alone, A and B exist simultaneously, or B exists alone.
[0066] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion.
[0067] Example 1
[0068] This embodiment mainly describes a preoperative and intraoperative three-dimensional spinal registration method based on multi-scale reliable occlusion fusion, such as... Figure 1 As shown, it specifically includes:
[0069] S1. Construct a network framework integrating dual-view 2D and 3D image reconstruction and registration;
[0070] S2. Input the intraoperative anteroposterior X-ray image and lateral X-ray image into the three-dimensional reconstruction network to generate intraoperative reconstructed three-dimensional data;
[0071] S3. The preoperative CT body data is used as a fixed image and the intraoperative reconstructed 3D data is used as a moving image. The data is input into a 3D registration network based on reliable occlusion. Multi-scale features are extracted through a dual-branch encoder, and the multi-layer cross-modal fusion module MCF is used to realize the interaction of scale-by-scale features.
[0072] S4. The encoded and fused features are decoded step by step. Through the multi-scale aggregation module MSA, the occlusion / credible estimation module OB / OC, and the credible guided fusion module COGF, enhanced features for deformation field prediction are obtained.
[0073] S5. After channel stitching of the enhanced features and moving image features, input them into the deformation field prediction layer to obtain the three-dimensional deformation field. Then, the spatial transformation module STN performs spatial transformation on the intraoperative reconstructed three-dimensional data to obtain the registered three-dimensional data aligned with the preoperative CT body data.
[0074] Furthermore, such as Figure 2 As shown, the patient's intraoperative anteroposterior X-ray image was acquired. Lateral X-ray images The two-dimensional X-ray images from both perspectives were input into a three-dimensional reconstruction network to recover the intraoperative reconstructed three-dimensional data. ;Preoperative CT scan data With intraoperative reconstruction of three-dimensional data A common input-based 3D registration network with reliable occlusion predicts 3D deformation fields through multi-scale feature extraction, reliable occlusion modeling, and guided fusion. It also uses a spatial transformation module to align intraoperative reconstructed 3D data to preoperative CT space, outputting registered 3D data. Preoperative CT scan data provides stable three-dimensional anatomical information of the patient before surgery, while intraoperative reconstructed three-dimensional data reflects the spatial distribution of bony structures in the patient's current intraoperative state. Spatial alignment of these two data provides a reliable spatial reference for segment identification, instrument positioning, and operative path planning in spinal surgery navigation.
[0075] In this invention, a dual-view 3D reconstruction network is used to reconstruct intraoperative 3D structures based on anteroposterior and lateral X-ray images, and is expressed as follows:
[0076]
[0077] in, Represents a 3D reconstruction network. This represents the intraoperative reconstructed three-dimensional data. Since anteroposterior and lateral X-ray images reflect bony structural information from different directions, the two viewpoints are complementary in terms of vertebral body contours, edge orientation, and visibility of overlapping areas. Utilizing the combined information from both views to reconstruct intraoperative three-dimensional data helps improve the structural integrity and spatial consistency of moving images during subsequent registration.
[0078] like Figure 3 As shown, the 3D registration network based on reliable occlusion includes a dual-branch encoder, a multi-layer cross-modal fusion module MCF, a decoder, a multi-scale aggregation module MSA, an occlusion / reliable estimation module OB / OC, a reliable guided fusion module COGF, a deformation field prediction layer, and a spatial transformation module STN. Preoperative CT volumetric data. and intraoperative reconstruction of three-dimensional data The data are input separately into a dual-branch encoder for multi-scale feature extraction. The two encoders are structurally identical but do not share parameters to accommodate the distribution differences between preoperative CT volume data and intraoperative reconstructed 3D data. Let the first... The encoding features of the fixed image branch and the moving image branch are respectively and Then we have:
[0079]
[0080]
[0081] in, and These represent the fixed image branch and the moving image branch respectively in the first... Layer encoding operations. Through the above encoding process, layer-level feature representations of preoperative CT volume data and intraoperative reconstructed 3D data at different scales can be extracted respectively.
[0082] After extracting encoded features at each scale, the preoperative CT branch features and intraoperative reconstructed 3D branch features are further interacted and fused on a scale-by-scale using the multi-layer cross-modal fusion module MCF, thereby enhancing the structural correspondence between different modalities and obtaining fused features. :
[0083]
[0084] After obtaining the bi-branch coding features, the multi-layer cross-modal fusion module (MCF) is further used to interact and fuse features at the same scale to enhance the structural correspondence between different modalities. The MCF module first performs channel mapping on the bi-branch features, which is expressed as follows:
[0085]
[0086] in, and These represent the fixed image branch features and the moving image branch features respectively in the first... The channel mapping function of the layer. This step maps the dual-branch features to a unified feature space, facilitating subsequent fusion.
[0087] After completing the channel mapping, the two feature channels are concatenated, and the result is obtained by using a convolutional fusion function. Layer fusion features, which are expressed as follows:
[0088]
[0089] in, This indicates a channel splicing operation. This represents the convolutional fusion function. Through this fusion process, structural information from both preoperative CT volume data and intraoperative reconstructed 3D data can be preserved at each scale, thus providing a more stable cross-modal feature foundation for subsequent decoding and deformation field prediction.
[0090] The fused features are fed into the decoder for step-by-step recovery to obtain multi-scale decoded features. , , , and bottleneck characteristics Shallower features retain more local edge and detail information, while deeper features retain more global semantic and structural context information. For example... Figure 4 As shown, the Multi-Scale Aggregation Module (MSA) receives... , , and As input, aggregated features are obtained through scale alignment, channel concatenation, and convolutional fusion. :
[0091]
[0092] To more clearly describe the processing flow of the MSA module, features at different scales are first resampled and channel aligned, which is expressed as follows:
[0093]
[0094] in, , , and This represents a mapping function that resamples and aligns features at different scales. This step maps multi-scale features with different resolutions and channel dimensions to a unified representation space.
[0095] After alignment, the above features are then spliced through channels and further aggregated and merged, as expressed as:
[0096]
[0097] in, This represents the aggregation and fusion function. Through this multi-scale aggregation process, deep semantic information and mid-level structural information can be effectively integrated, thereby enhancing the ability to discriminate complex anatomical regions in subsequent occlusion / trust modeling stages.
[0098] like Figure 5 As shown, the occlusion / reliability estimation module OB / OC receives aggregated features. As input, occlusion information and reliability information are estimated through two parallel branches. Specifically, the occlusion weight map... With Trusted Weights It can be represented as:
[0099]
[0100] in, and These represent the occlusion estimation branch and the confidence estimation branch, respectively. This step allows us to obtain the occlusion degree and confidence degree of each region in the current aggregated features. After obtaining the two types of weight maps, a confidence occlusion guidance map is further constructed to highlight high-confidence and low-occlusion regions, expressed as:
[0101]
[0102] in, This represents element-wise multiplication. This formula is used to highlight high-confidence, low-occlusion areas while suppressing unreliable information in occluded, overlapping, or artifact-affected areas. Its purpose is to provide a larger response to high-confidence areas while suppressing heavily occluded areas, thus generating guiding information that is more suitable for subsequent fusion.
[0103] After obtaining the reliable occlusion guidance map, it is then applied to the aggregated features. The occlusion / credibility perception features are obtained through convolutional mapping, and are expressed as follows:
[0104]
[0105] in, This represents the convolution mapping function. The mechanism of this module lies in its explicit construction... This enables the network to respond more effectively to regions with high credibility and low occlusion at the feature level, while suppressing regions with occlusion, structural overlap, and artifact interference, thereby improving the robustness of subsequent fusion and registration.
[0106] Furthermore, the Trusted Guidance Fusion Module (COGF) integrates occlusion / trusted awareness features. shallow decoding features Mapped to respectively and This is to facilitate subsequent complementary integration, and it can be expressed as:
[0107]
[0108] in, and This represents the feature mapping function. Through this mapping process, the two feature streams can be integrated into a unified fusion space. After obtaining the two mapped features, guiding weights are further generated based on their concatenation result. Its expression is:
[0109]
[0110] in, This represents the guiding weight generation function. This represents the activation function. The weights are used to control the relative contributions of the two features during fusion. After obtaining the guiding weights, the activation function is then used... Weighting of highly reliable semantic features, while utilizing Complementary weighting of shallow detail features is expressed as follows:
[0111]
[0112]
[0113] The above two equations yield the enhanced results of the credible semantic branch and the shallow detail branch, respectively. Then, the two weighted features are added together to obtain a temporary fusion feature, expressed as:
[0114]
[0115] This step achieves complementary fusion of the two feature streams, enabling the network to enhance the representation of reliable regions while preserving local anatomical details. After obtaining the temporary fused features, the final enhanced features are obtained through residual enhancement. Its expression is:
[0116]
[0117] in, This represents the residual enhancement mapping function. This step further enhances the expressive power of the fused features, thereby improving registration stability in complex occlusion scenarios. After obtaining the final enhanced features... Then, it is concatenated with the moving image features through channels and input into the deformation field prediction layer. The three-dimensional deformation field is obtained through three-dimensional convolutional regression. :
[0118]
[0119] in, This represents the deformation field prediction function. This indicates a channel stitching operation. Subsequently, the intraoperative reconstructed 3D data will be... With three-dimensional deformation field The data are input into the STN spatial transformation module to transform the intraoperative reconstructed 3D data into the preoperative CT space, resulting in the final registered 3D data. :
[0120]
[0121] in, This indicates the registration result after spatial alignment with the preoperative CT body data.
[0122] To improve the spatial alignment accuracy between preoperative CT volumetric data and intraoperative reconstructed 3D data, the registration network can be trained using a joint loss function based on image similarity and deformation field smoothness. During training, the loss function consists of two main components: image similarity loss and deformation smoothness loss.
[0123] Image similarity loss is used to measure the similarity of the registered 3D data. Compared with preoperative CT data The degree of similarity between them. Considering the gray-level consistency characteristics in medical image registration, this invention uses the normalized cross-correlation coefficient as the similarity measure function, the expression of which is:
[0124]
[0125] in, Indicates intraoperative reconstruction of three-dimensional data In the deformation field The transformation result obtained under this action encourages the network to learn a deformation field that maximizes the local similarity between two images. The deformation smoothness loss is used to constrain the spatial continuity and physical plausibility of the predicted deformation field. This loss achieves the smoothness constraint on deformation by minimizing the first-order gradient of the displacement vector in three-dimensional space, and is defined as follows:
[0126]
[0127] in, Represents the positions of all voxels within the image domain. Indicates the deformation field at the voxel position The gradient at that point.
[0128] In summary, the total loss function is expressed as a weighted sum of image similarity loss and deformation smoothness loss, that is:
[0129]
[0130] in, and These are the corresponding weighting coefficients, used to balance the relationship between image matching accuracy and deformation field smoothing constraints.
[0131] This embodiment cascades dual-view 2D X-ray 3D reconstruction with preoperative and intraoperative 3D registration. During the registration process, it introduces multi-layer cross-modal fusion, multi-scale aggregation, occlusion / reliable estimation, and reliable guided fusion mechanisms. This effectively reduces the impact of occlusion, inconsistent viewpoint quality, missing local structures, and reconstruction artifacts on registration accuracy. Compared with existing technologies, this invention maintains good 3D structural recovery capability and spatial registration accuracy even with inconsistent dual-view image quality, local occlusion, or structural degradation during surgery, thus improving the spatial positioning reliability and clinical application value in spinal navigation scenarios.
[0132] The above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any changes, modifications, substitutions, integrations, and parameter changes made to these embodiments within the spirit and principles of the present invention, without departing from the principles and spirit of the present invention, through conventional substitutions or to achieve the same function, fall within the scope of protection of the present invention.
Claims
1. A preoperative and intraoperative three-dimensional spinal registration method based on multi-scale reliable occlusion fusion, characterized in that, Includes the following steps: S1. Construct a network framework integrating dual-view 2D and 3D image reconstruction and registration; S2. Input the intraoperative anteroposterior X-ray image and lateral X-ray image into the three-dimensional reconstruction network to generate intraoperative reconstructed three-dimensional data; S3. The preoperative CT body data is used as a fixed image and the intraoperative reconstructed 3D data is used as a moving image. The data is input into a 3D registration network based on reliable occlusion. Multi-scale features are extracted through a dual-branch encoder, and the multi-layer cross-modal fusion module MCF is used to realize the interaction of scale-by-scale features. S4. The encoded and fused features are decoded step by step. Through the multi-scale aggregation module MSA, the occlusion / credible estimation module OB / OC, and the credible guided fusion module COGF, enhanced features for deformation field prediction are obtained. S5. After channel stitching of the enhanced features and moving image features, input them into the deformation field prediction layer to obtain the three-dimensional deformation field. Then, the spatial transformation module STN performs spatial transformation on the intraoperative reconstructed three-dimensional data to obtain the registered three-dimensional data aligned with the preoperative CT body data.
2. The preoperative and intraoperative three-dimensional spinal registration method based on multi-scale reliable occlusion fusion according to claim 1, characterized in that, The integrated network framework constructed in S1 specifically includes a dual-view X-ray input terminal, a three-dimensional reconstruction terminal, a preoperative CT input terminal, and a three-dimensional registration terminal based on reliable occlusion. The dual-view X-ray input terminal is used to input the anteroposterior X-ray image AP and lateral X-ray image LA acquired during the operation of the patient. The three-dimensional reconstruction terminal is used to restore the three-dimensional structure during the operation based on the dual-view two-dimensional image. The preoperative CT input terminal is used to input the patient's preoperative CT volume data. The three-dimensional registration terminal is used to perform spatial alignment between the preoperative CT volume data and the intraoperative reconstructed three-dimensional data, and finally outputs the registered three-dimensional result.
3. The preoperative and intraoperative three-dimensional spinal registration method based on multi-scale reliable occlusion fusion according to claim 2, characterized in that, The three-dimensional registration network in S3 uses a dual-branch encoder with the same structure but different parameters to extract multi-scale features from both the fixed and moving images. The fixed image branch coding features and the moving image branch coding features of the i-th layer are respectively represented as follows: in, Encode features for fixed image branches in layer i. For the encoding operation of the i-th layer of the fixed image branch, Preoperative CT scan data; Encode features for the i-th layer moving image branch. For the encoding operation of the i-th layer of the moving image branch, To reconstruct three-dimensional data during the operation.
4. The preoperative and intraoperative three-dimensional spinal registration method based on multi-scale reliable occlusion fusion according to claim 3, characterized in that, In step S3, after the feature extraction of each layer is completed, the multi-layer cross-modal fusion module MCF is used to perform scale-wise interaction and fusion of the same-scale dual-branch features to obtain the fused features of the i-th layer: in, For the cross-modal fusion feature of the i-th layer, For the i-th layer multi-modal fusion operation, For the i-th layer, the image encoding features are fixed. Encode features for the i-th layer of the moving image; The multi-layer cross-modal fusion module MCF first maps the two features to a unified feature space, and then performs channel splicing and convolution fusion to enhance the structural correspondence between modalities.
5. The preoperative and intraoperative three-dimensional spinal registration method based on multi-scale reliable occlusion fusion according to claim 4, characterized in that, The decoder in S4 outputs multi-scale decoding features. Bottleneck characteristics ,Will and Input the Multiscale Aggregation Module (MSA); The module first performs resampling and channel alignment on features at different scales, then performs channel concatenation and convolution fusion to output multi-scale aggregated features: in, This is a multi-scale aggregation feature. For multi-scale aggregation operations, Features for mid-level and deep decoding This is a bottleneck characteristic.
6. The preoperative and intraoperative three-dimensional spinal registration method based on multi-scale reliable occlusion fusion according to claim 5, characterized in that, The S4 section will incorporate multi-scale aggregation features. The input occlusion / reliability estimation module OB / OC generates an occlusion weight map through two parallel branches. With Trusted Weights And construct a reliable occlusion guidance graph: in, For a reliable occlusion guide map, For a reliable weighted graph, To occlude the weight map, This is an element-wise multiplication operation; By using guided maps, high-confidence, low-occlusion areas are highlighted, while invalid information from occluded and artifact areas is suppressed.
7. The preoperative and intraoperative three-dimensional spinal registration method based on multi-scale reliable occlusion fusion according to claim 6, characterized in that, The S4 section will include a reliable occlusion guidance map. Multi-scale aggregation features Element-wise multiplication followed by convolution mapping yields occlusion / credibility-aware features: in, For occlusion / reliable perception features, For convolution mapping operations, This is a multi-scale aggregation feature. For a reliable occlusion guide map, This is an element-wise multiplication operation.
8. The preoperative and intraoperative three-dimensional spinal registration method based on multi-scale reliable occlusion fusion according to claim 7, characterized in that, The COGF receiver obstruction / trusted perception feature in the S4 signaling and guidance fusion module shallow decoding features First, map the two feature paths to a unified fusion space to obtain... Then based on adaptive guiding weights Perform complementary weighted fusion: right Perform residual augmentation to obtain the final enhanced features used for deformation prediction. The formula is: ,in, Features after mapping For adaptive guiding weights, As a temporary fusion feature, This represents the residual enhancement mapping function.
9. The preoperative and intraoperative three-dimensional spinal registration method based on multi-scale reliable occlusion fusion according to claim 8, characterized in that, The enhanced features will be described in S5. With moving image features Channel splicing is performed, and the input deformation field prediction layer is used to obtain the three-dimensional deformation field through three-dimensional convolutional regression: in, For a three-dimensional deformation field, For the deformation field prediction function, This is for channel splicing operations. To ultimately enhance features, To reconstruct three-dimensional data during the operation.
10. The preoperative and intraoperative three-dimensional spinal registration method based on multi-scale reliable occlusion fusion according to claim 9, characterized in that, The three-dimensional deformation field in S5 With intraoperative reconstruction of three-dimensional data Input the spatial transformation module STN to complete spatial alignment and obtain the registered 3D data: The network is trained using a joint loss function of image similarity and deformation smoothness. in, For the registered 3D data, For spatial transformation operations, For the total loss, For similarity loss, For deformation smoothing loss, This is the loss weighting coefficient.