Spinal medical image registration method and system based on semantic reconstruction

Through a semantic reconstruction-based method, the 2D-3D and 3D-3D feature extraction models combined with cross-attention parallel dual U-type network for registration of spinal medical images is solved, and the problems of both accuracy and efficiency in the prior art are achieved, and high-precision 3D/2D registration is achieved.

CN120047499AActive Publication Date: 2025-05-27SHANGHAI JIAOTONG UNIV

Patent Information

Application Number
CN202311586648.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-05-27
Estimated Expiration
2043-11-24

AI Technical Summary

Technical Problem

The prior art has problems in the 3D/2D registration of spinal medical images that large-scale and high-precision cannot be met simultaneously, insufficient utilization of joints/key semantic information, and inability to compatible with efficiency and accuracy.

Method used

Using a semantic reconstruction method, the spine 3D-CT image data and its corresponding 2D image data are constructed, and the feature map is extracted using the 2D-3D reconstruction model and the 3D-3D feature extraction model, and the feature map is fused and registered by a parallel dual U-shaped network model based on cross attention. A multi-weight loss function optimization model based on pixels and semantic information is adopted.

Benefits of technology

It effectively avoids registration failure and error problems caused by information and accuracy losses, and makes full use of the joint-semantic information extracted from the 2D-3D reconstruction network and the 3D-3D feature extraction network, which improves the accuracy and accuracy of registration, and reduces the number of image shootings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047499A_ABST
    Figure CN120047499A_ABST
Patent Text Reader

Abstract

The invention provides a spinal medical image registration method and system based on semantic reconstruction. Constructing spinal column 3D-CT image data and corresponding 2D image data of the spinal column 3D-CT image data; providing a 2D-3D reconstruction model, and reconstructing the 2D image data to obtain a 3D feature map b; providing a 3D-3D feature extraction model, and taking 3D-CT image data as input to obtain a feature map s with the same dimension as the 3D feature map; providing a parallel double-U-shaped network model based on cross attention, and fusing and registering the feature map b and the feature map s; a multi-weight loss function based on pixel and semantic information is adopted for optimization; and training to obtain a registration model, and performing spine medical image registration by using the registration model. According to the method, the information and precision loss caused by dimension reduction is avoided, the 2D / 3D information fusion capability is improved, the attention weight of the key semantic region is improved, and accurate spine image registration can be realized while the image shooting times can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image registration. Specifically, it relates to a method and system for spinal medical image registration based on semantic reconstruction, and also relates to a corresponding computer terminal and a computer-readable storage medium. Background Art

[0002] Currently, during spinal surgery, navigation is usually carried out at the surgical site relying on multi-dimensional image guidance. Before surgery, high-quality 3D-CT images are used for surgical planning and simulation. During surgery, X-rays generated by a C-arm are used for surgical guidance. An essential technology is to rigidly register the high-quality pre-scanned 3D-CT images with the X-ray images generated by the intraoperative C-arm. By guiding the physician to obtain the correct anatomical specific standard projections, the registered intraoperative imaging can further guide, monitor, and evaluate the surgical results. However, the existing C-arm positioning is usually performed manually by doctors, which not only requires doctors to have rich experience, but also requires doctors to perform multiple fluoroscopies in most cases, increasing not only the surgical time but also the radiation received by patients during the operation. In previous work, some used external tracking hardware for auxiliary positioning, software assistance, and optimization algorithms, but there were certain limitations and there was still a certain distance from clinical practice.

[0003] With the development of deep learning, there has been substantial progress in the intelligent processing of medical images using convolutional neural networks. Automatic registration based on deep learning has been widely studied thanks to the Digitally Reconstructed Radiograph (DRR) technology. Researchers can obtain a large amount of simulated X-ray data, which meets the problem that deep learning model training requires a large amount of data. Automatic registration technology based on deep learning is mainly divided into iterative registration and direct registration. Early work generally adopted reinforcement learning or iterative registration based on conditional judgment, and used image similarity as a criterion to measure the quality of registration. Although the registration accuracy has been improved, the problems of more X-ray exposure of patients caused by multiple scans required for registration and the registration time have not been effectively solved. In addition, due to the non-convex nature of the similarity metric, when the initial pose of the 3D model exceeds the capture range, inaccurate registration results may occur for such technologies.

[0004] In contrast to iterative registration, direct registration theoretically only requires one registration, which greatly alleviates the problems of radiation dose and time. For example, Li et al. inferred code vectors from the input X-ray images and further generated DRR images, enabling the registration of lateral skull images under unsupervised conditions (Li, Peixin, et al. "Non-rigid 2D-3D registration using convolutional autoencoders." 2020 IEEE 17th International Symposium on Biomedical Imaging (ISBI). IEEE, 2020.). Wu et al. constructed an architecture that combines deep learning and machine learning to achieve high-precision patella registration (Wu, Jing, Emam E. Abdel Fatah, and Mohamed R. Mahfouz. "Fully automatic initialization of two-dimensional–three-dimensional medical image registration using hybrid classifier." Journal of Medical Imaging 2.2 (2015): 024007-024007.). Further, work related to spinal registration has also been explored. Esfandiari et al. proposed a deep learning-based method to remove implants in the spine (Esfandiari, Hooman, et al. "Deep learning-based X-ray inpainting for improving spinal 2D-3D registration." The International Journal of Medical Robotics and Computer Assisted Surgery 17.2 (2021): e2228). While Kausch et al. considered k-wire and screw implants in spinal registration, increasing the robustness of the network (Kausch, Lisa, et al. "C-arm positioning for standard projections during spinal implant placement." Medical Image Analysis 81 (2022): 102557.).

[0005] In the above-mentioned technology, the registration result is generally evaluated by the error of 6DOF. However, registration failures still occur in complex situations or when the initial error is large. In contrast, feature-based technologies have also been extensively studied. Features include semantic information such as segmentation maps and fiducial points to assist registration. On the other hand, previous direct registration requires reducing the 3DCT image to 2D before registering it with the 2D X-ray. Inevitably, information loss occurs during the dimensionality reduction process, so it is also limited by a smaller capture range. With the improvement of hardware capabilities, technologies for upsampling 2D to 3D have gradually been attempted. Shen et al. mapped the patient's projection radiographs to the corresponding 3D anatomical structures, and the trained network can generate the patient's 3D tomographic X-ray images from a single projection view (Shen, Liyue, Wei Zhao, and Lei Xing. "Patient-specific reconstruction of volumetric computed tomography images from a single projection view via deep learning." Nature biomedical engineering 3.11(2019):880-888.). Mi et al. proposed a method called SGReg to reconstruct 2D and 3D images and predict the segmentation map, so that the 3D / 2D data has a dimensional correspondence (Mi, Jia, et al. "SGReg: segmentation guided 3D / 2D rigid registration for orthogonal x-ray and CT images in spine surgery navigation." Physics in Medicine and Biology(2023).). Although extracting 3D features from 2D images avoids information loss, it also poses a challenge to the hardware requirements. How to achieve a balance between the two remains an urgent problem to be solved in this field.

[0006] In summary, although there are currently medical image registration technologies based on different methods, the 3D / 2D registration of CT and X-ray images of the same patient still faces problems such as the inability to simultaneously meet a large range and high accuracy, insufficient utilization of joint / critical semantic information, and incompatibility between efficiency and accuracy. Summary of the Invention

[0007] In view of the above deficiencies in the prior art, the present invention provides a spine medical image registration method and system based on semantic reconstruction, and also provides a corresponding computer terminal and computer-readable storage medium.

[0008] According to one aspect of the present invention, there is provided a method for registering spinal medical images based on semantic reconstruction, including:

[0009] Construct spinal 3D-CT image data and its corresponding 2D image data as a training data set;

[0010] Provide a 2D-3D reconstruction model, and use the 2D-3D reconstruction model to reconstruct the 2D image data to obtain a 3D feature map b corresponding to the 2D image data;

[0011] Provide a 3D-3D feature extraction model, take the 3D-CT image data as the input of the 3D-3D feature extraction model, and obtain a feature map s corresponding to the 3D-CT image data with the same dimension as the 3D feature map;

[0012] Provide a parallel dual U-shaped network model based on cross-attention, and use the parallel dual U-shaped network model to fuse and register the feature map b and the feature map s;

[0013] Optimize the parallel dual U-shaped network model by using a multi-weight loss function based on pixel and semantic information;

[0014] Through the above steps, a registration model is trained, and this registration model is used to obtain the registration result of spinal medical images.

[0015] Preferably, the construction of the spinal 3D-CT image data and its corresponding 2D image data includes:

[0016] Obtain spinal CT data, unify the data format of the spinal CT data, and perform central cropping preprocessing to obtain preprocessed spinal CT data;

[0017] Perform rigid body transformation on the preprocessed spinal CT data to generate floating CT data; wherein, the rigid body transformation is represented by six parameters, and the components of translation and rotation are respectively the displacement components of translation in three axial directions (t x , t y , t z ) and the rotation components around three axial directions (r x , r y , r z ); different combinations of 6 DOF parameters are randomly generated for each parameter within a set range; then the DOF parameters are converted into a rigid body transformation matrix, and finally, according to the DOF parameters and the preprocessed spinal CT data, the corresponding floating CT data is generated;

[0018] Perform DRR projection on the floating CT data, and perform bitwise inversion and adaptive equalization on the image pixel values to obtain the final simulated C-arm X-ray data; finally, perform DRR projection on the preprocessed spinal CT data without rigid body transformation to generate standard registration reference 2D image data.

[0019] Preferably, provide a 2D-3D reconstruction model, and use the 2D-3D reconstruction model to reconstruct the 2D image data to obtain a 3D feature map b corresponding to the 2D image data, including:

[0020] Construct a 2D-3D reconstruction model, use the 2D image data and the manually segmented spinal mask as the input of the 2D-3D reconstruction model, and use the 2D-3D reconstruction model to reconstruct the 3D-CT volume image by projecting the 2D image data at a single angle, that is, obtain the 3D feature map b.

[0021] Preferably, the 2D-3D reconstruction model includes: a representative network, a generation network, and a conversion layer connected between the representative network and the generation network; where:

[0022] The representative network is used layer by layer to extract multi-scale features of the 2D image data, convert high-dimensional data into an embedded representation, and obtain the semantic information of the 3D structure hidden in the input 2D image data.

[0023] The conversion layer is used to learn the manifold mapping function corresponding to the extracted multi-scale features, so that the extracted multi-scale features cross dimensions.

[0024] The generation network is mainly composed of 3D deconvolution blocks, and is used to reconstruct a high-dimensional image from the feature information obtained from the conversion layer, that is, the 3D volume image corresponding to the 2D image data projection.

[0025] Preferably, provide a 3D-3D feature extraction model, use the 3D-CT image data as the input of the 3D-3D feature extraction model, and obtain a feature map s corresponding to the 3D-CT image data with the same dimension as the 3D feature map, including:

[0026] Use a 3D Res-NET network to construct a 3D-3D feature extraction model;

[0027] Use the 3D-CT image data and the manually segmented spinal mask as the input of the 3D-3D feature extraction model. After downsampling the 3D-CT image data, under the action of convolution and deconvolution, extract the key spinal information in the 3D-CT image data and output to obtain the feature map s.

[0028] Preferably, a parallel dual U-shaped network model based on cross-attention is provided, and the parallel dual U-shaped network model is used to fuse and register the feature map b and the feature map s, including:

[0029] Combine the 3D feature map b and the feature map s with the 2D image data and the 3D-CT image data respectively to form a key region feature map;

[0030] Perform window partitioning and window region partitioning on the key region feature map;

[0031] Construct a window-based multi-head cross-attention mechanism for calculating new features with corresponding correlation degrees between the input key region feature maps;

[0032] Provide a parallel dual U-shaped network, use the multi-head cross-attention mechanism as the convolutional layer of the parallel dual U-shaped network, and construct a parallel dual U-shaped network model, which is used to output the registration result of the spinal medical image.

[0033] Preferably, optimizing the parallel dual U-shaped network model by using a multi-weight loss function based on pixel and semantic information includes:

[0034] The loss function for measuring the parallel dual U-shaped network model is composed of the inter-pixel loss in the image domain and the perceptual loss. Among them: the inter-pixel loss is used to measure the difference between the predicted image and the target image, and the perceptual loss is used to solve the problem of image over-smoothing caused by the inter-pixel loss. Then, a multi-weight loss function L total is:

[0035]

[0036] Among them: x is the original image, x * is the predicted image, Cov(·) is the covariance of two images, Var(·) is the variance of the image itself, μ is the class sigmoid function; E(x,y) is, whc is, ρ is, is feature extraction.

[0037] Preferably, the above method of the present invention further includes:

[0038] Process the spinal medical image to be registered by using the registration model, and output the registration result of the spinal medical image.

[0039] According to another aspect of the present invention, a spinal medical image registration system based on semantic reconstruction is provided, including:

[0040] A data processing module, which is used to construct spinal 3D-CT image data and its corresponding 2D image data as a training data set;

[0041] A 2D-3D reconstruction module, which is used to provide a 2D-3D reconstruction model, and reconstruct the 2D image data by using the 2D-3D reconstruction model to obtain a 3D feature map b corresponding to the 2D image data;

[0042] A 3D-3D feature extraction module, which is used to provide a 3D-3D feature extraction model, take the 3D-CT image data as the input of the 3D-3D feature extraction model, and obtain a feature map s corresponding to the 3D-CT image data with the same dimension as the 3D feature map;

[0043] A registration module, which is used to provide a parallel dual U-shaped network model based on cross-attention, fuse and register the feature map b and the feature map s by using the parallel dual U-shaped network model; optimize the parallel dual U-shaped network model by using a multi-weight loss function based on pixel and semantic information, and is used to output the optimal registration result of the spinal medical image.

[0044] According to the third aspect of the present invention, there is provided a computer terminal, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it can be used to execute the method described in any one of the above of the present invention, or run the system described in any one of the above of the present invention.

[0045] According to the fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can be used to execute the method described in any one of the above of the present invention, or run the system described in any one of the above of the present invention.

[0046] Due to the adoption of the above technical solutions, compared with the prior art, the present invention has at least one of the following beneficial effects:

[0047] The method and system for registering spinal medical images based on semantic reconstruction provided by the present invention are different from the traditional technology of reducing 3DCT to 2D and then registering it with 2D X-ray. The 2D information is used to extract a 3D feature map through a 2D-3D reconstruction network, which can effectively avoid the problems of registration failure and large errors caused by information and accuracy loss.

[0048] The method and system for registering spinal medical images based on semantic reconstruction provided by the present invention make full use of the bone joint-semantic information extracted by the 2D-3D reconstruction network and the 3D-3D feature extraction network, reduce the size of the 3D image on the premise of avoiding the loss of key information, and ensure the accuracy and accuracy of registration without increasing the computing power required by the network. Description of the Drawings

[0049] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments read in conjunction with the accompanying drawings:

[0050] Figure 1 It is a flowchart of the spine medical image registration method based on semantic reconstruction in an embodiment of the present invention.

[0051] Figure 2 It is a schematic diagram of the working principle of 2D-3D spine medical image registration based on semantic reconstruction in a preferred embodiment of the present invention.

[0052] Figure 3 It is a schematic diagram of the feature fusion based on cross-attention in a preferred embodiment of the present invention.

[0053] Figure 4 It is a schematic diagram of the parallel dual U-shaped network model based on cross-attention in a preferred embodiment of the present invention.

[0054] Figure 5 It is a spine registration effect diagram in a specific application example of the present invention; among them, (a) and (d) are fixed images respectively, (b) and (e) are floating images respectively, and (c) and (f) are registration prediction images respectively.

[0055] Figure 6 It is a schematic diagram of the component modules of the spine medical image registration system based on semantic reconstruction in an embodiment of the present invention. Detailed implementation manners

[0056] The following details the embodiments of the present invention: These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

[0057] An embodiment of the present invention provides a spine medical image registration method based on semantic reconstruction. This method uses a deep learning method to replace the traditional method to achieve 2D-3D registration of the spine 3D-CT medical image and 2D-X-ray image of the same patient, solving the problems that have existed in this work, such as the inability to simultaneously meet large range and high precision, insufficient utilization rate of bone joint / critical semantic information, and incompatibility between efficiency and accuracy, so as to assist doctors in obtaining more accurate detection effects and improving work efficiency.

[0058] As Figure 1 shown, the spine medical image registration method based on semantic reconstruction provided in this embodiment may include the following operations:

[0059] S1. Construct the 3D-CT image data of the spine and its corresponding 2D image data as the training dataset;

[0060] S2. Provide a 2D-3D reconstruction model, and use the 2D-3D reconstruction model to reconstruct the 2D image data to obtain the 3D feature map b corresponding to the 2D image data;

[0061] S3. Provide a 3D-3D feature extraction model, take the 3D-CT image data as the input of the 3D-3D feature extraction model, and obtain the feature map s corresponding to the 3D-CT image data with the same dimension as the 3D feature map;

[0062] S4. Provide a parallel dual U-shaped network model based on cross-attention, and use the parallel dual U-shaped network model to fuse and register the feature map b and the feature map s;

[0063] S5. Optimize the parallel dual U-shaped network model by using a multi-weight loss function based on pixel and semantic information;

[0064] S6. Obtain a registration model through the above S1-S5 training, and this registration model is used to obtain the registration result of the spine medical image.

[0065] In some preferred embodiments, the above S1, constructing the 3D-CT image data of the spine and its corresponding 2D image data, includes:

[0066] S11. Obtain the spine CT data, unify the data format of the spine CT data, and perform central cropping preprocessing to obtain the preprocessed spine CT data;

[0067] S12. Perform rigid body transformation on the preprocessed spine CT data to generate floating CT data; among them, the rigid body transformation represents translation and rotation components through six parameters. The six parameters of the rigid body transformation are the displacement components of translation in three axial directions (t x , t y , t z ) and the rotation components around three axial directions (r x , r y , r z ), and different combinations of 6 DOF parameters are randomly generated for each parameter within the set range; the rigid body transformation matrix represents the rigid body transformation as a 4x4 matrix, which transforms a point from the initial position to a new position. In the present invention, the rigid body transformation matrix is usually represented as T, and its form is as follows:

[0068]

[0069] Among them, t is a 3x1 translation vector R is a 3x3 rotation matrix, which is obtained by multiplying three rotation matrices R x 、R y and R z to get R = R x *R y *R z , where

[0070]

[0071]

[0072]

[0073] Subsequently, the DOF parameters are converted into a rigid body transformation matrix according to the above formula, and finally the corresponding floating CT data is generated according to the DOF parameters and the preprocessed spinal CT data;

[0074] S13, perform DRR projection on the generated floating CT data, and perform bitwise inversion and adaptive equalization processing on the image pixel values to obtain the final simulated C-arm X-ray data. Finally, perform DRR projection on the preprocessed spinal CT data without rigid body transformation to generate standard registration reference 2D image data.

[0075] In some preferred embodiments, in S2 above, a 2D-3D reconstruction model is provided, and the 2D image data is reconstructed using the 2D-3D reconstruction model to obtain a 3D feature map b corresponding to the 2D image data, including:

[0076] S21, construct a 2D-3D reconstruction model;

[0077] S22, use the 2D image data and the manually segmented spinal mask as the input of the 2D-3D reconstruction model, and reconstruct the 3D-CT volume image by projecting the 2D image data at a single angle using the 2D-3D reconstruction model, that is, obtain the 3D feature map b.

[0078] In some preferred embodiments, in S21 above, the 2D-3D reconstruction model includes: a representative network, a generation network, and a conversion layer connected between the representative network and the generation network; where:

[0079] The representative network is used layer by layer to extract the multi-scale features of the 2D image data, convert the high-dimensional data into an embedded representation, and obtain the semantic information of the 3D structure hidden in the input 2D image data;

[0080] The conversion layer is used to learn the manifold mapping function corresponding to the extracted multi-scale features, so that the extracted multi-scale features cross dimensions;

[0081] The generation network is mainly composed of 3D deconvolution blocks, which are used to reconstruct a high-dimensional image from the feature information obtained from the conversion layer, that is, the 3D volume image corresponding to the projection of the 2D image data.

[0082] In some preferred embodiments, in S3 above, a 3D-3D feature extraction model is provided, and the 3D-CT image data is used as the input of the 3D-3D feature extraction model to obtain a feature map s corresponding to the 3D-CT image data with the same dimension as the 3D feature map, including:

[0083] S31, constructing a 3D-3D feature extraction model using a 3D Res-NET network;

[0084] S32, using the 3D-CT image data and the manually segmented spine mask as the input of the 3D-3D feature extraction model. After downsampling the 3D-CT image data, under the action of convolution and deconvolution, the key spine information in the 3D-CT image data is extracted and output to obtain the feature map s.

[0085] In some preferred embodiments, in S4 above, a parallel dual U-shaped network model based on cross-attention is provided, and the feature map b and the feature map s are fused and registered using the parallel dual U-shaped network model, including:

[0086] S41, performing window partitioning and window region partitioning on the key region feature maps b and s;

[0087] S42, constructing a window-based multi-head cross-attention mechanism for calculating new features with corresponding correlation degrees between the input key region feature maps;

[0088] S43, providing a parallel dual U-shaped network, using the multi-head cross-attention mechanism as the convolutional layer of the parallel dual U-shaped network, and constructing a parallel dual U-shaped network model, which is used to output the registration result of the spine medical image.

[0089] In some preferred embodiments, in S5 above, a multi-weight loss function based on pixel and semantic information is used to optimize the parallel dual U-shaped network model, including:

[0090] The loss function for measuring the parallel dual U-shaped network model is composed of the inter-pixel loss in the image domain and the perceptual loss. Among them: the inter-pixel loss is used to measure the difference between the predicted image and the target image, and the perceptual loss is used to solve the problem of image over-smoothing caused by the inter-pixel loss. Then, a multi-weight loss function L total is:

[0091]

[0092] where: x is the original image, x *For the predicted image, Cov(·) is the covariance of two images, Var(·) is the variance of the image itself, μ is the class sigmoid function; E(x,y) is, whc is, ρ is, For feature extraction.

[0093] In some preferred embodiments, the method provided in the above embodiments of the present invention may further include the following operations:

[0094] S7. Use the registration model to process the spine medical image to be registered, and output the registration result of the spine medical image.

[0095] The following further details the technical solutions provided in the above embodiments of the present invention in conjunction with a preferred embodiment.

[0096] As Figure 1 shown, the spine medical image registration method based on semantic reconstruction provided in this preferred embodiment includes the following steps:

[0097] Step 1. Generation and preprocessing of spine 3D-CT image data and corresponding 2D image data;

[0098] Step 2. Use the 2D-3D reconstruction model to reconstruct the 2D image data to obtain a 3D feature map;

[0099] Step 3. Use the 3D-3D feature extraction model to process the 3D-CT image data to obtain a feature map of the same dimension;

[0100] Step 4. Use the parallel double U-shaped network model based on the cross-attention mechanism (cross-attention) to register after fusing the two groups of feature maps;

[0101] Step 5. Use the multi-weight loss function based on pixel and semantic information to further guide the parallel double U-shaped network model to obtain the optimal registration result;

[0102] Step 6. Train the registration model through Steps 1 to 5;

[0103] Step 7. Use the registration model to process the spine medical image to be registered to obtain the registration result of the spine medical image.

[0104] The above Step 1 includes the following steps:

[0105] Step 11. Obtain the spine CT data, unify the data format of the spine CT data, save it in the NifTI format, and perform central cropping preprocessing.

[0106] Step 12. For the preprocessed spine CT data, make each case of data at t x,y,z ∈ (-45mm, 45mm), r x,y,zRandomly generate different combinations of 6 parameters within the range of (-45°, 45°), where (t x , t y , t z ) are the displacement components of translation in three axial directions and (r x , r y , r z ) are the rotation components around three axial directions; Subsequently, convert the DOF parameters into a 4x4 rigid body transformation matrix according to the formula, and finally generate the corresponding floating CT data according to different parameters and the original spinal CT data.

[0107] Step 13: Use the DeepDRR tool to perform DRR projection on the randomly transformed CT data to generate C-arm X-ray data for simulation, and perform DRR projection on the original spinal CT data to generate standard registration reference 2D image data. Finally, perform bitwise inversion on the pixel values of the C-arm X-ray data and perform adaptive equalization processing on the obtained image.

[0108] The above step 2 includes the following steps:

[0109] Reconstruct the 3D-CT volume image by projecting the 2D image data at a single angle, that is, obtain the 3D feature map; where:

[0110] The 2D-3D reconstruction model used includes: a representative network, a generation network, and a conversion layer connected between the representative network and the generation network. The representative network extracts multi-scale features layer by layer and converts high-dimensional data into an embedded representation. Its purpose is to obtain the semantic representation of the hidden 3D structure in the input projection, such as the size and position of the spine; the representative network and the generation network are connected by a conversion layer. The purpose of the conversion layer is to learn the manifold mapping function corresponding to the extracted features so that the features can cross dimensions. The generation network is composed of 3D deconvolution blocks, and finally reconstructs the high-dimensional image from the feature information obtained from the conversion layer, that is, the 3D volume image corresponding to the projection. Through this architecture, the 2D projection and the 3D image share the same semantic feature representation in the feature domain, enabling the model to learn how to generate a 3D image from a 2D projection; during the training phase, input the artificial segmentation mask corresponding to its reconstructed volume at the same time to obtain the mask of the significant region in the X-ray reconstruction image data at the same time.

[0111] The above step 3 includes the following steps:

[0112] Use the 3D Res-NET network to construct a 3D-3D feature extraction model. Where:

[0113] After the input 3D-CT data is downsampled, under the action of convolution and deconvolution, the key information of the spine is extracted and the final key features are output. During the training process, the manually segmented spine mask is simultaneously used as the input of the network to obtain the feature information and the significant region mask of the three-dimensional image. Since the input of the feature extraction model is set to 3D information, the 3D image generated by CT does not need to be upsampled or downsampled, and only needs to obtain a feature map with the same dimension as the 3D feature map obtained by the reconstruction model after feature extraction.

[0114] Step 4 above includes the following steps:

[0115] Step 4.1: Combine the 3D feature map corresponding to the 2D data and the feature map corresponding to the 3D-CT data with the 2D image data and the 3D-CT image data respectively, and perform central cropping to form the key region feature maps b and s, which are used as the input of the registration network.

[0116] Step 4.2: Perform window partitioning and window region partitioning.

[0117] Step 4.3: Construct a window-based multi-head cross-attention mechanism; the multi-head cross-attention mechanism aims to calculate new features with corresponding correlation degrees between the input feature b and the feature s through the attention mechanism; the features b and s are used for window-based attention calculation after generating the basic window and the search window through Step 4.2; each basic window S ba is linearly projected and layer-normalized into the query set, and each search window S se is linearly projected and layer-normalized into the key set and the value set, and then the cross-attention between the two windows is calculated by the window-based multi-head cross-attention algorithm; finally, the new output set is sent to a two-layer MLP with Gelu non-linear mapping after passing through the LayerNorm (LN) layer, and finally outputs new features with corresponding correlation degrees between the feature b and the feature s after passing through the multi-head cross-attention mechanism;

[0118] Step 4.4: Construct a parallel dual U-shaped network model based on cross-attention mechanism as the registration network. The registration network uses a dual U-shaped network based on the cross-attention feature fusion module in parallel to connect the network floating data b and the standard data s, and outputs two feature maps b' and s' of the same size after feature fusion. The two parallel U-shaped networks adopt the structure of U-NET in the encoding and decoding parts, downsample features in the encoder, and upsample features and perform skip connections between the encoder and the decoder in the decoder. Use the window-based multi-head attention mechanism in Step 4.3 to replace convolution, enabling the registration network to exchange cross-image information. Finally, append a classification head at the end of the network. The two feature maps b' and s' are concatenated in the channel dimension and then averaged and input into the classification head. After linear mapping and activation function, 6 rigid body transformation parameters are finally output.

[0119] The above Step 5 includes the following steps:

[0120] The loss function for measuring the registration network is composed of the pixel and perceptual loss in the image domain. The pixel-wise loss is used to measure the difference between the predicted image and the target image, while the perceptual loss alleviates the problem of image over-smoothing caused by the pixel loss, and can further improve the structural similarity and perceptual similarity between the predicted image and the target image. The formula is:

[0121]

[0122] In the formula, L total is the multi-weight loss function based on pixel and semantic information, x is the original image, x * is the predicted image, Cov(·) is the covariance of the two images, Var(·) is the variance of the image itself, μ is the class sigmoid function; E(x,y) is, whc is, ρ is, is feature extraction.

[0123] Next, a specific application example is used to further elaborate on the technical solutions provided in the above embodiments of the present invention in detail.

[0124] The method for registering spinal medical images based on semantic reconstruction adopted in this specific application example, as Figure 1 shown, specifically includes the following contents:

[0125] Step S1: Generation and preprocessing of 3D-CT images and corresponding 2D data;

[0126] Step S2: Use the 2D-3D reconstruction model to reconstruct the 2D image to obtain a 3D feature map;

[0127] Step S3: Use the 3D-3D feature extraction model to process the 3D image to obtain a feature map of the same dimension;

[0128] Step S4: Use a parallel dual U-shaped network model based on cross-attention mechanism to register the two sets of feature maps after fusion;

[0129] Step S5: Use a multi-weight loss function based on pixel and semantic information to further guide the parallel dual U-shaped network model to obtain the optimal registration result;

[0130] Step S6: Obtain a registration model through training in Steps S1 to S5;

[0131] Step S7: Use the registration model to process the spine medical image to be registered to obtain the registration result of the spine medical image

[0132] To achieve the registration of 2D-3D spine registration, the present invention constructs a registration model based on a deep learning network architecture as shown in Figure 2 The following will detail this specific application example.

[0133] Let the 2D x-ray image M x-ray be taken from the anterior-posterior (AP) view of the C-arm, while the 3D image F CT is taken preoperatively by 3D CT. They respectively represent the floating image and the fixed image in the spatial domain The present invention mainly focuses on 2D-3D direct registration. Therefore, for M x-ray n = 2, and for F CT n = 3. 2D / 3D registration is to match the 2D projection with the 3D grayscale image. For a rigid structure such as the spine, rigid body transformation is mostly used for registration. Therefore, first reconstruct the three-dimensional feature map M x-ray with n = 3 from M CT . The goal is to learn the rigid body transformation matrix from M CT to F CT . Specifically, use the registration model provided in the above embodiment of the present invention to parameterize the rigid body registration problem as a function f θ (M CT , F CT ) = T, where θ is a set of rigid body transformation parameters, and T represents the predicted rigid body transformation matrix.

[0134] In Step S1, for the preprocessing of the spine image and the production of the experimental dataset, it specifically includes:

[0135] Select the currently largest publicly available annotated spinal CT dataset, CTSpine1K, as the dataset for the experiment. CTSpine1K collects and annotates a large-scale spinal CT dataset from multiple sources and different manufacturers, totaling 1,005 CT volumes with different appearances (more than 500,000 labeled slices and more than 11,000 vertebrae). The CTSpine1K dataset is divided into a training dataset (610) and a public test dataset (395). All data is uniformly saved in the NIfTI format, with a slice size of 512*512, and the number of slices is centrally cropped to 512, so the overall size of all data is 512 3 Select 279 cases from CTSpine1K as the training set and 81 cases as the test set.

[0136] For each existing CT data, at t x,y,z ∈ (-45mm, 45mm0, r x,y,z ∈ (-45°, 45°), different combinations of 6 parameters are randomly generated. Subsequently, the DOF parameters are converted into a 4x4 rigid body transformation matrix according to the formula. Finally, according to the affine_transform method in the ndimage library of the scipy toolkit, corresponding floating CT data is generated based on different parameters and the original CT data, totaling 1,674 groups of data. Considering the actual clinical situation, the parameter range in the test set is t x,y,z ∈ (-20±15mm, 20±15mm), r x,y,z ∈ (-20±15°, 20±15°), and 4 different combinations of data are generated for each case, totaling 324 groups of data.

[0137] DeepDRR is a deep learning-based method for generating high-quality digital radiography (DRR). DRR is an image generated by simulating the passage of X-rays through human tissues at different projection angles and is commonly used in the fields of computer-aided diagnosis (CAD) and medical image processing. Different from traditional physics-based DRR methods, DeepDRR uses deep learning techniques such as deep convolutional neural networks (CNNs) to generate more realistic and accurate DRR images by learning the features and structures of a large amount of medical image data.

[0138] Use the DeepDRR tool to perform DRR projection on the randomly transformed CT data to generate simulated C-arm X-ray data, and perform DRR projection on the original CT data to generate a standard registration reference image. The size of the simulated X-ray data is 128 2。To be closer to clinically realistic C-arm X-ray data, the data generated by DeepDRR is preprocessed in two steps. The preprocessing process includes first performing bitwise inversion of image pixel values using the bitwise_not function in the OpenCV library, and then performing adaptive equalization on the image using the createCLAHE function in the OpenCV library.

[0139] In step S2, for the feature extraction of 2D X-ray images, different from the traditional 3D / 2D registration method that first projects the 3D CT image into 2D and then registers it with the 2D image, a single-angle 2D X-ray projection is used to reconstruct the volume image of the 3D CT. The volume reconstruction network can be specifically divided into a representative network and a generation network. The network inputs the preprocessed simulated X-ray image with a size of 128 2 The representative network structure is a 2D 4*4 convolutional layer. The first convolution increases the number of feature channels to 64, and then under the action of convolution, the features increase sequentially at a rate of 2 times per layer to 4096. The representative network extracts multi-scale features layer by layer, converting high-dimensional data into an embedded representation, aiming to obtain the semantic representation of the hidden 3D structure in the input projection, such as the size and position of the spine. The representative network and the generation network are connected through a conversion layer. The purpose of the conversion layer is to learn the manifold mapping function corresponding to the extracted features, enabling the features to cross dimensions. The conversion layer is specifically obtained by connecting a 2D convolution operation with a kernel size of 1×1 and a 3D transposed convolution layer with a kernel size of 1×1×1. It can connect the 2D and 3D feature space conversion layers while keeping the feature size unchanged. The generation network is composed of 3D transposed convolution blocks. Similar to the structure of the generation network, under the action of transposed convolution, it decreases the number of channels at a rate of 2 times per layer to 64, and finally reconstructs a high-dimensional image with a size of 128 3 from the feature information obtained from the conversion layer, that is, the volume image corresponding to the projection. Through this architecture, the 2D projection and the 3D image share the same semantic feature representation in the feature domain, enabling the network to learn how to generate a 3D image from a 2D projection. During the training phase, the artificial segmentation mask corresponding to its reconstructed volume is input simultaneously to obtain the significant region mask in the X-ray reconstructed image data.

[0140] In step S3, for the feature extraction of three-dimensional CT images, a 3D Res-NET network is used. The difference is that in this specific application example, the input size is 512 3 downsampled 128 3 image data. Under the action of convolution and transposed convolution, the key information of the spine is extracted and the final key features are output, and its size is still 128 3. During the training process, the manually segmented spine mask is used as the input of the network to obtain the feature information and the significant region mask of the 3D image. Since the input of the registration network is set to 3D information, the 3D image generated by CT does not need to be upsampled or downsampled, and only needs to be feature-extracted to obtain a feature map with the same dimension as the 2D reconstruction network.

[0141] In step S4, as Figure 3 shown, the parallel dual U-shaped network based on the cross-attention mechanism is used to register the two groups of feature maps after fusion.

[0142] Step 4.1: After the 2D X-ray and 3D CT images pass through the 2D reconstruction network and the 3D feature extraction network respectively, they will extract intermediate feature maps with different dimensions and combine them with the image data of the standard size 512 3 and center-crop them into significant regions b and s of 128 3 for use in the subsequent registration network. It should be noted that the feasibility of doing this lies in the fact that the single joint region accounts for a relatively small proportion of the overall image. Finally, the semantic information containing only the significant region obtained is input into the registration network to obtain an accurate registration result.

[0143] Step 4.2: Window partitioning and window region partitioning

[0144] The multi-scale window division includes two different methods: window partitioning and window region partitioning. After step 4.1, the input features b and s are divided into windows of different sizes. Window partitioning directly divides the features into the basic window set S of size n×h×w×d ba , and window region partitioning uses magnification factors α, β, and γ to enlarge the window size, where h = w = d = 2, and α = β = γ = 3. Therefore, the calculation formulas for the basic window and the search window sizes are:

[0145] h ba , w ba , d ba = h, w, d

[0146] h se , w se , d se = αh, β·w, γ·d

[0147] where h ba , w ba , d ba are the sizes of the basic window, and h se , w se , d se are the sizes of the search window. In order to obtain two window sets with the same number, window region partitioning uses a sliding window and sets the step size to the basic window size, so that Sse has a size of n×α•h×β•w×γ•d.

[0148] Step 4.3: Window-based multi-head cross-attention mechanism

[0149] The multi-head cross-attention mechanism aims to calculate new features with corresponding correlations between the input feature b and the feature s through the attention mechanism. The feature b and the feature s are used for window-based attention calculation after generating the basic window and the search window in Step 4.2. Each basic window S ba is linearly projected and layer-normalized into the query set, and each search window S se is linearly projected and layer-normalized into the key and value sets. Subsequently, the cross-attention degree between the two windows is calculated by the window-based multi-head cross-attention algorithm, as Figure 4 shown, and the calculation formula is as follows:

[0150]

[0151] where Q ba 、K se and V se represent the query, key, and value matrices respectively.

[0152] Finally, the new feature output set is sent to a two-layer MLP with Gelu non-linear mapping after passing through the LayerNorm (LN) layer to enhance the learning ability.

[0153] Step 4.4: Parallel dual U-shaped network based on cross-attention mechanism

[0154] The parallel dual U-shaped network based on the cross-attention mechanism is used as the registration network. The dual U-shaped network is used to extract the features of the moving image and the fixed image respectively, and they are connected through the cross-attention feature fusion module, as Figure 2 shown, and the network input size is 128 3For the floating data b and the standard data s, two feature maps b′ and s′ of the same size after feature fusion are output. The two parallel networks adopt the U-NET structure in the encoding and decoding parts, downsample features in the encoder, upsample features in the decoder, and perform skip connections between the encoder and the decoder, achieving the purpose of continuously refining features. The window-based multi-head attention mechanism in step 4.3 is used to replace convolution, enabling the registration network to exchange cross-image information. Finally, a classification head is appended at the end of the network, which is implemented by two consecutive multi-layer perceptron (MLP) layers and uses the hyperbolic tangent (Tanh) activation function. The two feature maps b′ and s′ are concatenated and averaged in the channel dimension and then input into the classification head. After linear mapping and activation function, 6 rigid body transformation parameters are finally output. Finally, the corresponding standard CT data is generated by performing rigid body transformation on the rigid body transformation parameters and the floating CT data according to the affine_transform method in the ndimage library of the scipy toolkit.

[0155] In step S5, the loss function for measuring the network is composed of the pixels in the image domain and the perceptual loss.

[0156] The predicted 3DCT is obtained by combining the obtained predicted pose with the original 3DCT, and the error between it and the standard 3DCT is calculated. This error can be further divided into the NCC loss at the pixel level and the perceptual loss at the structure-semantic level. The pixel-wise loss is used to measure the difference between the predicted image and the target image, and the algorithm is optimized by making the predicted image closer to the target image:

[0157]

[0158] where x represents the original image, x * represents the predicted image, Cov(·) represents the covariance of the two images, and Var(·) represents the variance of the image itself.

[0159] Opposed to the pixel-level loss is the structural-semantic perceptual loss. This loss function is often used in the field of image denoising, aiming to alleviate the over-smoothing problem of images caused by pixel-level loss. Since the introduction of perceptual loss can further improve the structural similarity and perceptual similarity between the predicted image and the target image. Because it is more in line with human intuition, it has been used as an auxiliary or even a substitute for pixel-level loss in recent years. In the scenario of rigid spine registration, the main concern of physicians will focus on the spine to be registered, so it is also an ideal application scenario for perceptual loss. The traditional perceptual loss uses vgg16 as the backbone network for feature extraction. However, vgg16 is trained using conventional images, and the weights of its feature extraction are more suitable for general fields, and there will be certain deviations when applied to the medical field. In contrast, the encoder in the 3D feature extraction network is used as the feature extraction architecture, and the trained hyperparameters are directly migrated. Since this architecture is originally for the 3D reconstruction of 2D spine images, it can better extract the features of the spine, and a much better result than vgg16 was observed in a specific application experiment. When calculating the structural-perceptual loss, the predicted image and the target image are respectively fed into the frozen encoder for forward propagation. Specifically, the L2 loss is used to calculate the perceptual similarity of the feature maps obtained from the two images, and the specific formula is as follows:

[0160] Perceptual Loss (Formula)

[0161]

[0162] Represents feature extraction.

[0163] According to the different focuses of the two different LOSS and the experimental process, the mutual information loss is usually used to measure the degree of information association between images, while the perceptual loss is usually used to measure the similarity of image quality or content. The perceptual loss is usually related to the high-level feature representation of the neural network, and these features are more sensitive to the content and structure of the image. The mutual information loss can help the network better handle situations such as deformation because it focuses on the alignment of information. Combining them can comprehensively consider the information retention and visual quality of image alignment or generation, which helps to complete more accurate registration. By combining these two losses, the robustness of the algorithm can be improved, making it perform better in various situations. To sum up, this specific application example proposes an adaptive multi-dimensional loss function, and the specific formula is as follows:

[0164] L total = L NCC + μL perceptual

[0165] Among them, μ is a sigmoid-like function that adaptively allocates the weights of the NCC Loss and the perceptual Loss according to the set total number of network training rounds, so that the network focuses on the alignment error at the beginning of training and on the image visual quality in the later stage.

[0166] Combining the above methods, the specific application example inputs the moving position 2D-C arm X-ray and outputs the rigid body transformation matrix from the moving position to the standard position. Further, the standard position 3D-CT can be obtained, and then the 2D-C arm X-ray at the standard position can be obtained through projection, thereby realizing the direct rigid registration of the 2D C-arm X-ray and 3D CT image data, as Figure 5 shown in (a)-(f) in, effectively assisting the doctor to obtain a more accurate detection effect and improve work efficiency.

[0167] An embodiment of the present invention provides a spine medical image registration system based on semantic reconstruction, as Figure 6 shown, the system may include the following modules:

[0168] The data processing module is used to construct the spine 3D-CT image data and its corresponding 2D image data as the training data set;

[0169] The 2D-3D reconstruction module is used to provide a 2D-3D reconstruction model and reconstruct the 2D image data using the 2D-3D reconstruction model to obtain the 3D feature map b corresponding to the 2D image data;

[0170] The 3D-3D feature extraction module is used to provide a 3D-3D feature extraction model, take the 3D-CT image data as the input of the 3D-3D feature extraction model, and obtain the feature map s corresponding to the 3D-CT image data with the same dimension as the 3D feature map;

[0171] The registration module is used to provide a parallel dual U-shaped network model based on cross-attention, fuse and register the feature map b and the feature map s using the parallel dual U-shaped network model; optimize the parallel dual U-shaped network model using a multi-weight loss function based on pixel and semantic information for outputting the optimal registration result of the spine medical image.

[0172] In some preferred embodiments:

[0173] The data processing module: used for the generation and preprocessing of 3D-CT images and corresponding 2D data;

[0174] The 2D-3D reconstruction module: uses a 2D-3D reconstruction network to reconstruct the 2D image to obtain a 3D feature map;

[0175] 3D-3D Feature Extraction Module: Process 3D images using a 3D-3D feature extraction network to obtain feature maps of the same dimension;

[0176] Registration Module: Use a parallel dual U-shaped network based on cross-attention mechanism to fuse and register two sets of feature maps; Use a multi-weight loss function based on pixel and semantic information to further guide the network to obtain the optimal registration result.

[0177] The data processing module specifically includes the following: Select the spinal CT dataset and save it in NifTI format uniformly, perform central cropping, generate corresponding floating CT data according to different parameters and the original CT data, and use the DeepDRR tool to perform DRR projection on the randomly transformed CT data to generate simulated C-arm X-ray data, and perform DRR projection on the original CT data to generate a standard registration reference image. Finally, perform two preprocessing operations: bitwise inversion of the pixel values of the image generated by DeepDRR and adaptive equalization of the image.

[0178] The 2D-3D Reconstruction Module specifically includes the following: For the feature extraction of 2D X-ray images, use single-angle 2D X-ray projections to reconstruct the volume image of 3D CT. The volume reconstruction network can be specifically divided into a representative network and a generation network. The representative network extracts multi-scale features layer by layer, converts high-dimensional data into an embedded representation, while the generation network is composed of 3D deconvolution blocks, and reconstructs a high-dimensional image from the feature information obtained from the conversion layer connecting the two parts of the network, that is, the volume image corresponding to the projection. This network can learn how to generate 3D images from 2D projections, and simultaneously input the artificial segmentation mask corresponding to its reconstructed volume during the training phase to obtain the mask of the significant region in the X-ray reconstructed image data at the same time.

[0179] The 3D-3D Feature Extraction Module specifically includes the following details: For the feature extraction of three-dimensional CT images, use a 3DRes-NET network. During the training process, simultaneously use the artificially segmented spinal mask as the input of the network to obtain the feature information and significant region mask of the three-dimensional image. Since the input of the registration network is set to 3D information, the 3D images generated by CT data do not need to be upsampled or downsampled, and only need to obtain feature maps of the same dimension as the 2D reconstruction network after feature extraction.

[0180] The registration module is specifically divided into the following sub-modules:

[0181] Feature Map Extraction Sub-module: This sub-module combines the feature maps of the same dimension in the middle of the 2D X-ray and 3D CT images that have passed through the 2D reconstruction network and the 3D feature extraction network respectively, together with the image data of the standard size 512 3 and perform central cropping, and then put it into the next registration network;

[0182] Window Partition Sub-module: This sub-module is used to perform window partitioning and window region partitioning on the feature map;

[0183] Cross-Attention Mechanism Construction Sub-module: This sub-module is based on the multi-head cross-attention mechanism of the window;

[0184] Registration Network Sub-module: This sub-module is a parallel dual U-shaped network based on the cross-attention mechanism;

[0185] Optimization Sub-module: This sub-module is used to optimize the parallel dual U-shaped network; among them, the loss function for measuring the network is composed of the pixels in the image domain and the perceptual loss. The pixel-wise loss is used to measure the difference between the predicted image and the target image, while the perceptual loss alleviates the over-smoothing problem of the image caused by the pixel loss, and can further improve the structural similarity and perceptual similarity between the predicted image and the target image.

[0186] It should be noted that the steps in the method provided by the present invention can be implemented by corresponding modules, devices, units, etc. in the system. Those skilled in the art can refer to the technical solution of the method to implement the composition of the system. That is, the embodiments in the method can be understood as the preferred examples for constructing the system, and will not be elaborated here.

[0187] An embodiment of the present invention provides a computer terminal, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it can be used to execute the method of any one of the above embodiments of the present invention, or, run the system of any one of the above embodiments of the present invention.

[0188] Optionally, the memory is used to store programs; the memory can include volatile memory (English: volatile memory), such as random access memory (English: random-access memory, abbreviation: RAM), such as static random access memory (English: static random-access memory, abbreviation: SRAM), double data rate synchronous dynamic random access memory (English: Double Data Rate Synchronous Dynamic Random Access Memory, abbreviation: DDR SDRAM), etc.; the memory can also include non-volatile memory (English: non-volatile memory), such as flash memory (English: flash memory). The memory is used to store computer programs (such as application programs, functional modules, etc. for implementing the above method), computer instructions, etc. The above computer programs, computer instructions, etc. can be partitioned and stored in one or more memories. And the above computer programs, computer instructions, data, etc. can be called by the processor.

[0189] The computer programs, computer instructions, etc. described above can be stored in partitions in one or more memories. And the computer programs, computer instructions, data, etc. described above can be called by a processor.

[0190] A processor, configured to execute the computer program stored in the memory to implement each step in the method involved in the above embodiments or each module of the system. For details, reference can be made to the relevant descriptions in the foregoing method and system embodiments.

[0191] The processor and the memory can be of an independent structure or an integrated structure integrated together. When the processor and the memory are of an independent structure, the memory and the processor can be coupled and connected through a bus.

[0192] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can be used to execute the method of any one of the above embodiments of the present invention, or, run the system of any one of the above embodiments of the present invention.

[0193] The spine registration method, system, terminal and medium based on semantic reconstruction provided by the above embodiments of the present invention comprehensively utilize deep learning network technology. By reconstructing 2D projections into 3D feature maps, it avoids information and accuracy loss caused by dimensionality reduction, and adopts a dual U-shaped network based on cross-attention mechanism to improve the network's ability to fuse 2D / 3D information. The multi-weight loss function further enhances the attention weight of key semantic regions in the network. The spine registration method, system, terminal and medium based on semantic reconstruction provided by the above embodiments of the present invention can reduce the number of image captures and achieve accurate spine image registration at the same time.

[0194] Those skilled in the art know that in addition to implementing the system and its various devices provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system and its various devices provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same functions. Therefore, the system and its various devices provided by the present invention can be regarded as a hardware component, and the devices included therein for implementing various functions can also be regarded as the structure within the hardware component; the devices for implementing various functions can also be regarded as both software modules for implementing the method and the structure within the hardware component.

[0195] Matters not described in detail in the above embodiments of the present invention are all well-known technologies in the art.

[0196] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which does not affect the essence of the present invention.

Claims

1. A spine medical image registration method based on semantic reconstruction, characterized in that, it includes: Constructing spine 3D-CT image data and its corresponding 2D image data as a training data set; Providing a 2D-3D reconstruction model, and using the 2D-3D reconstruction model to reconstruct the 2D image data to obtain a 3D feature map b corresponding to the 2D image data; Providing a 3D-3D feature extraction model, taking the 3D-CT image data as the input of the 3D-3D feature extraction model to obtain a feature map s corresponding to the 3D-CT image data with the same dimension as the 3D feature map; Providing a parallel dual U-shaped network model based on cross attention, and using the parallel dual U-shaped network model to fuse and register the feature map b and the feature map s; Optimizing the parallel dual U-shaped network model by using a multi-weight loss function based on pixel and semantic information; Training a registration model through the above steps.

2. The spine medical image registration method based on semantic reconstruction according to claim 1, characterized in that, The constructing of the spine 3D-CT image data and its corresponding 2D image data includes: Obtaining spine CT data, unifying the data format of the spine CT data, and performing central cropping preprocessing to obtain preprocessed spine CT data; Perform a rigid body transformation on the preprocessed spinal CT data to generate floating CT data; wherein, the rigid body transformation is represented by six parameters, and the components of translation and rotation are respectively the displacement components of translation in three axial directions (t x , t y , t z ) and the rotation components about three axial directions (r x , r y , r z ); different combinations of 6 DOF parameters are randomly generated for each parameter within a set range; then the DOF parameters are converted into a rigid body transformation matrix, and finally, the corresponding floating CT data is generated according to the DOF parameters and the preprocessed spinal CT data; Performing DRR projection on the floating CT data, and performing bitwise inversion and adaptive equalization processing on the image pixel values to obtain the final simulated C-arm X-ray data; finally, performing DRR projection on the preprocessed spine CT data without rigid body transformation to generate standard registration reference 2D image data.

3. The spine medical image registration method based on semantic reconstruction according to claim 1, characterized in that, The providing of a 2D-3D reconstruction model, and using the 2D-3D reconstruction model to reconstruct the 2D image data to obtain a 3D feature map b corresponding to the 2D image data, includes: Constructing a 2D-3D reconstruction model, taking the 2D image data and the manually segmented spine mask as the input of the 2D-3D reconstruction model, and using the 2D-3D reconstruction model to reconstruct a 3D-CT volume image by projecting the 2D image data at a single angle, that is, obtaining the 3D feature map b; Wherein: The 2D-3D reconstruction model includes: a representative network, a generation network, and a conversion layer connected between the representative network and the generation network; wherein: The representative network is used layer by layer to extract multi-scale features of the 2D image data, convert high-dimensional data into an embedded representation, and obtain the semantic information of the hidden 3D structure in the input 2D image data; The conversion layer is used to learn the manifold mapping function corresponding to the extracted multi-scale features, so that the extracted multi-scale features cross dimensions; The generation network is mainly composed of 3D deconvolution blocks, and is used to reconstruct a high-dimensional image from the feature information obtained from the conversion layer, that is, the 3D volume image corresponding to the 2D image data projection.

4. The spine medical image registration method based on semantic reconstruction according to claim 1, characterized in that, Providing a 3D-3D feature extraction model, taking the 3D-CT image data as the input of the 3D-3D feature extraction model, and obtaining a feature map s corresponding to the 3D-CT image data with the same dimension as the 3D feature map, including: Constructing a 3D-3D feature extraction model using a 3D Res-NET network; Taking the 3D-CT image data and the manually segmented spine mask as the input of the 3D-3D feature extraction model. After downsampling the 3D-CT image data, under the action of convolution and deconvolution, extracting the key information of the spine in the 3D-CT image data and outputting to obtain the feature map s.

5. The method for spine medical image registration based on semantic reconstruction according to claim 1, wherein, Providing a parallel dual U-shaped network model based on cross-attention, and using the parallel dual U-shaped network model to fuse and register the feature map b and the feature map s, including: Combining the 3D feature map b and the feature map s with the 2D image data and the 3D-CT image data respectively to form a key region feature map; Performing window partitioning and window region partitioning on the key region feature map; Constructing a window-based multi-head cross-attention mechanism for calculating new features with corresponding correlation degrees between the input key region feature maps; Providing a parallel dual U-shaped network, taking the multi-head cross-attention mechanism as the convolutional layer of the parallel dual U-shaped network, and constructing a parallel dual U-shaped network model, which is used to output the registration result of the spine medical image.

6. The method for spine medical image registration based on semantic reconstruction according to claim 1, wherein, Optimizing the parallel dual U-shaped network model by using a multi-weight loss function based on pixel and semantic information, including: The loss function for measuring the parallel double-U network model is composed of a comprehensive combination of the inter-pixel loss in the image domain and the perceptual loss. Among them: the inter-pixel loss is used to measure the difference between the predicted image and the target image, and the perceptual loss is used to solve the problem of image over-smoothing caused by the inter-pixel loss. Then, a multi-weight loss function L based on pixel and semantic information is constructed total as follows: where: x is the original image, x * is the predicted image, Cov(·) is the covariance of two images, Var(·) is the variance of the image itself, μ is a sigmoid-like function; E(x,y) is, whc is, ρ is, is feature extraction.

7. The method for spine medical image registration based on semantic reconstruction according to any one of claims 1-6, wherein, further including: Processing the to-be-registered spine medical image by using the registration model, and outputting the registration result of the spine medical image.

8. A spine medical image registration system based on semantic reconstruction, wherein, including: A data processing module, which is used to construct spine 3D-CT image data and its corresponding 2D image data as a training data set; A 2D-3D reconstruction module, which is used to provide a 2D-3D reconstruction model, and using the 2D-3D reconstruction model to reconstruct the 2D image data to obtain a 3D feature map b corresponding to the 2D image data; A 3D-3D feature extraction module, which is used to provide a 3D-3D feature extraction model, taking the 3D-CT image data as the input of the 3D-3D feature extraction model, and obtaining a feature map s corresponding to the 3D-CT image data with the same dimension as the 3D feature map; A registration module, which is used to provide a parallel dual U-shaped network model based on cross-attention, and use the parallel dual U-shaped network model to fuse and register the feature map b and the feature map s; optimize the parallel dual U-shaped network model by using a multi-weight loss function based on pixel and semantic information, and is used to output the optimal registration result of the spinal medical image.

9. A computer terminal, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the computer program, it can be used to execute the method described in any one of claims 1-7, or run the system described in claim 8.

10. A computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by the processor, it can be used to execute the method described in any one of claims 1-7, or run the system described in claim 8.

Citation Information

Patent Citations

  • Two-dimensional image and CT or MR image three-dimensional fusion method

    CN106204511A

  • Two-dimensional and three-dimensional medical image registration method and system based on deep learning

    CN112150524A

  • DR and DRR image cross-modal automatic registration method in image-guided radiotherapy based on EPID

    CN112785632A

  • Anatomical scanning, targeting, and visualization

    US20230015717A1

Cited By

  • Data registration method based on parallel variable window convolutional neural network

    CN120411179A

  • Method and system for reconstructing three-dimensional CT from single two-dimensional X-ray fluoroscopic image using reinforcement learning

    CN121616744A

  • Method and system for reconstructing three-dimensional ct from a single two-dimensional x-ray perspective image using reinforcement learning

    CN121616744B