Unsupervised multi-modal medical image registration method and system based on anatomical structure perception

By adopting an unsupervised deep learning method based on anatomical structure perception in multimodal medical image registration, the problems of high time overhead and label dependence in the prior art are solved, and efficient and accurate multimodal medical image registration is achieved, which is suitable for clinical applications.

CN120070521APending Publication Date: 2025-05-30SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510145558.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing multimodal medical image registration methods have the shortcomings of high time overhead and label dependence. Especially in 3D image processing, labeling segmentation labels requires a lot of manpower and material resources, and there are also problems of artifacts, fuzzy or unreal anatomical structure.

Method used

An unsupervised multimodal medical image registration method based on anatomical structure perception is proposed, using an end-to-end deep learning model, using the consistency of different modal anatomical structures to guide the registration process, reduce the information interference of modal differences through the anatomical structure perception module, and pay attention to edge information through the edge registration task to avoid dependence on anatomical structure segmentation labels.

Benefits of technology

It significantly improves the speed and accuracy of image registration, reduces the registration time to 0.42s, enhances the interpretability of the registration network, improves the ability to capture important anatomical details in CT and MRI images, and has strong versatility and clinical application potential.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070521A_ABST
    Figure CN120070521A_ABST
Patent Text Reader

Abstract

The invention provides an unsupervised multi-modal medical image registration method and system based on anatomical structure perception, and the method comprises the steps: S1, constructing a registration model, training the registration model through an unsupervised loss function, and obtaining a trained registration model; s2, registering the CT image by using the trained registration model according to the MRI image to obtain a deformed CT image aligned with the MRI; the registration model can realize registration of a CT image and an MRI image through an end-to-end deep learning network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image registration, and specifically, to an unsupervised multimodal medical image registration method and system based on anatomical structure perception. Background Art

[0002] In modern medical diagnosis, the application of multimodal medical imaging technology is becoming increasingly important. Different imaging modalities provide different aspects of human anatomical structure and physiological function information, providing rich and complementary data for clinical diagnosis. However, there are spatial position inconsistencies between multimodal medical images, such as patient position changes, equipment imaging angle differences, time differences, etc., making it difficult to directly compare the characteristic manifestations of the same anatomical structure. Therefore, accurate multimodal medical image registration has important clinical significance.

[0003] Image registration refers to mapping a moving image to a fixed image through a spatial transformation, so that the two images are as consistent as possible under a certain similarity metric. In other words, the optimization of the registration process can be expressed as:

[0004]

[0005] where L sim (·) represents a similarity function, which is used to measure the alignment degree between the deformed moving image M(φ) and the fixed image, and L smooth (·) represents a smoothness constraint function, which is used to ensure the smoothness of the spatial deformation field.

[0006] Traditional image registration involves complex non-linear optimization processes and large-scale deformation field solutions, with high time complexity and large computational overhead, making it difficult to meet the requirements of clinical real-time processing. Moreover, traditional methods rely on parameter settings, and the tuning process is cumbersome and lacks theoretical guidance. In recent years, the feasibility of deep learning methods has been verified, which can greatly shorten the registration time, and their performance has been proven to be comparable to traditional methods. However, when existing deep learning methods are directly used for multimodal registration, there are the following defects: (1) Existing multimodal registration methods mostly adopt weak supervision methods and rely on segmentation labels for supervision. However, the cost of obtaining labels is expensive, especially for 3D images. Compared with 2D images, their spatial structure is more complex, and annotating segmentation labels requires a lot of manpower and material resources. (2) Although the method of using a generative adversarial network can transform multimodal registration into a single-modal registration task to reduce the registration difficulty, its accuracy is limited by the performance of the synthesis network, and there may be artifacts, blurring, or unrealistic anatomical structures. For medical clinical applications, false details will bring potential risks.

[0007] Considering that the differences between different modalities mainly exist in texture, while the spatial positions and shapes of anatomical structures change relatively little, and anatomical structures have high stability, and the anatomical structures of the same individual and different individuals have high spatial similarity, anatomical structure information can serve as a reliable reference in multimodal registration.

[0008] Based on the limitations of existing multimodal registration algorithms, the present invention proposes an unsupervised multimodal medical image registration method based on anatomical structure perception. This method is an end-to-end deep learning model that innovatively uses the consistency of anatomical structures in different modalities to guide the registration process, proposes an anatomical structure perception module to reduce the information interference of modality differences, and ensures the alignment of anatomical structures; and by introducing an edge registration task, guides the network to focus on edge information. In addition, an unsupervised learning paradigm is adopted to avoid dependence on anatomical structure segmentation labels, and has stronger generality and clinical application potential. Based on such a strategy, we can obtain CT and MRI image pairs with high-precision registration. Summary of the Invention

[0009] Aiming at the defects in the prior art, the purpose of the present invention is to provide an unsupervised multimodal medical image registration method based on anatomical structure perception.

[0010] An unsupervised multimodal medical image registration method based on anatomical structure perception provided by the present invention includes:

[0011] Step S1: Construct a registration model and train the registration model through an unsupervised loss function to obtain a trained registration model;

[0012] Step S2: The CT image is registered according to the MRI image by using the trained registration model to obtain a deformed CT image aligned with the MRI;

[0013] The registration model can realize the registration of CT images and MRI images through an end-to-end deep learning network.

[0014] Preferably, the registration model is an end-to-end deep learning network including a feature encoder, an anatomical structure perception module, and a composite feature fusion module;

[0015] The feature encoder adopts a pyramid structure with two-stream shared weights to extract multi-scale feature pyramids F f and F m from the fixed image I f and the moving image I m respectively; among them, the CT image is used as the moving image I m ; the MR image is used as the fixed image I m ;

[0016] The anatomical structure perception module extracts the multi-scale feature pyramid F by implicitly embedding Sobel and Laplacian filters into the convolutional layer f and F m of the first-order and second-order differential features of the moving image and the first-order and second-order differential features of the fixed image;

[0017] The composite feature fusion module is used to fuse the features of the two modalities of the moving image and the fixed image.

[0018] Preferably, the feature encoder includes: a pyramid structure with two-stream shared weights; among them, the encoder in the pyramid structure with two-stream shared weights includes four cascaded convolutional modules;

[0019] Among them, each convolutional module includes: first, an average pooling operation is performed through the downsampling module to reduce the spatial resolution, then two cascaded 3×3×3 convolutional layers are used for feature extraction, and then through the normalization layer and the LeakyReLU activation function layer, the multi-scale feature pyramid of the image is output;

[0020] The number of feature channels of the downsampling modules in the four convolutional modules starts from the initial value c and increases to 2c, 4c, and 8c in each downsampling module in turn.

[0021] Preferably, the anatomical structure perception module includes: an ES module and an EL module;

[0022] The multi-scale feature pyramid F f and F m extract the first-order differential features and second-order differential features of the CT image and the first-order differential features and second-order differential features of the MRI image through the ES module and the EL module respectively;

[0023] Among them, the ES module includes four parallel branches; one standard 3×3×3 convolutional branch is used to extract basic features; three branches based on the Sobel operator respectively simulate the Sobel-x, Sobel-y, and Sobel-z operators to capture the edge gradient information in three directions; the first-order differential features are generated based on the extracted basic features and edge gradient information;

[0024] The EL module includes two parallel branches, one of which is a 3×3×3 standard convolutional branch for extracting basic features; one convolutional branch based on the Laplacian operator for obtaining enhanced second-order structural features; the second-order differential features are generated based on the extracted basic features and enhanced second-order structural features.

[0025] Preferably, the composite feature fusion module includes: three parallel branches; the three parallel branches include: two channel attention branches and one convolutional feature extraction branch;

[0026] The channel attention branch enhances the shared structural features of the two modalities through channel attention;

[0027] The enhancing of the shared structural features of the two modalities through channel attention includes: using average pooling and max pooling operations to extract the global statistical information of the features, adding the extracted features and then passing through a 1×1×1 convolutional layer and a Sigmoid activation function in sequence to generate an attention weight map in the channel dimension; then multiplying it with the original features for feature weighting to capture the anatomical structure features therein;

[0028] The convolutional feature extraction branch processes the input features through convolutional operations to capture the global feature representation of the dual-path input, and then performs feature weighting through multiplication operations;

[0029] Finally, the results of the three branch paths are fused to obtain the final features, and a 3D convolutional layer with a convolutional kernel size of 3×3×3 is used to obtain the final deformation field Δφ.

[0030] Preferably, the unsupervised loss function includes:

[0031]

[0032] Among them, f and m are the moving CT image and the original fixed MR image respectively, and m°φ is to superimpose the spatial deformation field φ on the moving CT image; is the similarity loss, which is used to measure the similarity between the moving image and the deformed image, and mutual information is used to optimize the similarity; represents the smoothness loss function, which is used to constrain the smoothness of the deformation field, and the L 2 norm is used for constraint; is the edge loss function, which is used to measure the similarity between the edges of the moving image and the deformed image, and is constrained by modality-independent neighborhood description; the calculation method of the image edge is to calculate the spatial gradients in the x, y, and z directions respectively ( and ); perform convolution operations on each direction using a 3D Sobel operator to obtain the corresponding gradient maps x sx 、x sy and x sz ; finally, the gradient maps in the three directions are weighted and fused with the same weight , and the specific formula is It is used to explicitly extract edge information as a constraint for image registration; λ and γ respectively represent the weights of the smoothness loss function and the edge loss function; the above three loss functions are jointly optimized during the training process, enabling the network to adaptively learn the spatial transformation between images in an unsupervised manner and effectively perform accurate registration of multimodal images; in the performance evaluation, the spatial deformation field is applied to the segmentation label of the fixed image to obtain the deformed segmentation label, and the Dice score is calculated with the segmentation label of the moving image to evaluate the registration performance, and the deformation field with the highest Dice score is retained.

[0033] An unsupervised multimodal medical image registration system based on anatomical structure perception provided by the present invention includes:

[0034] Module M1: Construct a registration model and train the registration model through an unsupervised loss function to obtain a trained registration model;

[0035] Module M2: The CT image is registered according to the MRI image using the trained registration model to obtain a deformed CT image aligned with the MRI;

[0036] The registration model can achieve the registration of CT images and MRI images through an end-to-end deep learning network.

[0037] Preferably, the registration model is an end-to-end deep learning network including a feature encoder, an anatomical structure perception module, and a composite feature fusion module;

[0038] The feature encoder adopts a pyramid structure with two-stream shared weights to extract multi-scale feature pyramids F f from the fixed image I m and the moving image I f respectively; among them, the CT image is used as the moving image I m ; the MR image is used as the fixed image I m ; m ;

[0039] The anatomical structure perception module extracts the first-order and second-order differential features of the moving image and the first-order and second-order differential features of the fixed image of the multi-scale feature pyramids F f and F m by implicitly embedding Sobel and Laplacian filters into the convolutional layer;

[0040] The composite feature fusion module is used to fuse the features of the two modalities of the moving image and the fixed image.

[0041] Preferably, the feature encoder includes: a pyramid structure with two-stream shared weights; among them, the encoder in the feature encoder with a pyramid structure with two-stream shared weights includes four cascaded convolutional modules;

[0042] Among them, each convolutional module includes: first, perform average pooling operation through the downsampling module to reduce the spatial resolution, then use two cascaded 3×3×3 convolutional layers for feature extraction, and then pass through the normalization layer and the LeakyReLU activation function layer to output the multi-scale feature pyramid of the image;

[0043] The number of feature channels of the downsampling module in the four convolutional modules starts from the initial value c and increases to 2c, 4c, and 8c in each downsampling module in sequence;

[0044] The anatomical structure perception module includes: an ES module and an EL module;

[0045] Multi-scale feature pyramid F f and F m Extract the first-order differential features and second-order differential features of the CT image and the first-order differential features and second-order differential features of the MRI image through the ES module and the EL module respectively;

[0046] Among them, the ES module includes four parallel branches; one standard 3×3×3 convolutional branch for extracting basic features; three branches based on the Sobel operator, respectively simulating the Sobel-x, Sobel-y, and Sobel-z operators for capturing the edge gradient information in three directions; generate the first-order differential features based on the extracted basic features and edge gradient information;

[0047] The EL module includes two parallel branches, one 3×3×3 standard convolutional branch for extracting basic features; one convolutional branch based on the Laplacian operator for obtaining enhanced second-order structural features; generate the second-order differential features based on the extracted basic features and enhanced second-order structural features;

[0048] The composite feature fusion module includes: three parallel branches; the three parallel branches include: two channel attention branches and one convolutional feature extraction branch;

[0049] The channel attention branch strengthens the shared structural features of the two modalities through channel attention;

[0050] The strengthening of the shared structural features of the two modalities through channel attention includes: using average pooling and max pooling operations to extract the global statistical information of the features, adding the extracted features and then passing through a 1×1×1 convolutional layer and a Sigmoid activation function in sequence to generate an attention weight map in the channel dimension; then multiply it with the original features for feature weighting to capture the anatomical structure features therein;

[0051] The convolutional feature extraction branch processes the input features through convolutional operations, captures the global feature representations of the dual-path input, and then performs feature weighting through multiplication operations;

[0052] Finally, the results of the three branch paths are fused to obtain the final features, and a 3D convolutional layer with a convolutional kernel size of 3×3×3 is used to obtain the final deformation field Δφ.

[0053] Preferably, the unsupervised loss function includes:

[0054]

[0055] where f and m are the moving CT image and the original fixed MR image respectively, is to superimpose the spatial deformation field φ onto the moving CT image; is the similarity loss, which is used to measure the similarity between the moving image and the deformed image, and mutual information is used to optimize the similarity; represents the smoothness loss function, which is used to constrain the smoothness of the deformation field, and the L 2 norm is used for constraint; is the edge loss function, which is used to measure the similarity between the edges of the moving image and the deformed image, and is constrained by the modality-independent neighborhood description; the edges of the image are calculated by calculating the spatial gradients in the x, y, and z directions respectively ( and ); the 3D Sobel operator is used for convolution operations in each direction to obtain the corresponding gradient maps x sx , x sy and x sz ; finally, the gradient maps in the three directions are weighted and fused with the same weight , and the specific formula is is used to explicitly extract edge information as a constraint for image registration; λ and γ respectively represent the weights of the smoothness loss function and the edge loss function; the above three loss functions are jointly optimized during training, and the network adaptively learns the spatial transformation between images in an unsupervised manner and effectively performs accurate registration of multimodal images; in performance evaluation, the spatial deformation field is applied to the segmentation label of the fixed image to obtain the deformed segmentation label, and the Dice score is calculated with the segmentation label of the moving image to evaluate the registration performance, and the deformation field with the highest Dice score is retained.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] 1. The present invention effectively solves the problem of high time consumption of traditional non-rigid registration methods. Compared with the average registration time of 190 s of traditional methods, the present invention reduces the time to 0.42 s, significantly improving the speed and accuracy of image registration, providing an efficient and reliable solution for clinical applications, being suitable for quickly processing a large number of medical images, and being able to support real-time or near-real-time clinical decisions.

[0058] 2. The network structure of the present invention is designed based on the mathematical models of the Sobel operator and the Laplacian operator, rather than relying entirely on data-driven methods; this design enhances the interpretability of the registration network and can effectively capture important anatomical structure details in CT and MRI images; this innovation can improve the registration accuracy, especially when dealing with modalities with complex anatomical structures and large morphological differences, significantly enhancing the accuracy of registration.

[0059] 3. The network proposed by the present invention adopts unsupervised learning technology, does not rely on any label information, avoids the labor cost and time consumption of manual annotation; uses the inherent properties of CT and MRI images as supervision signals, can perform registration adaptively, is suitable for CT and MRI registration tasks of all organs and parts, and has strong generality; does not need to design features specifically for different organs and can be applied to various organs and lesion sites throughout the body.

[0060] 4. The present invention adopts a multi-modal iterative network, combined with a layer-by-layer refinement optimization strategy, to improve the accuracy and robustness of cross-modal registration; through multi-stage iterative optimization, it can effectively overcome the differences between different modal images, especially in the case of complex anatomical structure details, improving the registration accuracy and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] By reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings, other features, objectives, and advantages of the present invention will become more apparent:

[0062] Figure 1 It is a flowchart for training an image inference model.

[0063] Figure 2 It is a schematic diagram of an anatomical structure perception module.

[0064] Figure 3 It is a schematic diagram of a composite feature fusion module. DETAILED DESCRIPTION OF THE INVENTION

[0065] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all fall within the protection scope of the present invention.

[0066] Embodiment 1

[0067] According to the unsupervised multimodal medical image registration method based on anatomical structure perception provided by the present invention, as Figures 1 to 3 shown, it includes:

[0068] Step S1: Construct a registration model, and train the registration model through an unsupervised loss function to obtain a trained registration model;

[0069] Step S2: Use the trained registration model to register the CT image according to the MRI image to obtain a deformed CT image aligned with the MRI;

[0070] The registration model is an end-to-end deep learning network model, including two processes: model training and image inference. During the model training process, the input is the CT and MR images of the same patient, with CT as the moving image and MR as the fixed image, and the output is the spatial deformation field. In image inference, the input is the CT image and the spatial deformation field obtained in the model training stage. The model applies the deformation field obtained in the training stage to the CT image features and gradually optimizes the deformation, thereby generating a deformed CT image aligned with the fixed image.

[0071] The registration model includes three modules: a feature encoder, an anatomical structure perception module, and a composite feature fusion module, as Figure 1 ; the model extracts the features of the fixed image and the moving image through the feature encoder, constructs a feature pyramid, and then gradually fuses and optimizes the features through the anatomical structure perception module and the composite feature fusion module to generate the spatial deformation field at each level. Among them, the anatomical structure perception module strengthens the structural information of each of the CT and MRI modalities, and the composite feature fusion module takes the strengthened features of the two as inputs, enhances the shared structural information through cross-modal interaction, and obtains the spatial deformation field at each level. The overall network adopts an iterative method, and at each level, the currently predicted spatial deformation field is applied to the moving image features for deformation, and the deformed features and the fixed image features are input together to the next level for decoding. This iterative structure ensures the optimization of the spatial deformation field from coarse to fine.

[0072] Among them, the feature encoder adopts a pyramid structure with dual-stream shared weights, respectively from the fixed image I f and the moving image Im Extract multi-scale features. The encoder is composed of four cascaded convolutional modules. Each convolutional module adopts the same structural design. First, the spatial resolution is gradually reduced through average pooling operations, and then two cascaded 3×3×3 convolutional layers are used for feature extraction, followed by a normalization layer and a LeakyReLU activation function layer. The number of feature channels starts from the initial value c and is sequentially increased to 2c, 4c, and 8c in each downsampling module, obtaining the multi-scale feature pyramids F f and F m . Through this pyramid structure, the network gradually extracts multi-scale features, while retaining the spatial details of the shallow layers and enhancing the representation ability of the deep semantic information.

[0073] The Anatomical Structure Perception Module (ASPM) extracts first-order and second-order differential features by implicitly embedding Sobel and Laplacian filters into convolutional layers, thereby enhancing the expression ability of structural features in the feature map, as Figure 2 shown. Its main modules are the ES module and the EL module. The ES module contains four parallel branches: a standard 3×3×3 convolutional branch for extracting basic features, and three branches based on the Sobel operator, which respectively simulate the Sobel-x, Sobel-y, and Sobel-z operators for capturing edge gradient information in three directions. Among them, the Sobel operator simulates the first-order partial derivative operations in the x, y, and z directions in three-dimensional space through the convolutional kernel weights; thus capturing the edge gradient information in three orthogonal directions and generating first-order differential features. The EL module contains two parallel branches: a 3×3×3 standard convolutional branch and a convolutional branch based on the Laplacian operator for enhancing second-order structural features. The Laplacian operator simulates the second-order partial derivative operations in three-dimensional space through the convolutional kernel weights; used to enhance the second-order structural features of the gray-level mutations in the image.

[0074] The Compound Feature Fusion Module (CFFM) is used to fuse the features of two modalities, CT and MRI. By assisting each other with the features of the two modalities, the representation ability of each modality is enhanced. After the features of the fixed image and the moving image pass through the feature encoder and the ASPM module, features with enhanced anatomical structure information are obtained. CFFM fuses the bimodal features through the following three branches, captures their similarities and differences, and realizes the generation of the spatial deformation field. The CFFM module contains three parallel branches: two channel attention branches and one convolutional feature extraction branch. The channel attention branches strengthen the shared structural features of the two modalities through channel attention, use average pooling (AVG) and max pooling (MAP) operations to extract the global statistical information of the features, add the extracted features, and then pass through a 1×1×1 convolutional layer and a Sigmoid activation function in sequence to generate an attention weight map in the channel dimension. Then it is multiplied by the original features for feature weighting to capture the anatomical structure features therein. The convolutional feature extraction branch processes the input features through convolutional operations, effectively captures the global feature representation of the dual-channel input, and then performs feature weighting through multiplication operations to highlight the differential information between the two-channel features and enhance the discriminative ability of the features. Finally, the results of the three branch paths are fused to obtain the final features, and a three-dimensional convolutional layer with a convolutional kernel size of 3×3×3 is used to obtain the final spatial deformation field Δφ. This structural design enables the network to adaptively focus on the structural information and achieve efficient feature selection and fusion.

[0075] In the model training stage, the present invention adopts an unsupervised learning method for training, avoiding the dependence on anatomical structure segmentation labels. The overall loss function is:

[0076]

[0077] where f and m are the moving CT image and the original fixed MR image respectively, and m°φ is to superimpose the spatial deformation field φ on the moving CT image. is the similarity loss, which is used to measure the similarity between the moving image and the deformed image, and mutual information (MI) is used to optimize the similarity. represents the smoothness loss function, which is used to constrain the smoothness of the deformation field, and the L 2 norm is used for constraint. is the edge loss function, which is used to measure the similarity between the edges of the moving image and the deformed image, and is constrained by modality-independent neighborhood description (MIND). The calculation method of the image edge is to calculate the spatial gradients in the x, y, and z directions respectively ( and ). The 3D Sobel operator is used for convolution operation in each direction to obtain the corresponding gradient maps x sx , x sy and xsz Finally, the gradient maps in three directions are weighted and fused with the same weights The specific formula is It is used to explicitly extract edge information as a constraint for image registration. λ and γ respectively represent the weights of the smoothness loss function and the edge loss function. The above three loss functions are jointly optimized during the training process, enabling the network to adaptively learn the spatial transformation between images in an unsupervised manner and effectively perform accurate registration of multimodal images. In the performance evaluation, the spatial deformation field is applied to the segmentation label of the fixed image to obtain the deformed segmentation label, and the Dice score is calculated with the segmentation label of the moving image to evaluate the registration performance. The deformation field with the highest Dice score is retained for subsequent inference.

[0078] In the image inference stage, taking the registration of CT to MRI images of the same patient as an example, with the CT as the moving image, applying the optimal spatial deformation field obtained in the model training stage to the CT image can obtain the deformed CT image. Doctors can view the deformed CT image and the patient's existing MRI image to observe the location, shape, size and other characteristics of the lesion or tumor from the same perspective, so as to achieve more accurate lesion analysis and treatment plan formulation. Through the registered images, doctors can intuitively compare the anatomical structures and lesion areas in different modality images, helping clinical staff better understand the spatial distribution of diseases, improving the diagnostic accuracy and formulating personalized treatment plans. Especially in the case where CT and MRI information needs to be combined to comprehensively evaluate the condition, it significantly improves the efficiency and accuracy of medical image analysis.

[0079] The present invention also provides an unsupervised multimodal medical image registration system based on anatomical structure perception. The unsupervised multimodal medical image registration system based on anatomical structure perception can be implemented by executing the process steps of the unsupervised multimodal medical image registration method based on anatomical structure perception. That is, those skilled in the art can understand the unsupervised multimodal medical image registration method based on anatomical structure perception as a preferred implementation manner of the unsupervised multimodal medical image registration system based on anatomical structure perception.

[0080] Those skilled in the art know that in addition to implementing the systems, devices, and their respective modules provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the systems, devices, and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to implement the same program. Therefore, the systems, devices, and their respective modules provided by the present invention can be considered as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware component; the modules for implementing various functions can also be regarded as either software programs for implementing the method or the structures within the hardware component.

[0081] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. An unsupervised multimodal medical image registration method based on anatomical structure perception, characterized in that: include: Step S1: construct a registration model, and train the registration model through an unsupervised loss function to obtain a trained registration model; Step S2: the CT image is registered with the MRI image using the trained registration model to obtain a deformed CT image aligned with the MRI; The registration model can realize the registration of CT images and MRI images through an end-to-end deep learning network.

2. The unsupervised multimodal medical image registration method based on anatomical structure perception according to claim 1, characterized in that: The registration model is an end-to-end deep learning network including a feature encoder, an anatomical structure perception module and a composite feature fusion module; The feature encoder adopts a two-stream shared weight pyramid structure to respectively extract the fixed image I f and moving image I m Extract a multi-scale feature pyramid F from f and F m ; Among them, the CT image is used as the moving image I m ; MR image as fixed image I m ; The anatomical structure perception module extracts a multi-scale feature pyramid F by implicitly embedding Sobel and Laplacian filters into the convolutional layer. f and F m The first-order and second-order differential features of moving images and the first-order and second-order differential features of fixed images; The composite feature fusion module is used to fuse the features of the two modalities of moving image and fixed image.

3. The unsupervised multimodal medical image registration method based on anatomical structure perception according to claim 2, characterized in that: The feature encoder comprises: a pyramid structure using dual streams sharing weights; wherein the encoder in the feature encoder using the pyramid structure using dual streams sharing weights comprises four convolution modules cascaded; Each convolution module includes: firstly, the spatial resolution is reduced by average pooling through the downsampling module, and then two 3×3×3 convolution layers are used in series to extract features, and then the multi-scale feature pyramid of the image is output through the normalization layer and the LeakyReLU activation function layer; The number of feature channels of the downsampling modules in the four convolutional modules starts from the initial value c and increases to 2c, 4c, and 8c in each downsampling module.

4. The unsupervised multimodal medical image registration method based on anatomical structure perception according to claim 2, characterized in that: The anatomical structure perception module includes: an ES module and an EL module; Multi-scale feature pyramid F f and F m The first-order differential features and second-order differential features of CT images and the first-order differential features and second-order differential features of MRI images are extracted through the ES module and the EL module respectively; The ES module includes four parallel branches; one of which is a standard 3×3×3 convolution branch for extracting basic features; three branches based on Sobel operators, which respectively simulate Sobel-x, Sobel-y and Sobel-z operators to capture edge gradient information in three directions; and first-order differential features are generated based on the extracted basic features and edge gradient information. The EL module includes two parallel branches, a 3×3×3 standard convolution branch for extracting basic features, and a convolution branch based on a Laplace operator for obtaining enhanced second-order structural features; and second-order differential features are generated based on the extracted basic features and the enhanced second-order structural features.

5. The unsupervised multimodal medical image registration method based on anatomical structure perception according to claim 2, characterized in that: The composite feature fusion module includes: three parallel branches; the three parallel branches include: two channel attention branches and one convolution feature extraction branch; The channel attention branch strengthens the shared structural features of the two modalities through channel attention; The method of strengthening the shared structural features of the two modalities through channel attention includes: extracting global statistical information of features using average pooling and maximum pooling operations, adding the extracted features and sequentially passing through a 1×1×1 convolutional layer and a Sigmoid activation function to generate an attention weight map of the channel dimension; then multiplying the map with the original features for feature weighting to capture the anatomical structural features therein; The convolutional feature extraction branch processes input features through convolution operations to capture the global feature representation of dual-path inputs, and then performs feature weighting through multiplication operations; Finally, the results of the three branch paths are fused to obtain the final features, and a three-dimensional convolution layer with a convolution kernel size of 3×3×3 is used to obtain the final deformation field Δφ.

6. The unsupervised multimodal medical image registration method based on anatomical structure perception according to claim 1, characterized in that: The unsupervised loss function includes: Among them, f and m are the moving CT image and the original fixed MR image, respectively. To superimpose the spatial deformation field φ onto the moving CT image; is the similarity loss, which is used to measure the similarity between the moved image and the deformed image, and uses mutual information to optimize the similarity; Represents the smoothness loss function, which is used to constrain the smoothness of the deformation field and uses the L2 norm to constrain; is an edge loss function, which is used to measure the similarity between the edges of moving images and deformed images, and is constrained by using a modality-independent neighborhood description. The image edge is calculated by calculating the spatial gradients in the x, y, and z directions respectively ( and ) ; Use the 3D Sobel operator to perform convolution operation on each direction to obtain the corresponding gradient map x sx 、x sy and x sz ; Finally, the gradient maps in the three directions are weighted equally Perform weighted fusion, the specific formula is: It is used to explicitly extract edge information as a constraint for image registration; λ and γ represent the weights of the smoothness loss function and the edge loss function, respectively; the above three loss functions are jointly optimized during the training process, and the network adaptively learns the spatial transformation between images in an unsupervised manner, and effectively performs accurate registration of multimodal images; in the performance evaluation, the spatial deformation field is applied to the segmentation label of the fixed image to obtain the deformed segmentation label, and the Dice score is calculated with the segmentation label of the moving image to evaluate the registration performance, and the deformation field with the highest Dice score is retained.

7. An unsupervised multimodal medical image registration system based on anatomical structure perception, characterized in that: include: Module M1: Construct a registration model and train the registration model through an unsupervised loss function to obtain a trained registration model; Module M2: The CT image is registered with the MRI image using the trained registration model to obtain a deformed CT image aligned with the MRI. The registration model can realize the registration of CT images and MRI images through an end-to-end deep learning network.

8. The unsupervised multimodal medical image registration system based on anatomical structure perception according to claim 7, characterized in that: The registration model is an end-to-end deep learning network including a feature encoder, an anatomical structure perception module and a composite feature fusion module; The feature encoder adopts a two-stream shared weight pyramid structure to respectively extract the fixed image I f and moving image I m Extract a multi-scale feature pyramid F from f and F m ; Among them, the CT image is used as the moving image I m ; MR image as fixed image I m ; The anatomical structure perception module extracts a multi-scale feature pyramid F by implicitly embedding Sobel and Laplacian filters into the convolutional layer. f and F m The first-order and second-order differential features of moving images and the first-order and second-order differential features of fixed images; The composite feature fusion module is used to fuse the features of the two modalities of moving image and fixed image.

9. The unsupervised multimodal medical image registration method based on anatomical structure perception according to claim 8, characterized in that: The feature encoder comprises: a pyramid structure using dual streams sharing weights; wherein the encoder in the feature encoder using the pyramid structure using dual streams sharing weights comprises four convolution modules cascaded; Each convolution module includes: firstly, the spatial resolution is reduced by average pooling through the downsampling module, and then two 3×3×3 convolution layers are used in series to extract features, and then the multi-scale feature pyramid of the image is output through the normalization layer and the LeakyReLU activation function layer; The number of feature channels of the downsampling modules in the four convolutional modules starts from the initial value c and increases to 2c, 4c, and 8c in each downsampling module; The anatomical structure perception module includes: an ES module and an EL module; Multi-scale feature pyramid F f and F m The first-order differential features and second-order differential features of CT images and the first-order differential features and second-order differential features of MRI images are extracted through the ES module and the EL module respectively; The ES module includes four parallel branches; one of which is a standard 3×3×3 convolution branch for extracting basic features; three branches based on Sobel operators, which respectively simulate Sobel-x, Sobel-y and Sobel-z operators to capture edge gradient information in three directions; and first-order differential features are generated based on the extracted basic features and edge gradient information. The EL module includes two parallel branches, one 3×3×3 standard convolution branch for extracting basic features; one convolution branch based on the Laplacian operator for obtaining enhanced second-order structural features; and second-order differential features are generated based on the extracted basic features and the enhanced second-order structural features. The composite feature fusion module includes: three parallel branches; the three parallel branches include: two channel attention branches and one convolution feature extraction branch; The channel attention branch strengthens the shared structural features of the two modalities through channel attention; The method of strengthening the shared structural features of the two modalities through channel attention includes: extracting global statistical information of features using average pooling and maximum pooling operations, adding the extracted features and sequentially passing through a 1×1×1 convolutional layer and a Sigmoid activation function to generate an attention weight map of the channel dimension; then multiplying the map with the original features for feature weighting to capture the anatomical structural features therein; The convolutional feature extraction branch processes input features through convolution operations to capture the global feature representation of dual-path inputs, and then performs feature weighting through multiplication operations; Finally, the results of the three branch paths are fused to obtain the final features, and a three-dimensional convolution layer with a convolution kernel size of 3×3×3 is used to obtain the final deformation field Δφ.

10. The unsupervised multimodal medical image registration system based on anatomical structure perception according to claim 7, characterized in that: The unsupervised loss function includes: Among them, f and m are the moving CT image and the original fixed MR image, respectively. To superimpose the spatial deformation field φ onto the moving CT image; is the similarity loss, which is used to measure the similarity between the moved image and the deformed image, and uses mutual information to optimize the similarity; Represents the smoothness loss function, which is used to constrain the smoothness of the deformation field and uses the L2 norm to constrain; is an edge loss function, which is used to measure the similarity between the edges of moving images and deformed images, and is constrained by using a modality-independent neighborhood description. The image edge is calculated by calculating the spatial gradients in the x, y, and z directions respectively ( and ) ; Use the 3D Sobel operator to perform convolution operation on each direction to obtain the corresponding gradient map x sx 、x sy and x sz ; Finally, the gradient maps in the three directions are weighted equally Perform weighted fusion, the specific formula is: It is used to explicitly extract edge information as a constraint for image registration; λ and γ represent the weights of the smoothness loss function and the edge loss function, respectively; the above three loss functions are jointly optimized during the training process, and the network adaptively learns the spatial transformation between images in an unsupervised manner, and effectively performs accurate registration of multimodal images; in the performance evaluation, the spatial deformation field is applied to the segmentation label of the fixed image to obtain the deformed segmentation label, and the Dice score is calculated with the segmentation label of the moving image to evaluate the registration performance, and the deformation field with the highest Dice score is retained.

Citation Information

Cited By

  • Medical image registration method based on large model robust features

    CN120782830A

  • Medical image registration method based on self-supervised representation learning and hierarchical feature optimization

    CN120997263A