A Medical Image Registration Method Based on Mixture of Multiple Low-Rank Experts

Through the multi-low-rank expert hybrid architecture, the low-rank adaptation module is applied to the deep learning framework for medical image registration, which solves the problem of insufficient registration accuracy and robustness of deep learning methods in multi-modal, multi-position and pathological states, and realizes the optimization of parameter quantity and the improvement of model capabilities.

CN120182334BActive Publication Date: 2025-07-18ANHUI PROVINCIAL HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510649337.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-07-18
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The existing deep learning medical image registration methods have problems with insufficient registration accuracy and robustness when dealing with multimodal, multi-position, and anatomical structural variations in pathological states.

Method used

The hybrid architecture of multiple low-rank experts is adopted to apply the efficient parameter fine-tuning capability of low-rank adaptation to a deep learning framework for medical image registration. By splitting a single large expert network into multiple dynamically activated lightweight low-rank experts, combined with multi-special strategies, the modeling capability of the model is improved.

Benefits of technology

It improves the accuracy and robustness of medical image registration, reduces the number of parameters of the neural network, and optimizes the processing ability of anatomical structural variation characteristics in multimodal, multi-position and pathological states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182334B_ABST
    Figure CN120182334B_ABST
Patent Text Reader

Abstract

The present invention discloses a medical image registration method based on a mixture of multiple low-rank experts, belonging to the technical field of medical image processing. Aiming at the problem of poor robustness of medical image registration, the present invention applies the efficient parameter fine-tuning ability adapted to low rank to the deep learning framework for medical image registration, combines it with the multi-expert strategy that leads to a substantial increase in the number of parameters, and by splitting a single large expert network into multiple dynamically activated lightweight low-rank experts, it can not only prevent the training parameter quantity from being too high through low-rank constraints, but also improve the model's modeling ability in the case of multi-modal and multi-position of medical images through a larger number of experts, aiming to improve the accuracy and robustness of medical image registration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical image processing, and particularly relates to a medical image registration method based on a mixture of multiple low-rank experts. Background Art

[0002] Medical image registration is the core technology for constructing a digital surgical navigation system, realizing image-guided radiotherapy, and carrying out multi-modal image fusion diagnosis. It is often used in clinical scenarios such as tumor target delineation, interventional treatment path planning, and longitudinal lesion tracking. However, due to the characteristics of medical images such as multi-modal, multi-position, multi-posture acquisition deviation, and anatomical structure variation under pathological conditions, traditional feature-matching-based registration methods face challenges in clinical applications and cannot effectively handle the non-linear correspondence of cross-modal image intensity distributions and the initial alignment errors under large-scale posture offsets.

[0003] In recent years, deep learning registration methods represented by VoxelMorph have made breakthroughs through an end-to-end displacement field prediction framework. This type of method uses a 3D U-Net architecture with multi-resolution skip connections and achieves a registration speed close to real-time inference by minimizing a composite loss function of dissimilarity and displacement field smoothing constraints. However, although this architecture has a large number of trainable parameters, the final registration accuracy and robustness are still lower than those of traditional optimization-based registration methods because deep learning methods are limited by the training dataset and the generalization ability for medical images with characteristics such as multi-modal, multi-position, multi-posture acquisition deviation, and anatomical structure variation under pathological conditions still needs to be improved.

[0004] With the rapid development of technologies in the field of artificial intelligence, large models based on Transformer have become a research and application hotspot. Among them, Low-Rank Adaptation (LoRA) as an efficient fine-tuning strategy has triggered a revolutionary change in the field and has become a key strategy for large models such as DeepSeek-R1 to improve training efficiency, enabling cross-task knowledge transfer by fine-tuning only about 0.1% of the parameters. Although the low-rank adaptation strategy has demonstrated powerful capabilities in the field of large models based on Transformer, it is not only applicable to the Transformer architecture with high parameter redundancy. There is also low-rank update space in each layer of convolutional neural networks. However, in basic task fields outside large models, such as the field of medical image registration based on convolutional neural networks, its role has not been explored because it usually does not involve a large number of parameters. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a medical image registration method based on a mixture of multiple low-rank experts. Through a mixture of multiple low-rank experts architecture, the efficient parameter fine-tuning ability adapted to low rank is applied to the deep learning framework for medical image registration, and it is combined with the multi-expert strategy that leads to a significant increase in the number of parameters. By splitting a single large expert network into multiple dynamically activated lightweight low-rank experts, it can not only prevent the number of training parameters from being too high through low-rank constraints, but also improve the model's ability to model features such as multi-modal, multi-position, multi-posture acquisition deviation, and anatomical structure variation under pathological conditions of medical images, aiming to improve the accuracy and robustness of medical image registration.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A medical image registration method based on a mixture of multiple low-rank experts, comprising the following steps:

[0008] Step S1: Generate a cost tensor for the floating image and the fixed image through a specific dissimilarity measure and input it into a cascaded convolutional neural network;

[0009] Step S2: Add multiple low-rank adaptation modules to the side paths of the cascaded convolutional neural network, and each low-rank adaptation module serves as a low-rank expert model;

[0010] Step S3: Based on a dynamic routing mechanism, assign different weights to each low-rank expert model through a gating network and activate the top K low-rank expert models;

[0011] Step S4: Based on the top K low-rank expert models, adjust the parameters of the cascaded convolutional neural network, convert the input cost tensor into a mixed displacement field, and extrapolate it to a dense displacement field;

[0012] Step S5: Use the dense displacement field to perform a spatial transformation on the floating image and output the registration result.

[0013] In a second aspect, the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the foregoing medical image registration method based on a mixture of multiple low-rank experts.

[0014] In a third aspect, the present invention provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor can implement the foregoing medical image registration method based on a mixture of multiple low-rank experts.

[0015] The beneficial effects of the present invention are as follows:

[0016] 1) Apply the low-rank adaptation module to the cascaded convolutional neural network with a large number of parameters, solve the defect that the number of parameters of the convolutional neural network is insufficient to use the low-rank adaptation module, and improve the model modeling ability through multiple experts;

[0017] 2) For the medical image registration task, the expert module designed in the present invention does not act on only one neural network layer for each expert like other multi-expert networks, but acts on the entire cascaded convolutional neural network. The finally output displacement field is obtained by weighting the displacement fields generated by each expert with the weights output by the gating network. Each expert has an advantage for a certain type of feature, and this advantage is completely obtained by deep learning, so the robustness to various medical image modalities, positions, etc. can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flowchart of a medical image registration method with a mixture of multiple low-rank experts;

[0019] Figure 2 It is a detailed logic diagram of a medical image registration method with a mixture of multiple low-rank experts. DETAILED DESCRIPTION OF THE INVENTION

[0020] The present invention will be further described below in conjunction with the drawings and embodiments.

[0021] The present invention provides a medical image registration method with a mixture of multiple low-rank experts, as Figure 1 shown, which mainly includes the following steps:

[0022] Step S1, generate a cost tensor for the floating image and the fixed image through a specific dissimilarity measure, and input it into the cascaded convolutional neural network. As Figure 2 shown, the specific process is as follows:

[0023] 1) By calculating the specific dissimilarity between the registration source - floating image and the registration target - fixed image , a cost tensor is formed. The specific dissimilarity measure preferably adopts a weighted combination of normalized mutual information (NMI) and gradient difference: , where and are weighting coefficients, and is a gradient operator.

[0024] 2) Further, use convolution and transposed convolution in the cascaded convolutional neural network to construct a symmetric encoder - decoder structure. The encoder processes the input Perform multi-level downsampling to generate a resolution sequence and generate multi-resolution feature maps. The symmetric encoder-decoder structure performs feature fusion on the multi-resolution feature maps and the outputs of each layer of the decoder through skip connections. The multi-level downsampling includes 4 resolution levels, and the downsampling rates are [1, 2, 4, 8] respectively. Each level of the encoder contains two convolutional layers with convolutional kernel sizes The convolutional layers have numbers of channels [64, 128, 256, 512] respectively, and the stride = 2. The decoder correspondingly uses transposed convolution to implement upsampling. The preferred activation function is LeakyReLU, and the negative slope coefficient is set to 0.2.

[0025] Step S2, add multiple low-rank adaptation modules to the side path of the cascaded convolutional neural network, and each low-rank adaptation module serves as a low-rank expert model. As Figure 2 shown, the specific process is as follows:

[0026] Perform low-rank decomposition on the optimization space of the convolutional kernel parameters of each convolutional layer in the original cascaded convolutional neural network , is the number of output channels, is the number of input channels, is the size of the convolutional kernel, represents the real number space. Introduce intermediate auxiliary matrices A and B to form a low-rank adaptation module, and use the low-rank adaptation module to perform low-rank decomposition on the optimization space : , where , , the rank , the superscript T represents transpose, and obtain the convolutional kernel parameters in the cascaded convolutional neural network optimized by the low-rank adaptation module, that is, regard each low-rank adaptation module as a low-rank expert model for fine-tuning the cascaded convolutional neural network. Preferably, the rank of the low-rank decomposition is set to . The cascaded convolutional neural network contains a total of 12 convolutional layers. After each convolutional layer of the encoder, batch normalization and LeakyReLU activation functions are used. Among them, the convolutional operation maintains the spatial dimension of the feature map through padding = 1, and then realizes a halving of the resolution through downsampling with a stride = 2. Each level of the decoder realizes a doubling of the resolution through transposed convolution (stride = 2, padding = 1, output_padding = 1), and also uses a chain structure of batch normalization and LeakyReLU activation functions. The skip connection between the encoder and the decoder fuses features in a channel concatenation manner. The low-rank adaptation module fuses the parameter update amounts through element-wise addition .

[0027] Step S3: Based on the dynamic routing mechanism, assign different weights to each low-rank expert model through the gating network, and select the Top-K activated low-rank expert models. As Figure 2 shown, the specific process is as follows:

[0028] Implement the dynamic routing mechanism through the sparse gating network which is a neural network composed of several convolutional and fully connected layers. The input is a floating image or a fixed image, and the output is a vector composed of N values , where N is the number of low-rank expert models. Then, adopt the Top-K sparse activation strategy to generate the weight of the th expert: , where is the temperature coefficient, is the indicator function, is the operator for selecting the top largest elements, i.e., the Top-K elements, in the input vector. Preferably, the sparse gating network contains 3 convolutional layers, 1 pooling layer, and 2 fully connected layers. The parameters of the first convolutional layer are: 64 channels, = 7, stride = 2. The parameters of the second convolutional layer are: 128 channels, = 5, stride = 2. The parameters of the third convolutional layer are: 256 channels, = 3, stride = 2. The pooling layer uses global average pooling, and the dimensions of the fully connected layers are 1024, 512, and N in sequence. The temperature coefficient adopts an annealing strategy, with the initial value set to 1.0 and the decay coefficient of 0.95 per batch. Preferably, the Top-K parameter = 4. When the total number of experts N = 8, the sparse activation ratio remains 50%.

[0029] Step S4: Invoke the Top-K low-rank expert models to generate the displacement field, obtain the mixed displacement field through the weights of each expert, and then extrapolate it to the dense displacement field and optimize it. As Figure 2 shown, the specific process is as follows:

[0030] 1) Generate the Top-K displacement fields based on the low-rank expert models, that is, map the cost tensor to the displacement field through the convolutional kernel parameters in the cascaded convolutional neural network optimized by each of the top K large-weight low-rank expert models.

[0031] 2) Further, perform the synthesis of the displacement fields based on the weights to obtain the mixed displacement field , where , is the The displacement field generated by a cascaded convolutional neural network optimized by a low-rank expert model. Then, the mixed displacement field is extrapolated to a dense displacement field by B-spline interpolation: where : where denotes the p-th B-spline basis function corresponding to the -th node in the direction, denotes the p-th B-spline basis function corresponding to the -th node in the direction, denotes the p-th B-spline basis function corresponding to the -th node in the direction, and

[0032] Figure 2

[0033] 1) Apply the dense displacement field to the floating image to obtain the displaced coordinates of each voxel of the floating image.

[0034] 2) Further, resample according to the intensity value of each voxel and the displaced coordinates into a regular lattice constrained by the original resolution to obtain the registration result . The resampling is implemented using differentiable bilinear interpolation.

[0035] Finally, use the "original-expert-joint" three-step training method to train the parameters of the model. As Figure 2 shown, the specific process is as follows:

[0036] 1) Use at least a dissimilarity metric that quantifies the dissimilarity between the registered result and the fixed image to construct the loss function for training. The loss function is specifically constructed as: where and NCC represents the normalized cross-correlation coefficient, is the displacement field gradient smoothing term. Preferably, the balance coefficient = 0.3, the optimizer is AdamW, the initial learning rate is 3e-4, and the weight decay is 1e-4.

[0037] 2) Further, in the first batch of training, all low-rank expert models and the gating network are ignored, and only the parameters of the cascaded convolutional neural network are trained, that is, the parameters of the basic model are trained.

[0038] 3) Further, in the second batch of training, the parameters of the cascaded convolutional neural network obtained from the first batch of training are frozen, and only the parameters of the gating network and the low-rank expert model are trained, that is, the specific advantages of each expert are trained.

[0039] 4) Further, in the third batch of training, all parameters including the cascaded convolutional neural network, the gating network, and all low-rank expert models are trained.

[0040] Preferably, mixed-precision training and gradient clipping techniques are adopted, and the gradient norm threshold is set to 1.0. In the inference stage, the number of optional expert activations K can be dynamically adjusted according to computing resources. When the video memory is limited, K = 1 is set, and when accuracy is prioritized, K = 4 is set.

[0041] As can be seen from the technical solutions of the above embodiments, the medical image registration method of the present invention combines the efficient parameter fine-tuning ability of low-rank adaptation with a multi-expert strategy that leads to a large increase in the number of parameters. By splitting a single large expert network into multiple dynamically activated lightweight low-rank experts, it can not only prevent the number of training parameters from being too high through low-rank constraints, but also improve the model's ability to model features such as multi-modal, multi-position, multi-body position acquisition deviation, and anatomical structure variation in pathological states of medical images through a larger number of experts, and can improve the accuracy and robustness of medical image registration.

[0042] When N = 8, = 4, the neural network with multi-expert low-rank mixing proposed by the present invention reduces the number of neural network parameters by 84.8% compared with the case without adding a low-rank adaptation module. Unsupervised training and testing were carried out on the DIR-Lab COPD dataset, and the error of 300 landmark points was 1.82 mm, which is better than 2.54 mm of using only a single cascaded convolutional neural network and better than 2.16 mm of the classical deep learning medical image registration method VoxelMorph, proving that the method proposed by the present invention can obtain encouraging results only by optimizing on a common cascaded convolutional neural network.

[0043] In a second aspect, the present invention provides an electronic device, including: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the foregoing medical image registration method with multi-low-rank experts mixing.

[0044] In a third aspect, the present invention provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is enabled to implement the foregoing medical image registration method with multi-low-rank experts mixing.

[0045] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc., made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A medical image registration method based on a mixture of multiple low-rank experts, characterized in that, Including the following steps: Step S1: Generate a cost tensor for the floating image and the fixed image through a specific dissimilarity metric and input it into a cascaded convolutional neural network; Step S2: Add multiple low-rank adaptation modules to the side paths of the cascaded convolutional neural network, and each low-rank adaptation module serves as a low-rank expert model; Step S3: Based on the dynamic routing mechanism, assign different weights to each low-rank expert model through the gating network and activate the top K low-rank expert models; Step S4: Based on the top K low-rank expert models, adjust the parameters of the cascaded convolutional neural network, transform the input cost tensor into a mixed displacement field, and extrapolate it to a dense displacement field; Step S5: Use the dense displacement field to perform a spatial transformation on the floating image and output the registration result.

2. The method for medical image registration by mixing multiple low-rank experts according to claim 1, characterized in that, The said Step S1 includes: Calculate the specific dissimilarity between the floating image and the fixed image to form a cost tensor The cascaded convolutional neural network uses convolution and transposed convolution to construct a symmetric encoder-decoder structure. Through multiple layers of encoders, the cost tensor is subjected to multi-scale feature extraction to generate multi-resolution feature maps. The symmetric encoder-decoder structure performs feature fusion on the multi-resolution feature maps and the outputs of each layer of the decoder through skip connections.

3. The method for registering medical images by mixing multiple low-rank experts according to claim 2, wherein The said Step S2 includes: The convolution kernel parameters of each convolution layer in the original cascaded convolutional neural network Optimization space Perform low-rank decomposition to obtain the convolution kernel parameters in the optimized cascaded convolutional neural network , is the number of output channels, is the number of input channels, is the size of the convolution kernel, represents the real number space.

4. The medical image registration method based on multi-low-rank expert mixture according to claim 3, wherein Introduce intermediate auxiliary matrices A and B to form a low-rank adaptation module, and use the low-rank adaptation module to perform low-rank decomposition on the optimization space : where , , the rank , and the superscript T represents the transpose, obtaining the convolutional kernel parameters in the cascaded convolutional neural network optimized by the low-rank adaptation module , that is, each low-rank adaptation module is regarded as a low-rank expert model for fine-tuning the cascaded convolutional neural network 5. The medical image registration method based on a mixture of multiple low-rank experts according to claim 4, characterized in that The said Step S3 includes: Implement a dynamic routing mechanism through a sparse gating network, where the sparse gating network is a neural network composed of several convolutional and fully connected layers, with the input being a floating image or a fixed image and the output being a vector composed of N values , where N is the number of low-rank expert models. Then, a Top-K sparse activation strategy is adopted to generate the weights of the th low-rank expert model: , where is the temperature coefficient, is the indicator function, is the operator for selecting the top largest elements, i.e., the Top-K elements, in the input vector, and and represent the i-th and j-th elements of the vector , respectively.

6. The method for registering medical images by mixing multiple low-rank experts as claimed in claim 5, wherein The said Step S4 includes: Generate the top-K displacement fields based on the low-rank expert models, that is, the convolution kernel parameters in the cascaded convolutional neural network optimized by the top-K low-rank expert models Map the cost tensor to the displacement field; Perform weight-based displacement field synthesis to obtain a hybrid displacement field , where , is the displacement field generated by the cascaded convolutional neural network optimized by the th low-rank expert model; The mixed displacement field is extrapolated to a dense displacement field by using B-spline interpolation : where denotes the p-th order B-spline basis function corresponding to the -th node in the direction, denotes the p-th order B-spline basis function corresponding to the -th node in the direction, denotes the p-th order B-spline basis function corresponding to the -th node in the direction, and denotes the element of the mixed displacement field at ( ). Optimize the smoothness constraint of the dense displacement field through implicit neural representation.

7. The medical image registration method of multi-low-rank expert mixture according to claim 6, characterized in that, The said Step S5 includes: Apply the dense displacement field to the floating image to obtain the floating image The displaced coordinates of each voxel are resampled to the regular lattice constrained by the original resolution according to the intensity value of each voxel and the displaced coordinates, and the registration result is obtained .

8. A medical image registration method based on a mixture of multiple low-rank experts according to claim 7, characterized in that The said method further includes training the parameters of the model and the neural network using the "original-expert-joint" three-step training method: Using a dissimilarity metric of the dissimilarity between the registration result and the fixed image to form a loss function Perform self-supervised training in three batches. In the training of the first batch, all low-rank expert models and the gating network are ignored, and only the parameters of the cascaded convolutional neural network are trained. In the training of the second batch, the parameters of the cascaded convolutional neural network obtained from the training of the first batch are frozen, and only the parameters of the gating network and the low-rank expert models are trained. In the training of the third batch, all parameters including the cascaded convolutional neural network, the gating network, and all low-rank expert models are trained.

9. An electronic device, characterized in that, Including: One or more processors; A memory for storing one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement a medical image registration method with multi-low-rank expert mixing according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, Stored thereon are executable instructions which, when executed by a processor, can enable the processor to implement a medical image registration method with multi-low-rank expert mixing according to any one of claims 1-8.

Citation Information

Patent Citations

  • Abdomen multi-organ registration method based on adaptive multi-gating hybrid expert model

    CN116993793A

  • Face forgery detection method and system based on reconstruction learning and hybrid expert mode

    CN119625812A