Multi-low-rank expert mixed medical image registration method
By adopting a hybrid architecture of multiple low-rank expert in medical image registration, combining low-rank adaptation module and dynamic routing mechanism, the alignment error problem of traditional methods when processing medical images is solved, and the accuracy and robustness of registration are improved.
Patent Information
- Application Number
- CN202510649337.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-20
AI Technical Summary
Traditional medical image registration methods have initial alignment errors when dealing with nonlinear correspondence of cross-modal image intensity distribution and large-scale position offset, and the registration accuracy and robustness of deep learning methods are also low.
Using a hybrid architecture of multiple low-rank experts, the low-rank adaptation module is applied to the cascaded convolutional neural network. The weight is assigned to each low-rank expert model through a dynamic routing mechanism and a gated network, the first K low-rank expert models are activated, and the mixed displacement field is generated and extrapolated to the dense displacement field.
It improves the accuracy and robustness of medical image registration, and can more effectively deal with multimodal, multi-position, and anatomical structural variations in medical images in more efficient manner.
Smart Images

Figure CN120182334A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical image processing, and particularly relates to a medical image registration method based on the mixture of multiple low-rank experts. Background Art
[0002] Medical image registration is the core technology for constructing a digital surgical navigation system, realizing image-guided radiotherapy, and performing multi-modal image fusion diagnosis, and is commonly used in clinical scenarios such as tumor target delineation, interventional treatment path planning, and longitudinal lesion tracking. However, due to the characteristics of multi-modal, multi-position, multi-posture acquisition deviation of medical images, and anatomical structure variation under pathological conditions, traditional feature-matching-based registration methods face challenges in clinical applications and cannot effectively handle the non-linear correspondence of cross-modal image intensity distributions and the initial alignment errors under large-scale posture offsets.
[0003] In recent years, deep learning registration methods represented by VoxelMorph have made breakthrough progress through an end-to-end displacement field prediction framework. This type of method uses a 3D U-Net architecture with multi-resolution skip connections and achieves a registration speed close to real-time inference by minimizing a composite loss function of dissimilarity and displacement field smoothing constraints. However, although this architecture has a large number of trainable parameters, the final registration accuracy and robustness are still lower than those of traditional optimization-based registration methods because deep learning methods are limited by the training dataset and the generalization ability for medical images with characteristics such as multi-modal, multi-position, multi-posture acquisition deviation, and anatomical structure variation under pathological conditions still needs to be improved.
[0004] With the rapid development of technologies in the field of artificial intelligence, large models based on Transformer have become a research and application hotspot. Among them, Low-Rank Adaptation (LoRA) as an efficient fine-tuning strategy has triggered a revolutionary change in the field and has become a key strategy for large models such as DeepSeek-R1 to improve training efficiency, enabling cross-task knowledge transfer by only fine-tuning about 0.1% of the parameters. Although the low-rank adaptation strategy has demonstrated powerful capabilities in the field of large models based on Transformer, it is not only applicable to the Transformer architecture with high parameter redundancy, and there is also low-rank update space in each layer of the convolutional neural network. However, in basic task fields outside large models, such as the field of medical image registration based on convolutional neural networks, its role has not been explored because it usually does not involve a large number of parameters. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a medical image registration method based on a mixture of multiple low-rank experts. Through a mixture of multiple low-rank experts architecture, the efficient parameter fine-tuning ability adapted to low rank is applied to the deep learning framework for medical image registration, and it is combined with the multi-expert strategy that leads to a substantial increase in the number of parameters. By splitting a single large expert network into multiple dynamically activated lightweight low-rank experts, it can not only prevent the training parameter quantity from being too high through low-rank constraints, but also improve the model's ability to model features such as multi-modal, multi-position, multi-body position acquisition deviation, and anatomical structure variation under pathological conditions of medical images, aiming to improve the accuracy and robustness of medical image registration.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A medical image registration method based on a mixture of multiple low-rank experts, comprising the following steps:
[0008] Step S1: Generate a cost tensor for the floating image and the fixed image through a specific dissimilarity measure, and input it into a cascaded convolutional neural network;
[0009] Step S2: Add multiple low-rank adaptation modules to the side paths of the cascaded convolutional neural network, and each low-rank adaptation module serves as a low-rank expert model;
[0010] Step S3: Based on a dynamic routing mechanism, assign different weights to each low-rank expert model through a gating network, and activate the top K low-rank expert models;
[0011] Step S4: Adjust the parameters of the cascaded convolutional neural network based on the top K low-rank expert models, convert the input cost tensor into a mixed displacement field, and extrapolate it to a dense displacement field;
[0012] Step S5: Use the dense displacement field to perform a spatial transformation on the floating image and output the registration result.
[0013] In a second aspect, the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned medical image registration method based on a mixture of multiple low-rank experts.
[0014] In a third aspect, the present invention provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor can implement the aforementioned medical image registration method based on a mixture of multiple low-rank experts.
[0015] The beneficial effects of the present invention are as follows:
[0016] 1) Apply the low-rank adaptation module to the cascaded convolutional neural network with a large number of parameters, solve the defect that the number of parameters in the convolutional neural network is insufficient to use the low-rank adaptation module, and improve the model modeling ability through multiple experts;
[0017] 2) For the medical image registration task, the expert module designed in the present invention does not act on only one neural network layer like other multi-expert networks, but acts on the entire cascaded convolutional neural network. The finally output displacement field is obtained by weighting the displacement fields generated by each expert with the weights output by the gating network. Each expert has an advantage for a certain type of feature, and this advantage is completely obtained by deep learning. Therefore, the robustness to various medical image modalities, positions, etc. can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flowchart of a medical image registration method with a mixture of multiple low-rank experts;
[0019] Figure 2 It is a detailed logic diagram of a medical image registration method with a mixture of multiple low-rank experts. DETAILED DESCRIPTION OF THE INVENTION
[0020] The present invention will be further described below with reference to the drawings and embodiments.
[0021] The present invention provides a medical image registration method with a mixture of multiple low-rank experts, as Figure 1 shown, which mainly includes the following steps:
[0022] Step S1, generate a cost tensor for the floating image and the fixed image through a specific dissimilarity measure, and input it into the cascaded convolutional neural network. As Figure 2 shown, the specific process is as follows:
[0023] 1) By calculating the specific dissimilarity between the registration source - floating image and the registration target - fixed image , a cost tensor is formed. The specific dissimilarity measure preferably adopts a weighted combination of normalized mutual information (NMI) and gradient difference: , where and are weighting coefficients, is the gradient operator.
[0024] 2) Further, use convolution and transposed convolution in the cascaded convolutional neural network to construct a symmetric encoder-decoder structure. The encoder processes the input Perform multi-level downsampling to generate a resolution sequence and generate multi-resolution feature maps. The symmetric encoder-decoder structure performs feature fusion on the multi-resolution feature maps and the outputs of each layer of the decoder through skip connections. The multi-level downsampling includes 4 resolution levels, and the downsampling rates are [1, 2, 4, 8] respectively. Each level of the encoder contains two convolutional layers with convolutional kernel sizes The convolutional layers have numbers of channels [64, 128, 256, 512] respectively, and the stride = 2. The decoder correspondingly uses transposed convolution to implement upsampling. The preferred activation function is LeakyReLU, and the negative slope coefficient is set to 0.2.
[0025] Step S2, add multiple low-rank adaptation modules to the side path of the cascaded convolutional neural network, and each low-rank adaptation module serves as a low-rank expert model. As Figure 2 shown, the specific process is as follows:
[0026] Perform low-rank decomposition on the optimization space of the convolutional kernel parameters of each convolutional layer in the original cascaded convolutional neural network where is the number of output channels, is the number of input channels, is the size of the convolutional kernel, represents the real number space, introduce intermediate auxiliary matrices A and B to form a low-rank adaptation module, and use the low-rank adaptation module to perform low-rank decomposition on the optimization space : where , The rank , and the superscript T represents transpose, obtaining the convolutional kernel parameters in the cascaded convolutional neural network optimized by the low-rank adaptation module, that is, each low-rank adaptation module serves as a low-rank expert model for fine-tuning the cascaded convolutional neural network. Preferably, the rank of the low-rank decomposition is set to . The cascaded convolutional neural network contains a total of 12 convolutional layers. After each convolutional layer in the encoder, batch normalization and LeakyReLU activation functions are used. Among them, the convolutional operation maintains the spatial dimension of the feature map through padding = 1, and then reduces the resolution by half through downsampling with stride = 2. Each level of the decoder doubles the resolution through transposed convolution (stride = 2, padding = 1, output_padding = 1), and also uses batch normalization and LeakyReLU activation function chain structures. The skip connection between the encoder and the decoder fuses features in a channel concatenation manner. The low-rank adaptation module fuses the parameter update amounts by element-wise addition.
[0027] Step S3: Based on the dynamic routing mechanism, assign different weights to each low-rank expert model through the gating network, and select the Top-K activated low-rank expert models. As Figure 2 shown, the specific process is as follows:
[0028] Implement the dynamic routing mechanism through the sparse gating network which is a neural network composed of several convolutional and fully connected layers. The input is a floating image or a fixed image, and the output is a vector composed of N values , where N is the number of low-rank expert models. Then, adopt the Top-K sparse activation strategy to generate the weights of the th expert: , where is the temperature coefficient, is the indicator function, is the operator for selecting the top largest (i.e., Top-K) elements in the input vector. Preferably, the sparse gating network contains 3 convolutional layers, 1 pooling layer, and 2 fully connected layers. The parameters of the first convolutional layer are: 64 channels, = 7, stride = 2. The parameters of the second convolutional layer are: 128 channels, = 5, stride = 2. The parameters of the third convolutional layer are: 256 channels, = 3, stride = 2. The pooling layer uses global average pooling, and the dimensions of the fully connected layers are 1024, 512, and N in sequence. The temperature coefficient adopts an annealing strategy, with the initial value set to 1.0 and the decay coefficient of 0.95 per batch. Preferably, the Top-K parameter = 4. When the total number of experts N = 8, the sparse activation ratio remains 50%.
[0029] Step S4: Call the Top-K low-rank expert models to generate the displacement field, obtain the mixed displacement field through the weights of each expert, and then extrapolate it to the dense displacement field and optimize it. As Figure 2 shown, the specific process is as follows:
[0030] 1) Generate Top-K displacement fields based on the low-rank expert models, that is, map the cost tensor to the displacement field through the convolutional kernel parameters in the cascaded convolutional neural network optimized by each of the top K large-weight low-rank expert models.
[0031] 2) Further, perform the synthesis of the displacement fields based on the weights to obtain the mixed displacement field , where , is the The displacement field generated by the cascaded convolutional neural network optimized by a low-rank expert model. Then, the hybrid displacement field is extrapolated to a dense displacement field using B-spline interpolation to : where denotes the p-th B-spline basis function corresponding to the -th node in the direction, denotes the p-th B-spline basis function corresponding to the -th node in the direction, denotes the p-th B-spline basis function corresponding to the -th node in the direction, and
[0032] Figure 2
[0033] 1) Apply the dense displacement field to the floating image to obtain the displaced coordinates of each voxel of the floating image.
[0034] 2) Further, resample according to the intensity value of each voxel and the displaced coordinates into the regular lattice constrained by the original resolution to obtain the registration result . The resampling is implemented using differentiable bilinear interpolation.
[0035] Finally, the parameters of the model are trained using the "original-expert-joint" three-step training method. As Figure 2 shown, the specific process is as follows:
[0036] 1) Construct the loss function for training using at least a dissimilarity metric term that quantifies the dissimilarity between the registered result and the fixed image . The loss function is specifically constructed as: where , NCC represents the normalized cross-correlation coefficient, is the displacement field gradient smoothing term. Preferably, the balance coefficient = 0.3, the optimizer is AdamW, the initial learning rate is 3e-4, and the weight decay is 1e-4.
[0037] 2) Further, in the first batch of training, all low-rank expert models and the gating network are ignored, and only the parameters of the cascaded convolutional neural network are trained, that is, the parameters of the basic model are trained.
[0038] 3) Further, in the second batch of training, the parameters of the cascaded convolutional neural network obtained from the first batch of training are frozen, and only the parameters of the gating network and the low-rank expert models are trained, that is, the specific advantages of each expert are trained.
[0039] 4) Further, in the third batch of training, all parameters including the cascaded convolutional neural network, the gating network, and all low-rank expert models are trained.
[0040] Preferably, mixed-precision training and gradient clipping techniques are adopted, and the gradient norm threshold is set to 1.0. In the inference stage, the number of optional expert activations K can be dynamically adjusted according to computing resources. When the video memory is limited, K = 1 is set, and when accuracy is prioritized, K = 4 is set.
[0041] As can be seen from the technical solutions of the above embodiments, the medical image registration method of the present invention combines the efficient parameter fine-tuning ability of low-rank adaptation with the multi-expert strategy that leads to a large increase in the number of parameters. By splitting a single large expert network into multiple dynamically activated lightweight low-rank experts, it can not only prevent the number of training parameters from being too high through low-rank constraints, but also improve the model's ability to model features such as multi-modal, multi-position, multi-body position acquisition deviation, and anatomical structure variation under pathological conditions of medical images through a larger number of experts, and can improve the accuracy and robustness of medical image registration.
[0042] When N = 8, = 4, the neural network with multi-expert low-rank mixing proposed by the present invention reduces the number of neural network parameters by 84.8% compared with the case without adding a low-rank adaptation module. Unsupervised training and testing were carried out on the DIR-Lab COPD dataset, and the error of 300 landmark points was 1.82 mm, which is better than 2.54 mm of using only a single cascaded convolutional neural network and better than 2.16 mm of the classic deep learning medical image registration method VoxelMorph, proving that the method proposed by the present invention can obtain encouraging results only by optimizing on a common cascaded convolutional neural network.
[0043] In a second aspect, the present invention provides an electronic device, including: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the foregoing medical image registration method with multi-low-rank expert mixing.
[0044] In a third aspect, the present invention provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor can be caused to implement the foregoing medical image registration method with multi-low-rank expert mixing.
[0045] The specific embodiments described above further elaborate on the object, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A medical image registration method based on a mixture of multiple low-rank experts, characterized in that: The following steps are involved: Step S1: Generate a cost tensor for the floating image and the fixed image through a specific dissimilarity metric and input it into the cascade convolutional neural network; Step S2: adding multiple low-rank adaptation modules to the cascade convolutional neural network side path, each low-rank adaptation module serves as a low-rank expert model; Step S3: Based on the dynamic routing mechanism, different weights are assigned to each low-rank expert model through the gating network, and the first K low-rank expert models are activated; Step S4: adjusting the parameters of the cascaded convolutional neural network based on the first K low-rank expert models, converting the input cost tensor into a mixed displacement field, and extrapolating it to a dense displacement field; Step S5: Use the dense displacement field to perform spatial transformation on the floating image and output the registration result.
2. The medical image registration method of multiple low-rank expert mixture as claimed in claim 1, characterized in that: The step S1 comprises: Calculate floating image and fixed images The specific dissimilarity between them constitutes the cost tensor The cascaded convolutional neural network uses convolution and transposed convolution to build a symmetrical encoder-decoder structure, and uses multiple layers of encoders to decode the cost tensor. Multi-scale feature extraction is performed to generate a multi-resolution feature map, and the symmetrical encoder-decoder structure performs feature fusion on the multi-resolution feature map and the output of each layer of the decoder through jump connections.
3. The medical image registration method of multiple low-rank expert mixtures as claimed in claim 1, characterized in that: The step S2 comprises: The convolution kernel parameters of each convolution layer in the original cascade convolutional neural network Optimization space Perform low-rank decomposition to obtain the optimized convolution kernel parameters in the cascade convolutional neural network , is the number of output channels, is the number of input channels, is the size of the convolution kernel, Represents the real number space.
4. The medical image registration method of multiple low-rank expert mixture as claimed in claim 3, characterized in that: The intermediate auxiliary matrices A and B are introduced to form a low-rank adaptation module, and the low-rank adaptation module is used to optimize the space Perform low-rank decomposition: ,in , ,rank , the superscript T represents transposition, and the convolution kernel parameters in the cascade convolutional neural network after low-rank adaptation module optimization are obtained. , that is, each low-rank adaptation module is used as a low-rank expert model to fine-tune the cascaded convolutional neural network.
5. The medical image registration method of multiple low-rank expert mixture as claimed in claim 1, characterized in that: The step S3 comprises: Through sparse gating network Implementing a dynamic routing mechanism, the sparse gating network It is a neural network composed of several convolutional and fully connected layers. The input is a floating image or a fixed image, and the output is a vector composed of N numerical values. , N is the number of low-rank expert models, and then the Top-K sparse activation strategy is used to generate the The weights of the low-rank expert models are: ,in is the temperature coefficient, is the indicator function, To select the first The operator of the top-K elements. and Represents vectors The i-th and j-th elements of .
6. The medical image registration method of multiple low-rank expert mixture as claimed in claim 5, characterized in that: The step S4 comprises: Generate the top-K displacement fields based on low-rank expert models, that is, the convolution kernel parameters in the cascaded convolutional neural network optimized by the first K low-rank expert models The cost tensor Mapped into a displacement field; Perform weight-based displacement field synthesis to obtain a mixed displacement field ,in , It is The displacement field generated by a cascaded convolutional neural network optimized by a low-rank expert model; B-spline interpolation is used to convert the mixed displacement field Extrapolation to dense displacement field : ,in express Direction The p-order B-spline basis function corresponding to the nodes, express Direction The p-order B-spline basis function corresponding to the nodes, express Direction The p-order B-spline basis function corresponding to the nodes, Represents the mixed displacement field exist( ) at the element; Optimizing smoothness constraints of dense displacement fields via implicit neural representation.
7. The medical image registration method of multiple low-rank expert mixtures as claimed in claim 1, characterized in that: The step S5 comprises: The dense displacement field Acting on floating images , get the floating image The displaced coordinates of each voxel are resampled to a regular lattice constrained by the original resolution according to the intensity value of each voxel and the displaced coordinates to obtain the registration result. .
8. The medical image registration method of multiple low-rank expert mixtures as claimed in claim 1, characterized in that: The method also includes training the parameters of the model and the neural network using a three-step training method of "original-expert-joint": Use the dissimilarity measure between the registration result and the fixed image The loss function Three batches of self-supervised training are performed. In the first batch of training, all low-rank expert models and gating networks are ignored, and only the parameters of the cascaded convolutional neural network are trained. In the second batch of training, the parameters of the cascaded convolutional neural network obtained by the first batch of training are frozen, and only the parameters of the gating network and the low-rank expert model are trained. In the third batch of training, all parameters including the cascaded convolutional neural network, the gating network and all low-rank expert models are trained.
9. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs; Wherein, when one or more programs are executed by the one or more processors, the one or more processors implement the medical image registration method of multiple low-rank expert mixtures as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that: Executable instructions are stored thereon, and when the instructions are executed by a processor, the processor can implement a medical image registration method of a mixture of multiple low-rank experts as described in any one of claims 1-8.
Citation Information
Patent Citations
Abdomen multi-organ registration method based on adaptive multi-gating hybrid expert model
CN116993793A
Face forgery detection method and system based on reconstruction learning and hybrid expert mode
CN119625812A
Method and system for analysing images of a retina
US20210319556A1
Image registration method, apparatus and device, and storage medium
WO2023207266A1