Multi-modal image registration processing method and system based on deep learning

By generating optimized sCT images through deep learning and combining dynamic features and weight selection, and using CNN and UNet models to generate global and local deformation fields, the problems of nonlinear accuracy and automated optimization in multimodal image registration are solved, achieving high-precision image registration and treatment plan optimization, which is suitable for radiotherapy in complex anatomical regions.

CN120894404APending Publication Date: 2025-11-04HANGZHOU NORMAL UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511124612.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing multimodal image registration techniques lack accuracy in handling complex nonlinear deformations and dynamic feature fusion, and lack automated optimization processes, making it difficult to meet the needs of high-precision clinical image processing.

Method used

This study utilizes a deep learning-based approach, employing a 2DcGAN model to generate optimized sCT images. By combining dynamic features and weight selection, a multi-scale feature set is generated. Global and local deformation fields are generated using CNN and UNet models. Registration parameters are optimized through mutual information loss and total loss to generate high-precision registered sCT images. The registration quality is verified by combining automatic segmentation and keypoint analysis, and a clinical image processing plan is generated.

Benefits of technology

It significantly improves the spatial alignment accuracy between sCT and reference CT and the reliability of clinical image processing, and is particularly suitable for radiotherapy planning in complex anatomical regions, achieving automated, high-precision image registration and treatment plan optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894404A_ABST
    Figure CN120894404A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal image registration processing method and system based on deep learning, and relates to the technical field of medical image processing, and the method comprises the steps: calculating a mutual information loss value after registration through a mutual information loss method, optimizing a CNN model and UNet model parameters in combination with the total loss generated by the fusion of global and local deformation fields, and obtaining a registered sCT image. A CNN model and a UNet model are combined to generate global and local deformation fields, overall rigid transformation and local nonlinear deformation are effectively captured, the spatial alignment precision of sCT and reference CT is improved, model parameters are optimized through mutual information loss and total loss, the intensity distribution consistency is ensured, feature weights and treatment plan parameters are automatically adjusted through registration quality feedback, and the accuracy of the treatment plan is improved. The coverage precision and efficiency of the radiotherapy plan are remarkably improved, automatic and high-precision image registration and treatment plan optimization are realized, and the reliability and practicability of clinical image processing are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical image processing, in particular to a multi-modal image registration processing method and system based on deep learning. BACKGROUND

[0002] In recent years, multi-modal medical image registration technology has made significant progress in the fields of radiotherapy, diagnostic imaging and surgical navigation. In recent years, the rise of deep learning technology has driven the development of registration technology, especially the application of generative adversarial networks (GAN) and convolutional neural networks (CNN) in image generation and transformation. However, the existing technology still has shortcomings in multi-modal image registration. Traditional rigid registration methods are difficult to handle complex nonlinear deformation, such as local distortion of tumors or organs, resulting in limited registration accuracy. Secondly, although sCT generation based on GAN can simulate CT intensity distribution, it lacks a dynamic feature fusion mechanism and is difficult to adaptively balance global and local deformation. Existing methods usually use fixed weights or single-scale features, ignoring the dynamic adjustment needs of high-level semantics (such as organ shape) and low-level texture (such as edge details), resulting in insufficient alignment accuracy of registered images in complex anatomical regions. In addition, the integration of registration quality assessment (such as DSC, TRE, MAE) and clinical treatment planning is insufficient, lacking an automated optimization process, limiting the operability of clinical applications. For example, current radiotherapy planning relies on manual parameter adjustment and fails to fully utilize registration quality feedback to optimize model parameters, affecting the accuracy and efficiency of treatment area positioning. The existing registration technology has limitations in handling complex nonlinear deformation and dynamic feature fusion, making it difficult to meet the high-precision clinical image processing needs. The present application generates dynamic features based on optimized sCT images and reference CT, combines dynamic weight selection to generate multi-scale feature sets, uses CNN and UNet models to generate global and local deformation fields, and optimizes registration parameters through mutual information loss and total loss to generate high-precision registered sCT images. Finally, based on the registration quality, a clinical image processing plan is generated, solving the shortcomings of existing technology in nonlinear registration accuracy and automated treatment plan optimization. SUMMARY

[0003] In view of the above existing problems, the present application is proposed.

[0004] Therefore, the present application provides a multi-modal image registration processing method and system based on deep learning, solving the problems of existing technology in nonlinear registration accuracy and automated treatment plan optimization.

[0005] To solve the above technical problems, the present application provides the following technical solutions: In a first aspect, the present application provides a deep learning-based multi-modal image registration processing method, which comprises collecting MRI and reference CT images, respectively preprocessing, deploying a 2DcGAN model to convert the MRI image into an sCT image, and outputting an optimized sCT image through a CREPs perceptual loss function; Based on the optimized sCT image and the reference CT, dynamic features are generated, and a multi-scale feature set is generated based on the dynamic features and dynamic weights; The multi-scale feature set is used to generate global and local deformation fields based on a CNN model and a UNet model, the mutual information loss value after registration is calculated through a mutual information loss method, the total loss generated by combining the global and local deformation fields is fused, the CNN model and UNet model parameters are optimized, and the registered sCT image is obtained; The alignment accuracy of the registered sCT image and the reference CT is evaluated, the registration quality is verified through automatic segmentation and key point analysis, and a clinical image processing plan is generated based on the registration quality result.

[0006] As a preferred scheme of the deep learning-based multi-modal image registration processing method of the present application, wherein: the collection of MRI and reference CT images, respectively preprocessing refers to collecting daily MRI and reference CT images, automatically generating a tumor mask by FSL, correcting non-uniformity of the MRI image using N4ITK algorithm, and applying Perona-Malik algorithm to remove noise, and performing intensity clipping and normalization on the reference CT image.

[0007] As a preferred scheme of the deep learning-based multi-modal image registration processing method of the present application, wherein: the deployment of a 2DcGAN model to convert the MRI image into an sCT image, and outputting an optimized sCT image through a CREPs perceptual loss function refers to deploying a 2DcGAN model, including a generator G, mapping a 2D MRI slice to an sCT slice, a discriminator D being a PatchGAN, judging the authenticity of the input image, and outputting a probability value, inputting a standardized MRI image into the 2DcGAN model, cutting into 128 2D slices according to the axis, and recombining the 128 2DsCT slices to output an sCT image; Optimizing the generator and the discriminator through a CREPs perceptual loss method to make the sCT image close to the CT image, extracting style features based on a ConvNext-Tiny model, including: using a linear algebra standard method to calculate the Gram matrix difference of the jth layer features, constructing a style loss , calculating an adversarial loss , calculating a total perceptual loss Updating the convolution layer weights and biases of the generator and the discriminator, calculating the KL divergence through the entropy function of SciPy, and if > threshold U, then , otherwise set .

[0008] As a preferred scheme of the multi-modal image registration processing method based on deep learning provided in the application, wherein: the dynamic features are generated based on the optimized sCT image and the reference CT, and the multi-scale feature set is generated based on the dynamic features and the dynamic weights, which uses ConvNext-Tiny to extract features from the optimized sCT and the reference CT respectively: the first layer: low-level texture features, the third layer: middle-level edge features, and the fifth layer: high-level semantic features; Calculate the tumor volume difference, if the difference is greater than the threshold , the weight distribution is from high weight to low weight, and the order is the fifth layer, the third layer, and the first layer, if the difference is less than the threshold , the weight distribution order is the first layer, the third layer, and the fifth layer, and the weight is obtained by weighted voting method, on the validation set, test different weight combinations, according to the peak signal-to-noise ratio of the generated sCT and the reference CT to evaluate the image quality, and select the weight combination with the highest PSNR as the optimal weight; ConvNext-Tiny directly extracts the first, third and fifth layer features from sCT and CT , to generate a multi-scale feature set .

[0009] As a preferred scheme of the multi-modal image registration processing method based on deep learning provided in the application, wherein: the multi-scale feature set is used to generate global and local deformation fields based on a CNN model and a UNet model, the mutual information loss value after registration is calculated by a mutual information loss method, and the total loss generated by combining the global and local deformation fields is fused to optimize the CNN model and UNet model parameters, and the registered sCT image is obtained ; The sCT and the reference CT image and the multi-scale feature set are composed of a channel splicing set , which is input into the UNet model , the UNet model fuses the global sCT, CT and multi-scale features by forward propagation, extracts spatial features, and generates a local deformation field , and the final deformation field is obtained by fusing the global and local deformation fields ; The transformation coordinates of the sCT voxels are calculated using a linear algebra standard method , the resampling formula parameters are calculated, and the registered sCT intensity is obtained ; Based on the registered Intensity and reference CT mutual information, registration is optimized using mutual information loss, maximizing MI, mutual information loss , local deformation field Regularization is performed to obtain a regularization loss , the total loss is calculated , based on the total loss Update CNN and UNet convolution layer weights and biases to generate the final deformation field , and output the registered sCT' image.

[0010] As a preferred scheme of the multi-modal image registration processing method based on deep learning according to the application, wherein: the evaluation of the alignment accuracy of the registered sCT' and the reference CT uses a fixed number of MRI / CT data verification sets, and the registration quality is verified by automatic segmentation and key point analysis. The FSL tool is used to automatically segment the registered sCT' and the reference CT to generate binary organ masks of white matter, gray matter, ventricle and tumor, and the mask overlap degree is calculated: the number of intersecting voxels of sCT' and CT masks is counted, and then divided by twice the sum of the number of voxels of the two masks to obtain the Dice similarity coefficient. The FLIRT tool of FSL is used to automatically extract a fixed number of key anatomical points of sCT' and CT, the Euclidean distance between the points is calculated to obtain the TRE, and the intensity values of sCT' and CT are compared voxel by voxel to calculate the average value of the absolute difference of each voxel to obtain the mean absolute error. If the DSC is less than the threshold Q, adjust the feature weight, output the updated weight distribution, update the ConvNext-Tiny 5th layer convolution parameters, regenerate the multi-scale feature set, and output the adjusted weight distribution , and input The UNet is updated to optimize the registration.

[0011] As a preferred scheme of the multi-modal image registration processing method based on deep learning according to the application, wherein: based on the registration quality result, a clinical image processing plan is generated, which is based on the registered image sCT' and the global DDF, and the feature weight , evaluation index, using the Eclipse system combined with the tumor and organ masks segmented by FSL to generate a treatment region positioning plan, calculating the region coverage distribution, optimizing the positioning parameters, if the coverage deviation is greater than the threshold T, adjusting the ConvNext-Tiny 5th layer convolution weight according to the DSC through grid search, regenerating sCT' and , output the optimized positioning plan and updated model parameters for clinical verification.

[0012] In a second aspect, the present application provides a deep learning-based multi-modal image registration processing system, comprising a data collection and preprocessing module for collecting daily MRI and reference CT images and outputting standardized MRI and CT images; an sCT generation module for deploying a 2DcGAN model, optimizing a generator and a discriminator through a CREPs perception loss, and generating high-quality sCT images; a dynamic feature extraction and multi-scale feature set generation module for extracting multi-layer features from the optimized sCT and the reference CT using a pre-trained ConvNext-Tiny and generating a multi-scale feature set; a deformation field generation and registration module for generating global and local deformation fields by CNN and UNet respectively, fusing them, optimizing mutual information loss and smoothing regularization loss, and generating registered images; a registration quality evaluation module for calculating DSC, extracting key points through FSLFLIRT, calculating TRE, and calculating MAE by voxel-by-voxel intensity comparison.

[0013] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, any step of the deep learning-based multi-modal image registration processing method according to the first aspect of the present application is implemented.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, any step of the deep learning-based multi-modal image registration processing method according to the first aspect of the present application is implemented.

[0015] The present application has the following advantages: through dynamic feature generation and dynamic weight selection of multi-scale feature sets, global and local deformation fields are generated by combining CNN and UNet models, effectively capturing overall rigid transformation and local nonlinear deformation, improving the spatial alignment accuracy of sCT and reference CT, optimizing model parameters through mutual information loss and total loss, ensuring intensity distribution consistency, and automatically adjusting feature weights and treatment planning parameters using registration quality feedback, significantly improving the coverage accuracy and efficiency of radiotherapy planning, compared with traditional manual adjustment or fixed weight method, realizing automatic, high-precision image registration and treatment planning optimization, enhancing the reliability and practicality of clinical image processing, and being particularly suitable for radiotherapy planning in complex anatomical regions. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.

[0017] Fig. 1 Flowchart of the deep learning-based multi-modal image registration processing method in embodiment 1.

[0018] Fig. 2 Structure diagram of the deep learning-based multi-modal image registration processing system in embodiment 1.

[0019] Fig. 3 Registration quality evaluation module view of the deep learning-based multi-modal image registration processing method in embodiment 1. DETAILED DESCRIPTION

[0020] In order to make the above-mentioned objects, features and advantages of the present application more apparent and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0021] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the scope of the present application, therefore the present application is not limited to the specific embodiments disclosed below.

[0022] Secondly, the "one embodiment" or "embodiment" referred to herein means that the specific features, structures or characteristics can be included in at least one implementation of the present application. "In one embodiment" appearing in different places in the specification does not mean the same embodiment, nor is it an independent or alternative embodiment.

[0023] Embodiment 1, refer to Figs. 1-3 , the first embodiment of the present application, the embodiment provides a deep learning-based multi-modal image registration processing method, comprising the following steps: S1, collect MRI and reference CT images, respectively, pre-process, deploy 2DcGAN model to convert MRI image into sCT image, and output optimized sCT image through CREPs perception loss function; Specifically, the MRI and reference CT images are collected and pre-processed. The daily MRI and reference CT images are collected, the tumor mask (indicating the position of the tumor area in the image (1 indicates the tumor area, and 0 indicates the non-tumor area)) is automatically generated by FSL (Jenkinson et al., 2012), the N4ITK algorithm is used to correct the non-uniformity of the MRI image, and the Perona-Malik algorithm is used to remove noise, the intensity of the reference CT image is clipped and normalized, (the original CT image (512x512x128, Hounsfield unit, HU) is read, the FSL tool is used for intensity clipping, the voxel intensity value is limited in the range of [-1000, 1600] HU: the voxel (the smallest unit in a three-dimensional image, representing a cubic region in a three-dimensional space) below -1000 HU is set to -1000 HU, and the voxel above 1600 HU is set to 1600 HU, and the [-1000, 1600] HU is mapped to [0, 1] by linear normalization, the calculation formula is: (intensity value+1000) / (1600+1000), and the normalized CT image (intensity [0, 1]) is output.

[0024] The tumor mask is automatically generated by using the FSL tool, the tumor area is accurately marked, the N4ITK algorithm is used to correct the non-uniformity of the MRI, and the Perona-Malik algorithm is used to remove noise, which effectively improves the quality of the MRI image, reduces the interference of artifacts, and ensures the consistency of the intensity distribution by clipping the intensity of the reference CT image and linearly normalizing it to [0, 1], which provides high-quality input for subsequent sCT generation and registration. The pre-processing procedure has high automation degree, and for the registration requirement of complex anatomical structure, the alignment accuracy of sCT and reference CT is significantly improved, which lays a foundation for subsequent dynamic feature extraction and treatment planning generation, and enhances the reliability and efficiency of clinical image processing.

[0025] Furthermore, a 2D cGAN model is deployed to convert MRI images into sCT images. The sCT images are optimized by outputting the CREPs perceptual loss function. The 2D cGAN (Goodfellow et al., 2014) model is deployed, which includes a generator G consisting of 6 residual networks (3×3 convolutional kernels, batch normalization, LeakyReLU activation, slope 0.2) to map 2D MRI slices (256×256×1) to sCT slices (256×256×1). The discriminator D is PatchGAN (binary cross-entropy loss) to judge the authenticity of the input image (CT / sCT) and output probability values ​​([0,1]). The standardized MRI image (256×256×128, intensity [0,1]) is input into the 2D cGAN model, which is divided into 128 2D slices by axis. The 128 2D sCT slices are recombined to output the sCT image (256×256×128, intensity [0,1]). The generator and discriminator are optimized using the CREPs perceptual loss method to make the sCT images resemble CT images. Style features are extracted based on the ConvNext-Tiny model, including: calculating the Gram matrix difference of the features at the j-th layer using the standard linear algebra method, with the formula as follows: , in, The Gram matrix captures the texture and intensity distribution of image I (sCT or CT). The feature extraction function for the j-th layer of ConvNext-Tiny is a network layer consisting of convolutional operations (7×7 convolutional kernel, stride 1, LayerNorm, GELU activation), which outputs a feature map ψ after taking an image as input. Let be the number of channels in the feature map of layer j, defined by the ConvNext-Tiny architecture. For feature map The transpose of the matrix, , Let the height and width of the feature map of the j-th layer be denoted as . Construct the style loss formula: , in, For style loss, sCT and CT are represented at the j-th layer of ConvNext-Tiny ( Gram matrix differences of layer features measure the consistency of texture and intensity distribution. It is the Frobenius norm; Calculate the adversarial loss To ensure the authenticity of sCT, the formula is: , in, The output probability of the discriminator for the reference CT, value range [0, 1], is obtained by the forward propagation calculation of the PatchGAN, The output probability of the discriminator for sCT (the MRI processed by the generator), is obtained by the forward propagation of the PatchGAN for the generator output sCT, and the Adam optimizer is iterated to obtain, and The probability distribution expectation of the CT and MRI samples is calculated by the training data batch (Monte Carlo estimation); The total perceptual loss is calculated The convolution layer weights and biases of the generator and discriminator are updated (the normalized MRI image is cut into 128 2D slices, the generator converts the MRI slice into an sCT slice, the discriminator receives the sCT and reference CT slice, and outputs a probability map of authenticity, the total perceptual loss is updated using the Adam optimizer to update the convolution layer weights and biases of the 2D cGAN generator and discriminator through backpropagation, a fixed number of epochs (such as 400) are trained, and high-quality sCT slices are generated), the formula is: , wherein, is the dynamic weight, the initial value , and β is determined by optimizing the MAE minimization of the validation set; The KL divergence is calculated by the entropy function of SciPy: , wherein, is the KL divergence, which measures the difference between the probability distribution of the MRI and CT intensity histogram and is a non-negative scalar, and are the probability values of the MRI and CT in the i-th bin, the MRI and CT intensity (the numerical value of each pixel (or voxel in the case of a three-dimensional image) in the image) histogram (statistical chart of the frequency of each intensity value in the image, 256 bins) is calculated by the histogram function of NumPy, and the probability distribution is obtained after normalizing the frequency; If > threshold U (based on the validation set of the MRI / CT training data, obtained by grid search), then , otherwise set , and are the high and low weight values, which are defined by particle swarm optimization (PSO) on the validation set (220 pairs of MRI / CT).

[0026] By deploying a 2D cGAN model combined with a CREPs perceptual loss function, the sCT image generation quality and registration accuracy are significantly improved. Compared to the traditional GAN method relying on a single adversarial loss, the 2D cGAN efficiently converts MRI slices into sCT slices, and the CREPs perceptual loss optimizes the consistency of texture and intensity distribution, reduces MAE, and dynamically adjusts the contribution of multi-scale features with adaptive weights, overcoming the imbalance between high-level semantics and low-level texture caused by fixed weights. The total perceptual loss based on style loss and adversarial loss iteratively updates the model parameters through the Adam optimizer, ensuring the superiority of sCT images in visual authenticity and anatomical consistency, significantly enhancing the robustness of sCT generation and registration accuracy, providing high-quality input for subsequent nonlinear registration and clinical treatment planning, especially suitable for radiotherapy planning in complex anatomical regions.

[0027] S2, generating a dynamic feature based on the optimized sCT image and the reference CT, and generating a multi-scale feature set based on the dynamic feature and the dynamic weight; Using the multi-scale feature set, a global and local deformation field is generated based on a CNN model and a UNet model, the mutual information loss value after registration is calculated by mutual information loss method, and the total loss generated by combining the global and local deformation field is used to optimize the CNN model and UNet model parameters to obtain the registered sCT image; Specifically, based on the optimized sCT image and the reference CT, a dynamic feature is generated, and a multi-scale feature set is generated based on the dynamic feature and the dynamic weight. ConvNext-Tiny (a pre-trained convolutional neural network) is used to extract features from the optimized sCT and the reference CT respectively: the first layer: low-level texture features (256x256x64), the third layer: middle-level edge features (64x64x128), and the fifth layer: high-level semantic features (16x16x256) (the optimized sCT and reference CT images are input into the pre-trained ConvNext-Tiny model, and the model gradually extracts the features of the images through multiple convolutional layers and down-sampling layers. In the first layer, the model extracts low-level texture features, and the output feature map size is 256x256x64. This layer mainly captures the basic texture information in the image. In the third layer, the model further extracts middle-level edge features, and the output feature map size is 64x64x128. This layer can capture more complex edge and structure information. In the fifth layer, the model extracts high-level semantic features, and the output feature map size is 16x16x256. This layer can capture high-level semantic information in the image, such as the shape and position of organs or tissues). The tumor volume difference is calculated (the number of voxels of the sCT tumor mask is denoted as The number of voxels of the CT tumor mask is denoted as The voxel number difference is calculated: ΔN=∣ − |, multiply the voxel number difference by the voxel volume (obtain the voxel volume based on the image resolution, such as 1x1x1mm3), to obtain the tumor volume difference), if the difference > threshold (based on the validation set and DSC optimization), the weight distribution is from high weight to low weight, and the order is the 5th layer, the 3rd layer, and the 1st layer (such as 0.6, 0.3, 0.1, obtained by genetic algorithm (GA)), if the difference < threshold , the weight distribution order is the 1st layer, the 3rd layer, and the 5th layer (such as 0.4, 0.3, 0.2), the weight obtained by weighted voting method, on the validation set, test different weight combinations (such as the distribution of the 1st, 3rd, and 5th layers), evaluate the image quality according to the peak signal-to-noise ratio (PSNR) of the generated sCT and the reference CT, and select the weight combination with the highest PSNR as the optimal weight; ConvNext-Tiny directly extracts the 1st, 3rd, and 5th layer features from sCT and CT , to generate a multi-scale feature set .

[0028] By generating dynamic features based on optimized sCT and reference CT, and combining dynamic weights to generate a multi-scale feature set, the registration accuracy and adaptability are significantly improved. The pre-trained ConvNext-Tiny model is used to extract the 1st layer low-level texture, the 3rd layer middle-level edge, and the 5th layer high-level semantic feature, which realizes the comprehensive capture of multi-scale features, overcomes the limitations of single-scale features in traditional methods, dynamically adjusts the weight distribution by calculating the tumor volume difference and optimizing the threshold based on the validation set, optimizes the weight combination using genetic algorithm and weighted voting method to maximize PSNR, and adaptively balances high-level semantics and low-level details, which enhances the anatomical consistency of sCT and reference CT, and provides high-quality feature input for subsequent global and local deformation field generation.

[0029] Further, the multi-scale feature set is used to generate global and local deformation fields based on CNN model and UNet model, the mutual information loss value after registration is calculated by mutual information loss method, and the total loss generated by combining global and local deformation fields is optimized to obtain the registered sCT image based on CNN model (5 layers of convolution, 3x3 kernel, stride 1, ReLU activation), spatial features are extracted through convolution layer, followed by fully connected layer to output 12 affine parameters (including 9 rotation / scale / shear parameters of 3x3 transformation matrix and 3 translation parameters), and then affine-to-field layer is used to convert affine parameters to global deformation field (global transformation (translation, rotation, scaling) of the entire image, provides a uniform (global deformation field provides the same 3D displacement vector for all voxels, represents translation and rotation of the entire image) 3D displacement vector for each voxel, ensures overall image alignment); composed by sCT and reference CT images and multi-scale feature sets through channel concatenation set, will input into the UNet model, which fuses global sCT, CT and multi-scale features through forward propagation, extracts spatial features, and generates local deformation field (local nonlinear deformation (such as subtle deformation or twisting of organs), provides an independent (3D displacement vector generated for each voxel is unique, personalized adjustment for local anatomical features (such as deformation of tumors or organs) of each voxel) 3D displacement vector, achieves finer local alignment by defining independent displacement of each voxel); fuses global and local deformation fields, balances global rigidity and local nonlinear deformation, formula: , wherein, is the final deformation field obtained after fusing global deformation field and local deformation field, represents the complete spatial transformation of sCT image to reference CT image, contains global (translation, rotation, scaling) and local (voxel-level nonlinear displacement) alignment information, used for accurate registration, is the fusion weight, on the validation set (20% data, 1100 pairs of MRI / CT), test different values, based on mean square error (MSE) to evaluate the registration quality of fused sCT and reference CT, select the with the lowest MSE as the final fusion weight; using linear algebra standard method, calculate the transformation coordinates of sCT voxels (3D vector, [x', y', z']), formula: +b, wherein, is the sCT voxel coordinate (3D vector, [x, y, z]), based on the voxel grid of sCT image, automatically generates 3D coordinates of each voxel, is a 3x3 rotation matrix (9 parameters, contains rotation, scaling, shear), CNN convolution layer extracts spatial features of sCT and CT, after flattening input into fully connected layer, output 9 parameters to form A), b is a 3D translation vector ([x, y, z]), , , ], CNN convolutional layer features are mapped through a fully connected layer, and three translation parameters are output to form b; Resampling formula parameters, formula: , Where, sCT original coordinate intensity (standardized MRI image is divided into 128 2D slices along the axis, input into the generator G of 2D cGAN, and the generator maps each MRI slice to a 2D sCT slice through convolution operation, and the pixel value is the voxel intensity, all 128 2D sCT slices are stacked in the original axis order to reconstruct a 3D sCT image, and each voxel coordinate Intensity directly comes from the generator output intensity value of the corresponding 2D slice), r is the voxel index, is the ITK library trilinear interpolation function (Ibanez et al., 2005), is the registered sCT intensity, is the coordinate of the target voxel (3D vector, [x, y, z]), and the target voxel coordinate is automatically generated from the output image grid using the stereomicroscope coordinate system method, is the sCT original voxel coordinate , and the registered coordinate is obtained after global transformation ; Based on the mutual information between the registered intensity and the reference CT, the statistical correlation of the intensity distribution of the two is measured, the mutual information (MI) loss is used to optimize the registration, the MI is maximized to ensure that sCT' and CT are highly consistent in intensity distribution, and the image alignment is promoted, and the formula is: , Where, is the mutual information loss (scalar), which measures the correlation of the intensity distribution of the registered sCT' and the reference CT, and is used to optimize the convolutional layer parameters of the registration model (CNN and UNet) to ensure that sCT' and CT are aligned, is the voxel intensity value of the registered sCT' and CT, and the CT intensity value of c is directly obtained from the standardized CT image, is the joint probability distribution of sCT' and CT (two-dimensional probability distribution matrix, describing the probability of all intensity pairs (s, c)), which is obtained using Parzen window estimation, is the marginal probability distribution of sCT' and CT (reflecting the probability of occurrence of a single intensity value, the value is in the range of [0, 1]), which is calculated using Parzen window estimation; Regularize the local deformation field to constrain smoothness and improve registration accuracy, and the formula is: , where, is the regularization loss, is the local DDF gradient; the total loss is calculated The formula is: , where, is the balance weight, we test different gamma values (range [0, 1]) on the validation set (20% data, 1100 pairs of MRI / CT) using cross-validation, calculate the corresponding model performance indicators (such as mutual information (MI) or structural similarity (SSIM)), average the results of all folds in cross-validation, and select the value that makes the average performance indicator optimal ; ; ; Based on the total loss update the CNN and UNet convolutional layer weights and biases (update the CNN and UNet convolutional layer weights and biases through backpropagation through the Adam optimizer based on the total loss, train until the set number of iterations (400 times) is reached, and stop, output the optimized convolutional layer weights and biases), generate an accurate final deformation field (DF) ), for each voxel coordinate of sCT , get the corresponding three-dimensional displacement vector ( , , ) (the displacement vector of each voxel is directly stored in the corresponding channel of (the first channel is , the second channel is , and the third channel is )), calculate the new coordinates , using the trilinear interpolation of the ITK library (Ibanez et al., 2005), based on the 8 adjacent integer grid points of sCT (surrounding the cube vertices of ), extract the intensity values, and weight the intensity values of the 8 grid points to obtain the intensity value of , combine the intensity values of all voxels of to fill the new image grid to output the registered sCT' image.

[0030] The global and local deformation fields are generated based on a multi-scale feature set, and the model is optimized in combination with mutual information loss and total loss, which significantly improves the registration accuracy and clinical practicability. Compared with the traditional method which only relies on rigid transformation or a single feature scale, the present scheme uses a CNN model to generate a global deformation field to capture the translation and rotation of the entire image. At the same time, a UNet model is used to fuse multi-scale features to generate a local deformation field, which accurately describes the nonlinear deformation of complex anatomical structures such as tumors, and improves the registration accuracy. Through dynamic weight adjustment and multi-scale feature fusion, the semantic and detail imbalance problem caused by fixed weight is overcome, and the target registration error is further reduced. The mutual information loss ensures that the intensity distribution of sCT and reference CT is highly consistent, and the smooth regularization loss optimizes the rationality of the local deformation field, avoiding excessive distortion. The model parameter optimization driven by the total loss realizes efficient convergence, and the generated high-precision sCT' image shows superior spatial alignment and intensity consistency on the validation set. The three linear interpolation resampling technology based on the ITK library ensures the smoothness and anatomical authenticity of the registered image, providing a reliable foundation for subsequent radiotherapy plan generation. Through the automatic optimization process and multi-dimensional loss function design, the registration quality of complex anatomical regions and the efficiency of clinical treatment planning are significantly improved.

[0031] S3, evaluate the alignment accuracy of the registered sCT image and the reference CT, verify the registration quality through automatic segmentation and key point analysis, and generate a clinical image processing plan based on the registration quality result; Specifically, the alignment accuracy of the registered sCT image and the reference CT is evaluated, and the registration quality is verified through automatic segmentation and key point analysis. The validation set (220 pairs, 20% data) of a fixed number of MRI / CT data is used to verify the registration quality of the registered sCT' and the reference CT through the FSL tool (Jenkinson et al., 2012). The white matter, gray matter, ventricle and tumor binary organ masks (1 represents the target region and 0 represents the background) are generated by automatically segmenting the registered sCT' and the reference CT, and the mask overlap degree is calculated. The number of intersecting voxels of the mask is counted, divided by twice the sum of the number of voxels of the two masks, and the Dice similarity coefficient (DSC) is obtained. The FLIRT tool of FSL is used to automatically extract a fixed number of key anatomical points (such as the vertex of the ventricle and the center of the tumor) of the registered sCT' and the reference CT, calculate the Euclidean distance between the points, and obtain the TRE (target registration error). The intensity values of sCT' and CT are compared voxel by voxel, and the average value of the absolute difference of each voxel is calculated to obtain the mean absolute error (MAE), which is used to comprehensively evaluate the registration quality (spatial alignment and intensity consistency) of sCT' and CT. ​​​​If DSC is less than threshold Q (evaluate mask overlap of sCT' and CT after registration on validation set, determine DSC threshold that meets clinical accuracy (high overlap rate)), adjust feature weights, output adjusted weight distribution (test different combination, increase 5th layer weight (high-level semantic features) to 0.6, correspondingly reduce 1st and 3rd layer weights, update ConvNext-Tiny 5th layer convolution parameters through Adam optimizer (learning rate 0.0001), regenerate multi-scale feature set, output adjusted weight distribution , and input set to update UNet to optimize registration.

[0032] By comprehensively evaluating the alignment accuracy of the registered sCT image and the reference CT, combining automatic segmentation and key point analysis, the comprehensiveness and clinical applicability of registration verification are improved, the Dice similarity coefficient is calculated to evaluate the mask overlap degree, and the FLIRT tool is used to extract key anatomical points and calculate the mean absolute error of intensity difference voxel by voxel, realizing multi-dimensional evaluation of spatial alignment and intensity consistency. When DSC is lower than the threshold, the weights are dynamically adjusted through grid search, and the multi-scale feature set is regenerated to optimize the registration performance of UNet, overcoming the limitations of manual adjustment or static evaluation of traditional methods. The feedback mechanism ensures continuous improvement of registration quality, significantly improves the reliability of registration in complex anatomical regions, and provides high-quality support for the generation of precise radiotherapy plans.

[0033] Further, based on the registration quality results, a clinical image processing plan is generated based on the registered image sCT' and the global DDF, as well as the feature weights , evaluation indicators (DSC, TRE, MAE), using the Eclipse system (Varian Medical Systems) combined with FSL segmentation of tumor and organ masks, generating a treatment area positioning plan (determining the radiation coverage range and dose distribution of the treatment area (tumor)), calculating the area coverage distribution (Monte Carlo algorithm, accuracy <2% deviation), optimizing the positioning parameters (5-7 angles, intensity distribution), if the coverage deviation is greater than the threshold T (optimized by the validation set), adjust the 5th layer convolution weight of ConvNext-Tiny through grid search according to DSC, regenerate sCT' and , output the optimized positioning plan and updated model parameters (ConvNext-Tiny, CNN, UNet) for clinical verification.

[0034] By generating a clinical image processing plan based on the registration quality result, the accuracy and automation level of the radiotherapy plan are significantly improved. By using the registered sCT image, global deformation field and feature weight, combining the tumor and organ mask segmented by FSL, and using multi-dimensional evaluation indexes such as DSC, TRE and MAE to guide the registration quality, if the coverage deviation exceeds the threshold, the grid search is used to dynamically adjust the 5th layer convolution weight of ConvNext-Tiny, update the sCT and model parameters, realize the iterative optimization of the treatment plan, compared with the inefficient method of traditional manual adjustment of parameters, the scheme automatically integrates the registration quality feedback and treatment plan generation, significantly improves the tumor positioning accuracy and plan generation efficiency, provides reliable support for precise radiotherapy of complex anatomical structure, and is especially suitable for high-precision radiotherapy planning.

[0035] The embodiment also provides a deep learning-based multi-modal image registration processing system, which comprises: A data collection and preprocessing module is configured to collect daily MRI and reference CT images and output standardized MRI and CT images. An sCT generation module is configured to deploy a 2DcGAN model, optimize a generator and a discriminator through a CREPs perception loss, and generate a high-quality sCT image. A dynamic feature extraction and multi-scale feature set generation module is configured to extract multi-layer features from the optimized sCT and the reference CT using a pre-trained ConvNext-Tiny and generate a multi-scale feature set. A deformation field generation and registration module is configured to generate global and local deformation fields by CNN and UNet respectively, fuse the global and local deformation fields, optimize mutual information loss and smooth regularization loss, and generate a registered image. A registration quality evaluation module is configured to calculate DSC, extract key points by FSLFLIRT, calculate TRE, and calculate MAE by comparing intensities voxel by voxel.

[0036] The embodiment also provides a computer device suitable for the deep learning-based multi-modal image registration processing method, which comprises a memory and a processor.

[0037] The computer device can be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected by a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved by WIFI, an operator network, NFC (Near Field Communication) or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.

[0038] The embodiment also provides a storage medium having a computer program stored thereon, the program being executed by a processor to implement the method and system for multi-modal image registration based on deep learning as described above. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic storage, a flash memory, a magnetic disk or an optical disk.

[0039] In summary, the present application generates global and local deformation fields by dynamic feature generation and dynamic weight selection of multi-scale feature sets, combines CNN and UNet models, effectively captures overall rigid transformation and local nonlinear deformation, improves the spatial alignment accuracy of sCT and reference CT, optimizes model parameters through mutual information loss and total loss, ensures intensity distribution consistency, and automatically adjusts feature weights and treatment planning parameters using registration quality feedback, significantly improves the coverage accuracy and efficiency of radiotherapy planning, realizes automatic, high-precision image registration and treatment planning optimization compared with traditional manual adjustment or fixed weight method, enhances the reliability and practicality of clinical image processing, and is particularly suitable for radiotherapy planning of complex anatomical regions.

[0040] It should be noted that the above examples are only used to illustrate the technical solutions of the present application but not to limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the present application, and all modifications and equivalents should be included in the scope of the claims of the present application.

Claims

1. A multimodal image registration processing method based on deep learning, characterized in that: include, MRI and reference CT images are collected, preprocessed separately, and a 2DcGAN model is deployed to convert MRI images into sCT images. The sCT images are then optimized by using the CREPs perceptual loss function. Dynamic features are generated based on optimized sCT images and reference CT, and a multi-scale feature set is generated based on dynamic features and dynamic weights. Global and local deformation fields are generated based on CNN and UNet models using multi-scale feature sets. The mutual information loss value after registration is calculated by the mutual information loss method. The total loss generated by fusing global and local deformation fields is combined to optimize the CNN and UNet model parameters and obtain the registered sCT image. The alignment accuracy between the registered sCT image and the reference CT is evaluated. The registration quality is verified through automatic segmentation and key point analysis. Based on the registration quality results, a clinical image processing plan is generated.

2. The deep learning-based multimodal image registration processing method as described in claim 1, characterized in that: The process of collecting MRI and reference CT images and performing preprocessing involves collecting daily MRI and reference CT images, automatically generating tumor masks using FSL, correcting non-uniformity in MRI images using the N4ITK algorithm, removing noise using the Perona-Malik algorithm, and performing intensity cropping and normalization on the reference CT images.

3. The deep learning-based multimodal image registration processing method as described in claim 2, characterized in that: The 2DcGAN model is deployed to convert MRI images into sCT images. The sCT images are optimized by outputting the CREPs perceptual loss function. The 2DcGAN model includes a generator G that maps 2D MRI slices to sCT slices, and a discriminator D that is PatchGAN to judge the authenticity of the input image and output a probability value. The standardized MRI image is input into the 2DcGAN model and divided into 128 2D slices by axis. The 128 2D sCT slices are recombined to output the sCT image. The generator and discriminator are optimized using the CREPs perceptual loss method to make sCT images resemble CT images. Style features are extracted based on the ConvNext-Tiny model, including: calculating the Gram matrix difference of the features at the j-th layer using standard linear algebra methods, and constructing the style loss. Calculate the countermeasure loss Calculate the total perceived loss Update the convolutional layer weights and biases of the generator and discriminator, and calculate the KL divergence using SciPy's entropy function. If the threshold U is greater than the total perceived loss, then... Otherwise set .

4. The deep learning-based multimodal image registration processing method as described in claim 3, characterized in that: The process of generating dynamic features based on optimized sCT images and reference CT, and generating multi-scale feature sets based on dynamic features and dynamic weights, refers to using ConvNext-Tiny to extract features from optimized sCT and reference CT respectively: Layer 1: low-level texture features, Layer 3: mid-level edge features, Layer 5: high-level semantic features; Calculate the difference in tumor volume; if the difference is greater than a threshold... Weights are assigned from high to low weight in the order of layer 5, layer 3, and layer 1. If the difference is less than a threshold... The weight allocation order is layer 1, layer 3, and layer 5, with weights... The weighted voting method was used to test different weight combinations on the validation set. The image quality was evaluated based on the peak signal-to-noise ratio of the generated sCT and the reference CT. The weight combination with the highest PSNR was selected as the final weight. ConvNext-Tiny directly extracts features from layers 1, 3, and 5 of sCT and CT. , ), generating multi-scale feature sets .

5. The deep learning-based multimodal image registration processing method as described in claim 4, characterized in that: The process involves using multi-scale feature sets to generate global and local deformation fields based on CNN and UNet models. The mutual information loss value after registration is calculated using the mutual information loss method. The total loss generated by fusing the global and local deformation fields is then combined to optimize the CNN and UNet model parameters, resulting in the registered sCT image index. Based on the CNN model, spatial features are extracted through convolutional layers, followed by fully connected layers that output 12 affine parameters. These affine parameters are then converted into a global deformation field using an Affine-to-Field layer. ; It is composed of sCT and reference CT images and multi-scale feature sets stitched together via channels. Set, The UNet model is fed into the dataset. Through forward propagation, the UNet model fuses global sCT, CT, and multi-scale features to extract spatial features and generate local deformation fields. The final deformation field is obtained by fusing the global and local deformation fields. ; The transformed coordinates of the sCT voxels are calculated using the standard linear algebra method. Calculate the resampling formula parameters to obtain the registered sCT intensity. ; Based on registration The mutual information between intensity and reference CT is used to optimize registration, maximizing the inter-information ratio (MI), and thus obtaining the mutual information loss. For local deformation fields Perform regularization to obtain the regularization loss. Calculate the total loss Based on total loss Update the weights and biases of the CNN and UNet convolutional layers to generate the final deformation field. It outputs the registered sCT' image.

6. The deep learning-based multimodal image registration processing method as described in claim 5, characterized in that: The alignment accuracy of the registered sCT' and the reference CT is evaluated by automatic segmentation and key point analysis to verify the registration quality. A validation set of a fixed number of MRI / CT data is used. The registered sCT' and the reference CT are automatically segmented using the FSL tool to generate binary organ masks for white matter, gray matter, ventricles, and tumors. The mask overlap is calculated by counting the number of voxels in the intersection of the sCT' and CT masks and dividing it by twice the sum of the number of voxels in the two masks to obtain the Dice similarity coefficient. The FLIRT tool of FSL is used to automatically extract a fixed number of key anatomical points from the sCT' and CT, and the Euclidean distance between the points is calculated to obtain the TRE. The intensity values ​​of the sCT' and CT are compared voxel by voxel, and the average of the absolute differences of each voxel is calculated to obtain the mean absolute error. If DSC is less than the threshold Q, adjust the feature weights, output the adjusted weight allocation, update the parameters of the 5th layer of ConvNext-Tiny, regenerate the multi-scale feature set, and output the adjusted weight allocation. , and enter The collection updates UNet to optimize registration.

7. The deep learning-based multimodal image registration processing method as described in claim 6, characterized in that: The generation of a clinical image processing plan based on the registration quality results refers to the process of generating a clinical image processing plan based on the registered image sCT' and the global DDF, as well as feature weights. The evaluation metrics were assessed using the Eclipse system combined with FSL-segmented tumor and organ masks to generate a treatment area localization plan. The regional coverage distribution was calculated, and localization parameters were optimized. If the coverage deviation exceeded a threshold T, the weights of the 5th layer of the ConvNext-Tiny convolution were adjusted via grid search based on DSC, and the sCT' was regenerated. The optimized localization plan and updated model parameters are output for clinical validation.

8. A deep learning-based multimodal image registration and processing system, based on the deep learning-based multimodal image registration and processing method according to any one of claims 1 to 7, characterized in that: Includes a data collection and preprocessing module for collecting daily MRI and reference CT images and outputting standardized images; The sCT generation module is used to deploy a 2DcGAN model, which optimizes the generator and discriminator through CREPs perceptual loss to generate high-quality sCT images. The dynamic feature extraction and multi-scale feature set generation module is used to extract multi-layer features from optimized sCT and reference CT using pre-trained ConvNext-Tiny and generate multi-scale feature sets. The deformation field generation and registration module is used to generate global and local deformation fields for CNN and UNet respectively, and then fuse them to optimize mutual information loss and smoothing regularization loss, thereby generating the registered image. The registration quality assessment module is used to calculate DSC, extract key points through FSLFLIRT, calculate TRE, and calculate MAE by comparing voxel intensity.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the deep learning-based multimodal image registration processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the deep learning-based multimodal image registration processing method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • IOS image and CBCT image registration method based on particle swarm and single-tooth optimization

    CN121544676A

  • Multi-modal medical image segmentation method and device, computer equipment and storage medium

    CN121837614A

  • Multi-modal medical image segmentation method and device, computer device and storage medium

    CN121837614B