Multi-modal Medical Image Registration Method, Device, Equipment and Storage Medium

Through the combination of multimodal medical image registration method and segmentation neural network, the problem of MRI medical image aberration is solved, the accuracy of lesion detection and the accuracy of fusion images are improved, and more accurate quantitative analysis is supported.

CN119991750BActive Publication Date: 2025-07-18CARBON (SHENZHEN) MEDICAL DEVICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510459824.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-18
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Multimodal MRI medical images are prone to the physiological movement of patients due to the long shooting time and are susceptible to the patient's physiological movement, resulting in misalignment of image space and pixels, affecting the accuracy of lesion detection.

Method used

The multimodal medical image registration method is used to adjust the image through the registration network to generate deformation field data, and the segmented neural network is used to optimize the deformation field data to achieve image alignment and fusion.

Benefits of technology

It improves the lesion detection accuracy of multimodal MRI medical images, ensures the accuracy and reliability of the fusion image, and supports more accurate quantitative analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991750B_ABST
    Figure CN119991750B_ABST
Patent Text Reader

Abstract

The present application provides a multi-modal medical image registration method, apparatus, device and storage medium, relating to the technical field of medical devices. The method includes: obtaining an initial medical image set, using an initialization registration network to perform initialization training on the initial medical image set to generate first deformation field data and second deformation field data, adjusting a second initial medical image based on the first deformation field data to generate a first initial registered image, adjusting a third initial medical image based on the second deformation field data to generate a second initial registered image, fusing the first initial registered image, the second initial registered image and the first initial medical image to obtain a fused medical image, using a segmentation neural network to perform segmentation training on the fused medical image to obtain a segmentation result, calculating loss data according to the segmentation result to optimize the first deformation field data and the second deformation field data, so as to realize optimizing the segmentation neural network and the registration network by using the loss data while training the segmentation neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical devices, and particularly to a multi-modal medical image registration method, device, equipment, and storage medium. Background Art

[0002] Multi-modal nuclear magnetic resonance imaging (MRI) medical images play an important role in diagnosing and evaluating lesions. Common modalities in MRI medical images include diffusion-weighted imaging (DWI), apparent diffusion coefficient (ADC), T1-weighted imaging (T1W1), and T2-weighted imaging (T2W1), etc. Different modalities of MRI medical images provide different tissue information. Comprehensive analysis of multiple modalities of MRI medical images can improve the accuracy and reliability of lesion diagnosis.

[0003] For some MRI medical images with a long acquisition time and easily affected by physiological movements such as patient breathing and heartbeat, such as medical images with the modality of DWI, there are easily misalignment problems in the image space and pixels, and such misalignment problems will seriously affect the comprehensive lesion detection and analysis of multi-modal MRI medical images, thereby reducing the lesion detection accuracy of multi-modal MRI medical images.

[0004] Based on this, it is necessary to propose a solution to the misalignment problem of multi-modal MRI medical images. Summary of the Invention

[0005] In view of the above problems, the present application provides a multi-modal medical image registration method, device, equipment, storage medium, and program product.

[0006] To achieve the above object:

[0007] In a first aspect, an embodiment of the present application provides a multi-modal medical image registration method, including:

[0008] Obtain an initial medical image set, where the initial medical image set includes a first initial medical image, a second initial medical image, and a third initial medical image with different modalities;

[0009] Initialize and train the initial medical image set using a registration network, and generate first deformation field data and second deformation field data according to the initialization training, where the first deformation field data is the deformation field data between the first initial medical image and the second initial medical image, and the second deformation field data is the deformation field data between the first initial medical image and the third initial medical image;

[0010] Adjust the second initial medical image based on the first deformation field data to generate a first initial registered image, and adjust the third initial medical image based on the second deformation field data to generate a second initial registered image;

[0011] Fuse the first initial registered image, the second initial registered image, and the first initial medical image to obtain a fused medical image;

[0012] Use a segmentation neural network to perform segmentation training on the fused medical image, calculate loss data according to the segmentation results obtained from the segmentation training, and optimize the first deformation field data and the second deformation field data according to the loss data.

[0013] Further, before the step of initializing and training the initial medical image set using the registration network, the method further includes:

[0014] Normalize the first initial medical image, the second initial medical image, and the third initial medical image respectively.

[0015] Further, the step of adjusting the second initial medical image based on the first deformation field data to generate a first initial registered image includes:

[0016] Obtain a first standard sampling grid of the second initial medical image;

[0017] Obtain a first reference sampling grid according to the first standard sampling grid and the first deformation field data;

[0018] Normalize the first reference sampling grid to obtain a first target sampling grid;

[0019] Sample the second initial medical image according to the first target sampling grid and a preset sampling interpolation function to obtain the first initial registered image.

[0020] Further, the step of adjusting the third initial medical image based on the second deformation field data to generate a second initial registered image includes:

[0021] Obtain a second standard sampling grid of the third initial medical image;

[0022] Obtain a second reference sampling grid based on the second standard sampling grid and the second deformation field data;

[0023] Perform normalization processing on the second reference sampling grid to obtain a second target sampling grid;

[0024] Sample the third initial medical image according to the second target sampling grid and a preset sampling interpolation function to obtain the second initial registration image.

[0025] Further, the step of calculating loss data according to the segmentation result obtained by segmentation training includes:

[0026] Obtain a preset loss function, and calculate the loss data according to the loss function, the segmentation result, and the true annotation of the fused medical image.

[0027] Further, the step of optimizing the first deformation field data and the second deformation field data according to the loss data includes:

[0028] Judge whether the loss data converges to a first preset condition;

[0029] If not, calculate the gradient of the loss function with respect to the first deformation field data and calculate the gradient of the loss function with respect to the second deformation field data, and update the first deformation field data by backpropagation according to the gradient of the first deformation field data and a preset global learning rate, and update the second deformation field data by backpropagation according to the gradient of the second deformation field data and the global learning rate.

[0030] Further, the method further includes:

[0031] Judge whether the loss data converges to a first preset condition;

[0032] If not, calculate the gradient of the loss function with respect to the segmentation result, and update the segmentation result by backpropagation according to the gradient of the segmentation result.

[0033] Further, the method further includes:

[0034] Judge whether the loss data converges to a first preset condition;

[0035] If not, calculate the gradient of the loss function with respect to the segmentation parameters of the segmentation neural network, and update the segmentation parameters by backpropagation according to the gradient of the segmentation parameters and a preset global learning rate.

[0036] Further, the first initial medical image is a T2-weighted imaging medical image, and one of the second initial medical image and the third initial medical image is an apparent diffusion coefficient medical image, and the other of the second initial medical image and the third initial medical image is a diffusion-weighted imaging medical image;

[0037] Or, the first initial medical image is the apparent diffusion coefficient medical image, and one of the second initial medical image and the third initial medical image is the diffusion-weighted imaging medical image, and the other of the second initial medical image and the third initial medical image is the T2-weighted imaging medical image;

[0038] Or, the first initial medical image is a diffusion-weighted imaging medical image, and one of the second initial medical image and the third initial medical image is the T2-weighted imaging medical image, and the other of the second initial medical image and the third initial medical image is the apparent diffusion coefficient medical image.

[0039] In a second aspect, an embodiment of the present application provides a multi-modal medical image registration device, including:

[0040] An initial medical image set acquisition module, configured to acquire an initial medical image set, where the initial medical image set includes a first initial medical image, a second initial medical image, and a third initial medical image with different modalities;

[0041] A deformation field data generation module, configured to perform initialization training on the initial medical image set by using a registration network, and generate first deformation field data and second deformation field data according to the initialization training, where the first deformation field data is the deformation field data between the first initial medical image and the second initial medical image, and the second deformation field data is the deformation field data between the first initial medical image and the third initial medical image;

[0042] An initial registration image generation module, configured to adjust the second initial medical image based on the first deformation field data to generate a first initial registration image and adjust the third initial medical image based on the second deformation field data to generate a second initial registration image;

[0043] An image fusion module, configured to fuse the first initial registration image, the second initial registration image, and the first initial medical image to obtain a fused medical image;

[0044] A deformation field data optimization module, configured to perform segmentation training on the fused medical image by using a segmentation neural network, calculate loss data according to the segmentation result obtained from the segmentation training, and optimize the first deformation field data and the second deformation field data according to the loss data.

[0045] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0046] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0047] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes computer instructions. The computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of any one of the methods provided in the first aspect above.

[0048] In the multi-modal medical image registration method, device, equipment, storage medium, and program product provided by the present application, taking the first initial medical image of the initial medical image set as a reference, the second initial medical image is corrected by the first deformation field data of self-supervised learning, and the third initial medical image is corrected by the second deformation field data of automatic supervised learning, so as to realize the automatic alignment of the first initial registration image obtained based on the second initial medical image and the first initial medical image, and to realize the automatic alignment of the second initial registration image obtained based on the third initial medical image and the first initial medical image. Then, a segmentation neural network is used to segment the fused medical image obtained by fusing the first initial registration image, the second initial registration image, and the first initial medical image, and loss data is calculated according to the segmentation result to optimize the first deformation field data and the second deformation field data, so as to realize optimizing the segmentation neural network and the registration network by using the loss data while training the segmentation neural network, and further improve the lesion detection accuracy based on the fused medical image, facilitating accurate quantitative analysis and research for medical images including multi-modal fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic flowchart of a multi-modal medical image registration method provided by an embodiment of the present application;

[0050] Figure 2 It is a schematic flowchart of a multi-modal medical image registration method provided by another embodiment of the present application;

[0051] Figure 3 is according to Figure 1 or Figure 2 shown schematic diagram of multi-modal medical image registration;

[0052] Figure 4 Block diagram of a multimodal medical image registration device provided for an embodiment of the present application;

[0053] Figure 5 Structural schematic diagram of a computer device provided for an embodiment of the present application. Detailed implementation manners

[0054] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.

[0055] In addition, in some processes described in the specification, claims and the above-mentioned accompanying drawings of the present application, there are multiple operations that appear in a specific order. These operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish each different operation, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0056] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0057] The inventors of the present application have found through research that in MRI medical images, the shooting time of MRI medical images with modalities of DWI, ADC, and T2W1 is relatively long, and it is easily affected by physiological movements such as the patient's breathing and heartbeat, and the problem of misalignment of image space and pixels is likely to occur. This problem will seriously affect the comprehensive analysis of multimodal MRI medical images and reduce the detection accuracy of lesions. Especially for deep learning models, the design of multi-channel input can provide rich feature representations for fusing medical images of different modalities. However, misaligned pixels may cause the loss of some information or the incorrect superposition of some information when fusing medical images of different modalities. Misaligned pixels may also cause the boundaries of the fused medical images of different modalities to be blurred, resulting in inaccurate lesion detection due to the confusion of features at different positions in the subsequent fused medical images of different modalities.

[0058] In order to improve the accuracy of lesion detection in MRI medical images that are prone to pixel space and pixel misalignment problems during multi-modal image fusion input, the inventors of the present application propose to introduce a medical image registration method during the multi-modal MRI medical image data fusion input to improve the lesion detection accuracy of the fused multi-modal MRI medical images.

[0059] In a first aspect, in an embodiment of the present application, as Figure 1 shown, a multi-modal medical image registration method is provided. The method 100 includes:

[0060] Step S10, obtaining an initial medical image set.

[0061] Among them, the initial medical image set includes a first initial medical image, a second initial medical image, and a third initial medical image with different modalities.

[0062] In some embodiments, medical images of the same target object can be collected by different MRI imaging devices to obtain initial medical images of different modalities of the same target object. Correspondingly, different modalities of MRI medical images of the same target object can also be collected by the same MRI imaging device to obtain initial medical images of different modalities of the same target object.

[0063] In other embodiments, the initial medical image set can be stored in a storage device, so that when performing this step S10, the initial medical image set is obtained from the storage device in a data reading manner.

[0064] Step S30, using a registration network to perform initialization training on the initial medical image set, and generating first deformation field data and second deformation field data according to the initialization training.

[0065] Among them, the first deformation field data is the deformation field data between the first initial medical image and the second initial medical image. The second deformation field data is the deformation field data between the first initial medical image and the third initial medical image.

[0066] That is, in this step S30, taking the first initial medical image as a fixed image, and taking the second initial medical image and the third initial medical image as floating images respectively, outputting the first deformation field data corresponding to the second initial medical image and the second deformation field data corresponding to the third initial medical image.

[0067] Furthermore, in an embodiment, the first initial medical image is a T2-weighted imaging (T2WI) medical image, the second initial medical image is an apparent diffusion coefficient (ADC) medical image, and the third initial medical image is a diffusion-weighted imaging (DWI) medical image.

[0068] Since the T2W1 medical image, the ADC medical image, and the DWI medical image in the MRI medical image all take a relatively long time to capture and are easily affected by physiological movements such as the patient's breathing and heartbeat, any one of the three is prone to image spatial and pixel misalignment. And the three play an important role in the diagnosis and evaluation of prostate tumors in prostate multi-modal magnetic resonance imaging. In this embodiment, exemplarily, the T2W1 medical image is used as the fixed image, and the ADC medical image and the DWI medical image are used as the floating images respectively, so as to introduce an efficient and autonomous multi-modal image registration method to improve the detection accuracy of prostate tumor recognition.

[0069] It is worth mentioning that in another embodiment of the present application, the ADC medical image can also be used as the fixed image, and the T2W1 medical image and the DWI medical image are used as the floating images respectively. In yet another embodiment of the present application, the DWI medical image can also be used as the fixed image, and the ADC medical image and the T2W1 medical image are used as the floating images respectively.

[0070] Furthermore, in one embodiment, the "trained registration network" refers to a three-dimensional encoding and decoding model (also known as a 3D encoding and decoding model) trained with data. Specifically, the 3D encoding and decoding model includes a three-dimensional data encoder and a three-dimensional data decoder. Among them, the three-dimensional data encoder is used to extract three-dimensional features from the fixed image and the floating image, and the three-dimensional data decoder is used to upsample the three-dimensional features extracted by the three-dimensional data encoder to obtain the deformation field data between the fixed image and the floating image. Among them, the deformation field data includes the deformation field and the change parameters.

[0071] Exemplarily, the "extracting the first deformation field data" in step S30 may specifically include:

[0072] Step S311, gradually reducing the size of the second initial medical image and downsampling to extract the features of the second initial medical image based on the convolutional blocks and pooling layers of the three-dimensional data encoder.

[0073] In step S311, taking the size of the input second initial medical image as 1×C×H×W×D as an example, where 1 represents the batch size, C represents the number of channels, and H, W, and D respectively represent the height, width, and depth of the second initial medical image.

[0074] The three-dimensional data encoder gradually reduces the size of the second initial medical image and extracts its features through multiple convolutional blocks and pooling layers. Among them, each convolutional block includes multiple three-dimensional convolutional layers, and each convolutional layer is followed by a ReLU activation function for introducing non-linearity and a batch normalization layer. In this way, the new size output after passing through the convolutional block for a size of 1×C×H×W×D is 1×2C×H×W×D. That is to say, the number of channels of the second initial medical image doubles, while the height, width, and depth of the second initial medical image remain unchanged.

[0075] Furthermore, the pooling layer of the three-dimensional data encoder is used to reduce the size of the second initial medical image. Generally speaking, the pooling layer uses max pooling with a stride of 2. Thus, each time after passing through this pooling layer, the output second initial medical image is: .

[0076] That is, in the size of the second initial medical image each time after passing through this pooling layer, the number of channels of the second initial medical image remains unchanged, while the height, width, and depth of the second initial medical image are all reduced by half. The second initial medical image after N times of downsampling feature extraction operations is output to the bridging layer of the three-dimensional data encoder.

[0077] Step S312: Based on the bridging layer of the three-dimensional data encoder, directly transfer the features of the extracted second initial medical image to the corresponding layer of the three-dimensional data decoder.

[0078] Step S313: Based on the upsampling layer and convolutional blocks of the three-dimensional data decoder, gradually restore the size of the second initial medical image.

[0079] Corresponding to the N times of downsampling feature extraction operations in step S311, in step S313, N times of upsampling decoding are required. Subsequently, the second initial medical image after N times of upsampling decoding is output to the output layer of the three-dimensional data decoder.

[0080] Among them, N is a preset value, and N is a non-zero integer. When the floating image is an image with low quality, if the number of downsampling times of the floating image is too many, the main features in it will be lost. By limiting the number of downsampling times of the floating image to a set number, the three-dimensional data encoder can reduce the loss of features when downsampling the floating image, improve the accuracy of the first deformation field data between the obtained first initial medical image and the second initial medical image, and further improve the accuracy of the subsequent obtained registration result.

[0081] Step S314: Generate learnable first deformation field data based on the output layer of the three-dimensional data decoder.

[0082] Among them, the first deformation field data includes a first deformation field and first deformation field parameters.

[0083] In step S314, the output layer passes through a three-dimensional convolutional layer with 3 channels to generate the first deformation field. The first deformation field is a three-dimensional vector field, representing the displacement of each voxel in the second initial medical image in space.

[0084] As described above, according to steps S311 to S314, the second initial medical image generates a learnable first deformation field flow_1. The first deformation field flow_1 represents a five-dimensional tensor, and the dimensions of the first deformation field flow_1 are (N, 1, 3, H, W), where N is the batch size, 1 is the number of channels, 3 represents the components of the displacement vector of each voxel in three dimensions, and H and W are the height and width of the second initial medical image, respectively.

[0085] The "extracting the second deformation field data" in step S30 can be implemented with reference to the aforementioned steps S311 to S314, and will not be elaborated here.

[0086] It is worth mentioning that in one embodiment, as Figure 2 shown, before step S30, the method 100 may further include:

[0087] Step S20, performing normalization processing on the first initial medical image, the second initial medical image, and the third initial medical image respectively.

[0088] That is, in this step S20, preprocessing is performed on each initial medical image in the initial medical image set. The intensity values of each initial medical image in the initial medical image set are mapped to the range of 0 to 1 by linear normalization, and the size of each initial medical image in the initial medical image set is adjusted to the size of 1×C×H×W×D respectively.

[0089] Exemplarily, in one embodiment, minimum-maximum normalization is used to scale each initial medical image in the initial medical image set to between 0 and 1. Among them, for the same initial medical image, the minimum intensity value of the initial image is mapped to 0, and the maximum intensity value of the initial image is mapped to 1.

[0090] In this embodiment, when preprocessing the initial medical image set, each initial medical image in the initial medical image set is normalized, so as to reduce the difference between any two initial medical images, and further reduce the difference between the fixed image and any floating image, facilitating the registration network to quickly obtain the first deformation field data and the second deformation field data, and improving the working efficiency of multiple registrations of medical images of different modalities.

[0091] Please continue to refer to Figure 1 and Figure 2, after step S30, the method 100 includes:

[0092] Step S40, adjusting a second initial medical image based on first deformation field data to generate a first initial registered image, and adjusting a third initial medical image based on second deformation field data to generate a second initial registered image.

[0093] In step S40, by registering the second initial medical image with the first initial medical image and registering the third initial medical image with the first initial medical image so as to synchronously fuse all the initial medical images in the initial medical image set in subsequent steps, the working efficiency of fusing medical images of three or more modalities is improved.

[0094] Specifically, in one embodiment, "adjusting a second initial medical image based on first deformation field data to generate a first initial registered image" in step S40 specifically includes:

[0095] Step S411, obtaining a first standard sampling grid of the second initial medical image.

[0096] The first standard sampling grid is the standard sampling grid of the second initial medical image generated when initializing the first deformation field. In medical image registration, a "standard sampling grid" refers to a regular three-dimensional grid used to define the initial positions of each voxel in a medical image. In the forward propagation process of the deformation field, the standard sampling grid is combined with the deformation field to calculate new sampling positions to achieve the deformation of the medical image. Based on the introduction of the standard sampling grid, the deformation process of the medical image can be accurately carried out.

[0097] Therefore, in this step S411, the first standard sampling grid provides an initial spatial position for each voxel in the second initial medical image, and moreover, the first standard sampling grid is uniformly distributed and covers the entire spatial range of the second initial medical image.

[0098] Step S412, obtaining a first reference sampling grid according to the first standard sampling grid and the first deformation field data.

[0099] The first reference sampling grid is the new sampling position of the second initial medical image obtained based on the first standard sampling grid and the first deformation field, thereby realizing the deformation of the second initial image.

[0100] Specifically, in the forward propagation process of the first deformation field, the first standard sampling grid is added to the first deformation field. Since the first deformation field is a vector field, that is, the first deformation field indicates the displacement direction and distance of each voxel in the second initial medical image, therefore, by adding the vector of the first deformation field to each point of the first standard sampling grid, the new sampling position of the second initial medical image can be obtained.

[0101] Step S412 can be understood by a first preset relational expression:

[0102] new_locs1 = grid1 + flow1(1);

[0103] In the first preset relational expression (1), new_locs1 is the first reference sampling grid, grid1 is the first standard sampling grid, and flow1 is the first deformation field.

[0104] Step S413: Normalize the first reference sampling grid to obtain the first target sampling grid.

[0105] Exemplarily, step S413 specifically includes normalizing the value of each point of the first parameter sampling grid to the range of [-1, 1]. In the first parameter sampling grid, "-1" and "1" represent displacement vectors with different directions but the same vector value.

[0106] Since the deformation vectors of the first deformation field may have different dimensions and orders of magnitude, normalizing the first reference sampling grid can eliminate the differences between dimensions and orders of magnitude, enabling fair comparison and processing of deformations in different directions, and simplifying the subsequent calculation process based on the first target sampling grid. For medical image registration, the first target sampling grid can be directly used for coordinate transformation without additional scaling operations.

[0107] Step S413 can be understood by a second preset relational expression:

[0108] (2);

[0109] In the second preset relational expression (2), flow_shape1[i] is the size of the first deformation field in dimension i, where dimension i specifically refers to one of the x-axis, y-axis, and z-axis in the three-dimensional space. The x-axis is perpendicular to the plane where the y-axis and z-axis are located together, the y-axis is perpendicular to the plane where the x-axis and z-axis are located together, and the z-axis is perpendicular to the plane where the x-axis and y-axis are located together.

[0110] new_locs1[:, i,...] is the position of the first deformation field after deformation in dimension i.

[0111] Step S414: Sample the second initial medical image according to the first target sampling grid and a preset sampling interpolation function to obtain the first initial registered image.

[0112] Exemplarily, the sampling interpolation function is a trilinear interpolation function. The values of each point in the first target sampling grid are interpolated using the trilinear interpolation function to obtain new sampling positions of the second initial medical image. Then, based on the new sampling positions, the voxels in the second initial medical image are resampled to obtain the first initial registration image.

[0113] Specifically, in one embodiment, "generating the second initial registration image by adjusting the third initial medical image based on the second deformation field data" in step S40 specifically includes:

[0114] Step S421, obtaining a second standard sampling grid of the third initial medical image.

[0115] The second standard sampling grid is the standard sampling grid of the third initial medical image generated during initialization.

[0116] Step S422, obtaining a second reference sampling grid according to the second standard sampling grid and the second deformation field data.

[0117] Step S423, performing a normalization process on the second reference sampling grid to obtain a second target sampling grid.

[0118] Step S424, sampling the third initial medical image according to the second target sampling grid and a preset sampling interpolation function to obtain the second initial registration image.

[0119] For the specific principles of steps S421 to S424, reference can be made to steps S411 to S414 accordingly, and details are not elaborated in this application.

[0120] Please continue to refer to Figures 1 to 3 , after step S40, the method 100 further includes:

[0121] Step S50, fusing the first initial registration image, the second initial registration image, and the first initial medical image to obtain a fused medical image.

[0122] Exemplarily, taking the first initial medical image as a T2W1 medical image, the second initial medical image as an ADC medical image, and the third initial medical image as a DWI medical image as an example.

[0123] The first initial registration image is transformed into T(ADC; θ ADC ) according to the first deformation field parameter matrix of the first deformation field data, where θ ADC is the parameter weight of the first deformation field parameter. The second initial registration image is transformed into T(DWI; θ DWI ) according to the second deformation field parameter matrix of the second deformation field data, where θ DWIis the parameter weight of the second deformation field parameter. Then, according to step S50, the fused medical image I can be obtained based on the fusion function, the first initial registration image, the second initial registration image, and the first initial medical image f = Fuse(T(ADC; θ ADC ),T(DWI; θ DWI )),T2W1).

[0124] It is worth mentioning that the fusion function is not particularly limited in this embodiment. The fusion function can be, but is not limited to, a weighted average fusion function, a feature-based fusion function, a deep learning model-based fusion function, etc.

[0125] In step S50, by integrating different medical images of the three modalities, the advantages of each modality can be utilized to make up for the deficiencies of the medical images of a single modality, providing a more accurate detection result for lesion detection.

[0126] Step S60: Use a segmentation neural network to perform segmentation training on the fused medical image, calculate loss data according to the segmentation result obtained from the segmentation training, and optimize the first deformation field data and the second deformation field data according to the loss data.

[0127] In this step, the segmentation neural network is a deep learning model for image segmentation tasks. Its main goal is to divide the fused medical image into multiple regions, each region representing a different object or a specific category, and perform feature extraction to determine the segmentation result output from the segmentation neural network.

[0128] It is worth mentioning that the segmentation neural network is not specifically limited in this embodiment. The segmentation neural network can be, but is not limited to, a SegNet model, a fully convolutional network model, a Mask R-CNN model, a Faster R-CNN model, and a DeepLab series model, etc.

[0129] Taking the SegNet model as an example, the segmentation result can be represented by the following third preset relational expression:

[0130] S = SegmentationNet(I f )(3);

[0131] In the third preset relational expression (3), I f is the fused medical image, and S is the segmentation result.

[0132] Specifically, "calculating loss data according to the segmentation result" in step S60 specifically includes:

[0133] Obtain a preset loss function, and calculate loss data according to the loss function, the segmentation result, and the true annotation of the fused medical image.

[0134] In this embodiment, the loss function is an index that measures the difference between the segmentation result output by the segmentation neural network and the true annotation of the fused medical image.

[0135] Furthermore, in one embodiment, the Dice loss function is used as the loss function. Thus, the Dice loss function is used to evaluate the similarity between the segmentation result and the true annotation (Ground Truth, GT). The value range of the Dice loss function is between 0 and 1. When the segmentation result is exactly the same as the true annotation, the Dice loss function takes the minimum value of 0. The true annotation is the image segmentation result obtained by manual delineation by experts or using other reliable methods, which represents the "gold standard" of the segmentation task.

[0136] Among them, the Dice Loss function is specifically expressed by the following fourth preset relational expression:

[0137] (4);

[0138] In the fourth preset relational expression (4), is the loss data, is the segmentation result obtained by segmenting the fused medical image by the segmentation neural network. Among them, , is the first initial registration image, is the second initial registration image, is the first initial medical image.

[0139] Furthermore, in one embodiment, "optimizing the first deformation field data and the second deformation field data according to the loss data" in step S60 specifically includes:

[0140] Judge whether the loss data converges to the first preset condition. If not, calculate the gradient of the loss function with respect to the first deformation field data and calculate the gradient of the loss function with respect to the second deformation field data. Update the first deformation field data by backpropagation according to the gradient of the first deformation field data and the global learning rate, and update the second deformation field data by backpropagation according to the gradient of the second deformation field data and the global learning rate. Conversely, if the loss data converges to the first preset condition, the training of the segmentation neural network is completed.

[0141] Among them, the first preset condition can be set according to actual needs, and can be specifically determined according to the loss function threshold set by the called loss function or the set number of iterations. Exemplarily, still taking the Dice loss function as an example, if the loss data is closer to 0, it indicates that the segmentation result obtained by segmenting the fused medical image by the segmentation neural network is closer to the true annotation.

[0142] Among them, the gradient of the first deformation field parameter is: , θ ADC is the first deformation field parameter, and D ADC is the first deformation field. Update the first deformation field parameter by backpropagation: , where α is the global learning rate.

[0143] Among them, the gradient of the second deformation field parameter is: , θ DWI is the second deformation field parameter, and D DWI is the second deformation field. Update the second deformation field parameter by backpropagation: , where α is the global learning rate.

[0144] In the multi-modal medical image registration method provided in this embodiment, taking the first initial medical image of the initial medical image set as a reference, the second initial medical image is corrected by the first deformation field data of self-supervised learning, and the third initial medical image is corrected by the second deformation field data of automatic supervised learning, so as to realize the automatic alignment of the first initial registration image obtained based on the second initial medical image and the first initial medical image, and to realize the automatic alignment of the second initial registration image obtained based on the third initial medical image and the first initial medical image. Then, use the trained segmentation neural network to segment the fused medical image obtained by fusing the first initial registration image, the second initial registration image, and the first initial medical image, calculate the loss data according to the segmentation result to optimize the first deformation field data and the second deformation field data, thereby improving the lesion detection accuracy based on the fused medical image, and facilitating accurate quantitative analysis and research for medical images including multi-modal fusion.

[0145] Further, in an embodiment, in the method 100, after step S60, it further includes: step S70, determining whether the loss data converges to a first preset condition. If not, calculate the gradient of the loss function with respect to the segmentation result, and update the segmentation result by backpropagation according to the gradient of the segmentation result.

[0146] Specifically, the gradient of the segmentation result is: .

[0147] Similarly to the foregoing embodiment, when the value of the gradient of the segmentation result converges to the corresponding gradient threshold or the backpropagation update reaches the corresponding number of iterations, the training of the segmentation neural network is completed, and the segmentation accuracy of the fused medical image reaches a satisfactory level.

[0148] Further, in an embodiment, in the method 100, after step S60, it further includes: step S80, determining whether the loss data converges to a first preset condition. If not, calculate the gradient of the loss function with respect to the segmentation parameters of the segmentation neural network, and update the segmentation parameters by backpropagation according to the gradient of the segmentation parameters and the global learning rate.

[0149] Specifically, the parameter gradients of the segmentation neural network are as follows: . Update the segmentation parameters of the segmentation neural network by backpropagation: , where α is the global learning rate and θs are the segmentation parameters of the segmentation neural network.

[0150] Similar to the foregoing embodiments, when the value of the parameter gradient of the segmentation neural network converges to the corresponding gradient threshold or the backpropagation update reaches the corresponding number of iterations, the training of the segmentation neural network is completed, and the segmentation result of the fused medical image reaches a satisfactory segmentation accuracy.

[0151] In a second aspect, in an embodiment provided in the present application, as Figure 4 shown, a multimodal medical image registration device 200 is further provided, including:

[0152] An initial medical image set acquisition module 201, configured to acquire an initial medical image set. Among them, the initial medical image set includes a first initial medical image, a second initial medical image, and a third initial medical image with different modalities.

[0153] A deformation field data generation module 202, configured to perform initialization training on the initial medical image set by using a registration network, and generate first deformation field data and second deformation field data according to the initialization training.

[0154] Among them, the first deformation field data is the deformation field data between the first initial medical image and the second initial medical image, and the second deformation field data is the deformation field data between the first initial medical image and the third initial medical image.

[0155] An initial registration image generation module 203, configured to adjust the second initial medical image based on the first deformation field data to generate a first initial registration image and adjust the third initial medical image based on the second deformation field data to generate a second initial registration image.

[0156] An image fusion module 204, configured to fuse the first initial registration image, the second initial registration image, and the first initial medical image to obtain a fused medical image.

[0157] A deformation field data optimization module 205, configured to perform segmentation training on the fused medical image by using a segmentation neural network, calculate loss data according to the segmentation result obtained from the segmentation training, and optimize the first deformation field data and the second deformation field data according to the loss data.

[0158] For the specific limitations of the multimodal medical image registration device, reference can be made to the limitations of the multimodal medical image registration method in the foregoing text, which will not be elaborated here. Each module in the above multimodal medical image registration device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0159] In a third aspect, Figure 5 FIG. shows a schematic structural diagram of a computer device provided in an embodiment of the present application. As Figure 5 shown, the computer device 300 may include: a processor, a memory, a communication module, a display screen, an input module, a power supply module, and an audio module connected through a system bus.

[0160] Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication module of the computer device is used to communicate with external computer devices and servers in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. The computer program, when executed by the processor, implements the multimodal medical image registration method provided by the present application. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input module of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the computer device, or an external keyboard, a touchpad, or a mouse, etc. The audio module of the computer device is used to convert digital audio information into an analog audio signal for output, and is also used to convert analog audio input into digital audio signals. The audio module can also be used for encoding and decoding audio signals. In some embodiments, the audio module can be set in the processor, or some functional modules of the audio module can be set in the processor. Figure 5 only some components are schematically shown, and it does not mean that the computer device 300 only includes Figure 5 the components shown.

[0161] Correspondingly, in a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and the computer program, when executed by a computer, can implement the steps or functions of the multimodal medical image registration method provided by each of the above method embodiments.

[0162] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0163] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in the specification of this application.

[0164] As mentioned above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A multimodal medical image registration method, characterized in that, The method includes: Obtaining an initial medical image set, where the initial medical image set includes a first initial medical image, a second initial medical image, and a third initial medical image with different modalities; Using a registration network to perform initialization training on the initial medical image set, and generating first deformation field data and second deformation field data according to the initialization training, where the first deformation field data is the deformation field data between the first initial medical image and the second initial medical image, and the second deformation field data is the deformation field data between the first initial medical image and the third initial medical image; Adjusting the second initial medical image based on the first deformation field data to generate a first initial registered image, and adjusting the third initial medical image based on the second deformation field data to generate a second initial registered image; Fusing the first initial registered image, the second initial registered image, and the first initial medical image to obtain a fused medical image; Using a segmentation neural network to perform segmentation training on the fused medical image, obtaining a preset loss function, calculating loss data according to the loss function, the segmentation result obtained from the segmentation training, and the true annotation of the fused medical image, and determining whether the loss data converges to a first preset condition. If not, calculating the gradient of the loss function with respect to the first deformation field data and calculating the gradient of the loss function with respect to the second deformation field data, and updating the first deformation field data by backpropagation according to the gradient of the loss function with respect to the first deformation field data and a preset global learning rate, and updating the second deformation field data by backpropagation according to the gradient of the loss function with respect to the second deformation field data and the global learning rate.

2. The method according to claim 1, wherein Before the step of using the registration network to perform initialization training on the initial medical image set, the method further includes: Performing normalization processing on the first initial medical image, the second initial medical image, and the third initial medical image respectively.

3. The method according to claim 1, wherein The step of adjusting the second initial medical image based on the first deformation field data to generate a first initial registered image includes: Obtaining a first standard sampling grid of the second initial medical image; Obtaining a first reference sampling grid according to the first standard sampling grid and the first deformation field data; Performing normalization processing on the first reference sampling grid to obtain a first target sampling grid; Sampling the second initial medical image according to the first target sampling grid and a preset sampling interpolation function to obtain the first initial registered image.

4. The method according to claim 1, wherein The step of adjusting the third initial medical image based on the second deformation field data to generate a second initial registered image includes: Obtaining a second standard sampling grid of the third initial medical image; Obtaining a second reference sampling grid according to the second standard sampling grid and the second deformation field data; Performing normalization processing on the second reference sampling grid to obtain a second target sampling grid; Sampling the third initial medical image according to the second target sampling grid and a preset sampling interpolation function to obtain the second initial registered image.

5. The method according to claim 1, wherein The method further includes: Determining whether the loss data converges to a first preset condition; If not, calculating the gradient of the loss function with respect to the segmentation result, and updating the segmentation result by backpropagation according to the gradient of the loss function with respect to the segmentation result.

6. The method according to claim 1, wherein The method further includes: Determining whether the loss data converges to a first preset condition; If not, calculating the gradient of the loss function with respect to the segmentation parameters of the segmentation neural network, and updating the segmentation parameters by backpropagation according to the gradient of the segmentation parameters and a preset global learning rate.

7. The method according to claim 1, wherein The first initial medical image is a T2-weighted imaging medical image, one of the second initial medical image and the third initial medical image is an apparent diffusion coefficient medical image, and the other of the second initial medical image and the third initial medical image is a diffusion-weighted imaging medical image; Or, the first initial medical image is the apparent diffusion coefficient medical image, one of the second initial medical image and the third initial medical image is the diffusion-weighted imaging medical image, and the other of the second initial medical image and the third initial medical image is the T2-weighted imaging medical image; Or, the first initial medical image is a diffusion-weighted imaging medical image, one of the second initial medical image and the third initial medical image is the T2-weighted imaging medical image, and the other of the second initial medical image and the third initial medical image is the apparent diffusion coefficient medical image.

8. A multimodal medical image registration device, characterized in that, It includes: An initial medical image set acquisition module configured to acquire an initial medical image set, where the initial medical image set includes a first initial medical image, a second initial medical image, and a third initial medical image with different modalities; A deformation field data generation module configured to initialize and train the initial medical image set using a registration network, and generate first deformation field data and second deformation field data according to the initialization training, where the first deformation field data is the deformation field data between the first initial medical image and the second initial medical image, and the second deformation field data is the deformation field data between the first initial medical image and the third initial medical image; An initial registration image generation module configured to adjust the second initial medical image based on the first deformation field data to generate a first initial registration image and adjust the third initial medical image based on the second deformation field data to generate a second initial registration image; An image fusion module configured to fuse the first initial registration image, the second initial registration image, and the first initial medical image to obtain a fused medical image; The deformation field data optimization module is configured to perform segmentation training on the fused medical image by using a segmentation neural network, obtain a preset loss function, calculate loss data according to the loss function, the segmentation result obtained from the segmentation training, and the ground truth annotation of the fused medical image, determine whether the loss data converges to a first preset condition, and if not, calculate the gradient of the loss function with respect to the first deformation field data and calculate the gradient of the loss function with respect to the second deformation field data, and update the first deformation field data by backpropagation according to the gradient of the loss function with respect to the first deformation field data and a preset global learning rate, and update the second deformation field data by backpropagation according to the gradient of the loss function with respect to the second deformation field data and the global learning rate.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Medical image registration model training method and equipment

    CN115830016A