Target spot positioning method and system based on multimode image registration fusion
By integrating the registration and fusion of MR images and ultrasound images and establishing a target positioning model, the problem of segmentation accuracy and efficiency being difficult to coexist in medical images is solved, and lightweight and real-time diagnosis of high-precision tumor tissue segmentation is achieved.
Patent Information
- Application Number
- CN202510657279.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-16
AI Technical Summary
In the tissue segmentation process in medical images in the prior art, segmentation accuracy and segmentation efficiency cannot coexist, which leads to an increase in computational complexity and affects segmentation efficiency.
By registering and fusing MR images and ultrasound images, using Unet segmentation and convolutional neural networks to establish a target positioning model, and combining GAN networks and Transformer models for image registration and feature extraction, multimodal target positioning is achieved.
It achieves the goal of maintaining high-precision tumor tissue segmentation while reducing the amount of computation, improving segmentation efficiency, and meeting real-time diagnosis needs.
Smart Images

Figure CN120655707A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target positioning, and in particular to a target positioning method and system based on multi-modal image registration and fusion. Background Art
[0002] Multimodal image fusion can fuse the feature information of each modality to present the tumor and its surrounding tissue more comprehensively, thereby more accurately defining the boundary between the tumor and normal tissue. Therefore, tumor tissue segmentation based on the multimodal fused image can obtain more feature information during the segmentation process, which helps the machine learning model to more accurately identify and segment the tumor with higher precision.
[0003] In the existing technology, segmentation accuracy is improved by performing tissue segmentation in multimodal images. However, this makes the tissue segmentation calculation process complicated and the amount of calculation increased, thereby affecting the segmentation efficiency of tissue segmentation in medical images, resulting in the defect that segmentation accuracy and segmentation efficiency cannot coexist. Summary of the Invention
[0004] The purpose of the present invention is to provide a target positioning method and system based on multimodal image registration and fusion to solve the technical problem in the prior art that the segmentation accuracy and segmentation efficiency in the tissue segmentation process in medical images cannot coexist.
[0005] In order to solve the above technical problems, the present invention specifically provides the following technical solutions:
[0006] A target positioning method based on multimodal image registration and fusion, comprising the following steps:
[0007] Perform registration and fusion of MR images and ultrasound images to obtain multimodal fusion images;
[0008] Perform Unet segmentation in the multimodal fusion image to obtain the tumor area as the multimodal target location;
[0009] Unet segmentation is performed on the MR image and ultrasound image respectively to obtain the tumor area as the target location of the MR image and the target location of the ultrasound image;
[0010] The MR image target position, the ultrasound image target position, and the multimodal target position are nonlinearly mapped using a convolutional neural network to obtain a target positioning model for determining the multimodal target position based on the MR image target position and the ultrasound image target position.
[0011] As a preferred solution of the present invention, the method for acquiring the multimodal fusion image includes:
[0012] Step 1: perform edge detection and principal component analysis on the MR image and the ultrasound image respectively to obtain the MR main structure feature map and the ultrasound main structure feature map;
[0013] Step 2: The MR main structure feature map and the ultrasound main structure feature map are used to calculate the MR registration deformation field and the ultrasound registration deformation field respectively using the generator G in the GAN network;
[0014] Step 3: spatially transform the MR main structure feature map and the ultrasound main structure feature map using the MR registration deformation field and the ultrasound registration deformation field to generate an MR registration image and an ultrasound registration image respectively;
[0015] Step 4: Use the discriminator D in the GAN network to compare the MR registered image and the ultrasound registered image with the MR image and the ultrasound image to perform registration discrimination, and identify the MR registered image and the ultrasound registered image with the best registration effect;
[0016] Step 5: Replace the MR image and ultrasound image in step 1 with the MR registered image and ultrasound registered image with the best registration effect identified by the discriminator D, and repeat steps 1 to 5 until the structural feature similarity between the MR registered image and ultrasound registered image with the best registration effect identified by the discriminator D is maximized to obtain the optimal registered image of the MR image and ultrasound image;
[0017] Step 6: Fusing the optimal registration image of the MR image with the ultrasound image to obtain a first multimodal fusion image, and fusing the optimal registration image of the ultrasound image with the MR image to obtain a second multimodal fusion image.
[0018] As a preferred embodiment of the present invention, the determination of the multimodal target position includes:
[0019] Perform Unet segmentation in the first multimodal fusion image to obtain the tumor area as the first target location;
[0020] Perform Unet segmentation in the second multimodal fusion image to obtain the tumor area as the second target location;
[0021] The first target position and the second target position are weighted averaged to obtain the multimodal target position P best =w1P m1 +w2P m2 , where P best is the multimodal target position, P m1 is the first target position, P m2 is the second target position, w1 is P m1 The weight of P m2 The weight of
[0022] Where A1=1 / (1+|P m1 ―P MR |+|P m1 ―P B |);
[0023] B1=1 / (1+|P m2 ―P MR |+|P m2 ―P B |);
[0024] w1=A1 / (A1+B1);
[0025] w2=B1 / (A1+B1);
[0026] Where A1 is w1 before normalization, B1 is w2 before normalization, P MR is the target position on the MR image, P B is the target position of the ultrasound image, |P m1 ―P MR | for P m1 and P MR The Euclidean distance between |P m1 ―P B | for P m1 and P B The Euclidean distance between |P m2 ―P MR | for P m2 and P MR The Euclidean distance between |P m2 ―P B | for P m2 and P B The Euclidean distance between .
[0027] As a preferred embodiment of the present invention, the method for constructing the target location model includes:
[0028] The target position in the MR image and the target position in the ultrasound image are used as input items of a convolutional neural network, the multimodal target position is used as an output item of the convolutional neural network, and the convolutional neural network is trained to obtain the target positioning model;
[0029] The target location model is P best =CNN(P m1 ,P m2 );
[0030] Where, P best is the multimodal target position, P m1 is the first target position, P m2 is the second target position, and CNN is a convolutional neural network.
[0031] As a preferred solution of the present invention, the method for generating the MR registration deformation field and the ultrasound registration deformation field in step 3 includes:
[0032] The MR main structure feature map PSR_f and the ultrasound main structure feature map PSR_r are input into the generator G to obtain the MR registration deformation field F_f and the ultrasound registration deformation field F_r;
[0033] The generation process of the registration deformation field is:
[0034] (F_f, F_r) = G(PSR_f, PSR_r);
[0035] Where F_f and F_r are the MR registration deformation field and the ultrasound registration deformation field, PSR_f and PSR_r are the MR main structure feature map and the ultrasound main structure feature map, respectively, and G is the generator;
[0036] The generator G is formed by a Transformer model structure combined with graph processing, and the generator G includes two input channels.
[0037] The loss function for training the generator G is:
[0038] L G =λ1L G,adv +λ2L G,MSE +λ3L G,smooth ;
[0039] in,
[0040] Where, L G is the generator loss function, λ1, λ2 and λ3 are L G,adv , L G,MSE and L G,smoot h Penalty coefficient, L G,adv To combat the loss, L G,MSE is the registration structure loss, L G,smooth is the deformation field smoothing loss, D is the output discrimination probability of the discriminator D, D(PSR_f_reg) is the discrimination probability of PSR_f_reg, D(PSR_r_reg) is the discrimination probability of PSR_r_reg, PSR_f_reg is the main structural feature map of the MR registration image, PSR_r_reg is the main structural feature map of the ultrasound registration image, PSR_f and PSR_r are the MR main structural feature map and the ultrasound main structural feature map, respectively. represents the gradient of F_f and F_r, E represents the expectation, and F represents the Frobenius norm.
[0041] As a preferred solution of the present invention, the method for identifying the MR registered image and the ultrasound registered image with the best registration effect in step 4 includes:
[0042] The principal structural feature map PSR_f_reg of the MR registration image and the principal structural feature map PSR_r_reg of the ultrasound registration image are obtained by edge detection and principal component analysis respectively;
[0043] Input the main structural feature map PSR_f_reg of the MR registration image to the first discriminator D, and obtain the MR registration image corresponding to the discrimination result of PSR_f_reg being discriminated as PSR_r as the MR registration image with the best registration effect;
[0044] Input the main structural feature map PSR_r_reg of the ultrasound registration image to the second discriminator D, and obtain the ultrasound registration image corresponding to the discrimination result of PSR_r_reg being discriminated as PSR_f as the ultrasound registration image with the best registration effect;
[0045] The first discriminator D and the second discriminator D are both formed by the Transformer model structure combined with graph processing, and each includes one input channel;
[0046] The loss function for training the first discriminator D and the second discriminator D is:
[0047]
[0048] Where, L D is the loss function of the discriminator D, D(PSR_f_reg) is the discrimination probability of PSR_f_reg, D(PSR_f) is the discrimination probability of PSR_f, D(PSR_r_reg) is the discrimination probability of PSR_r_reg, D(PSR_r) is the discrimination probability of PSR_r, PSR_f_reg is the main structural feature map of the MR registration image, PSR_r_reg is the main structural feature map of the ultrasound registration image, PSR_f and PSR_r are the MR main structural feature map and the ultrasound main structural feature map respectively, E represents the expectation, and F represents the Frobenius norm.
[0049] As a preferred embodiment of the present invention, the method for generating the optimal registration image of the MR image and the ultrasound image in step 5 includes:
[0050] The optimization goal is to maximize the similarity of structural features between the MR registration image and the ultrasound registration image with the best registration effect. The optimization objective function is:
[0051]
[0052] Where M is the optimization target identifier, bset(PSR_f_reg) is the MR registered image with the best registration effect, bset(PSR_r_reg) is the ultrasound registered image with the best registration effect, E represents the expectation, and F represents the Frobenius norm;
[0053] Based on the optimization target M, the registration process of the GAN network that has completed the training of the generator G and the first discriminator D and the second discriminator D is repeated multiple times until the optimization target M is achieved. The MR registration image and the ultrasound registration image with the best registration effect generated by the generator G in the GAN network at this time are used as the optimal registration images of the MR image and the ultrasound image.
[0054] As a preferred embodiment of the present invention, the present invention provides a target positioning system based on multimodal image registration and fusion, which is applied to a target positioning method based on multimodal image registration and fusion. The system includes:
[0055] A registration and fusion unit is used to register and fuse the MR image and the ultrasound image to obtain a multimodal fusion image;
[0056] a segmentation and positioning unit, configured to perform Unet segmentation in the multimodal fusion image to obtain a tumor region as the multimodal target location; and to perform Unet segmentation in the MR image and the ultrasound image respectively to obtain a tumor region as the MR image target location and the ultrasound image target location;
[0057] The model construction unit is used to perform nonlinear relationship mapping on the MR image target position, the ultrasound image target position, and the multimodal target position using a convolutional neural network to obtain a target positioning model for determining the multimodal target position based on the MR image target position and the ultrasound image target position.
[0058] As a preferred solution of the present invention, the method for determining the multimodal target position by the segmentation and positioning unit includes:
[0059] Perform Unet segmentation in the first multimodal fusion image to obtain the tumor area as the first target location;
[0060] Perform Unet segmentation in the second multimodal fusion image to obtain the tumor area as the second target location;
[0061] The first target position and the second target position are weighted averaged to obtain the multimodal target position P best =w1P m1 +w2P m2 , where P best is the multimodal target position, P m1 is the first target position, P m2is the second target position, w1 is P m1 The weight of P m2 The weight of
[0062] Where A1=1 / (1+|P m1 ―P MR |+|P m1 ―P B |);
[0063] B1=1 / (1+|P m2 ―P MR |+|P m2 ―P B |);
[0064] w1=A1 / (A1+B1);
[0065] w2=B1 / (A1+B1);
[0066] Where A1 is w1 before normalization, B1 is w2 before normalization, P MR is the target position on the MR image, P B is the target position of the ultrasound image, |P m1 ―P MR | for P m1 and P MR The Euclidean distance between |P m1 ―P B | for P m1 and P B The Euclidean distance between |P m2 ―P MR | for P m2 and P MR The Euclidean distance between |P m2 ―P B | for P m2 and P B The Euclidean distance between .
[0067] As a preferred embodiment of the present invention, the method for constructing a target positioning model of a model positioning unit includes:
[0068] The target position in the MR image and the target position in the ultrasound image are used as input items of a convolutional neural network, the multimodal target position is used as an output item of the convolutional neural network, and the convolutional neural network is trained to obtain the target positioning model;
[0069] The target location model is P best =CNN(P m1 ,P m2 );
[0070] Where, P bestis the multimodal target position, P m1 is the first target position, P m2 is the second target position, and CNN is a convolutional neural network.
[0071] Compared with the prior art, the present invention has the following beneficial effects:
[0072] The present invention establishes a mapping relationship between the target localization results in the multimodal fusion image of MR images and ultrasound images and the target localization results in MR images and ultrasound images through a convolutional neural network, which can achieve multimodal segmentation results based on a single segmentation result. The target localization accuracy performance is improved while achieving model lightweighting. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other implementation drawings based on the provided drawings without inventive effort.
[0074] Figure 1 A flow chart of a target positioning method based on multimodal image registration and fusion provided in an embodiment of the present invention;
[0075] Figure 2 A block diagram of a target positioning system based on multimodal image registration and fusion provided by an embodiment of the present invention;
[0076] Figure 3 A schematic diagram of a multimodal image real-time registration process provided by an embodiment of the present invention;
[0077] Figure 4 Schematic diagram of the model structure of the generator and discriminator provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0078] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0079] like Figure 1 As shown, the present invention provides a target positioning method based on multimodal image registration and fusion, comprising the following steps:
[0080] Perform registration and fusion of MR images and ultrasound images to obtain multimodal fusion images;
[0081] Perform Unet segmentation in the multimodal fusion image to obtain the tumor area as the multimodal target location;
[0082] Unet segmentation is performed on the MR image and ultrasound image respectively to obtain the tumor area as the target location of the MR image and the target location of the ultrasound image;
[0083] The MR image target position, the ultrasound image target position, and the multimodal target position are nonlinearly mapped using a convolutional neural network to obtain a target positioning model for determining the multimodal target position based on the MR image target position and the ultrasound image target position.
[0084] Different modal images contain different types of tumor tissue information. The fused multimodal image can integrate the characteristic information of each modality to more comprehensively present the tumor and its surrounding tissue, thereby more accurately defining the boundary between the tumor and normal tissue. Therefore, tumor tissue segmentation based on the multimodal fused image can obtain more characteristic information during the segmentation process, helping the machine learning model to more accurately identify and segment the tumor with higher precision. Tumor tissue segmentation improves accuracy with the help of multimodal fusion. On the other hand, because tumor tissue segmentation requires the participation of the multimodal fusion process, the computational complexity of the segmentation process will increase significantly, which will inevitably affect segmentation efficiency.
[0085] To this end, the present invention first aligns and fuses the MR image and the ultrasound image to obtain an accurate multimodal fusion image, which can provide accuracy guarantee for the subsequent tumor tissue segmentation on the multimodal fusion image, so that high-precision tumor tissue segmentation results can be obtained based on the multimodal fusion image.
[0086] Then, while obtaining high-precision segmentation results, the present invention hopes that the tumor segmentation process can be lightweight. By mapping the segmentation results on a single modality image with the high-precision tumor tissue segmentation results obtained from the multimodal fusion image, a high-precision segmentation result on the multimodal image can be directly obtained based on the low-precision segmentation results on the single modality image. In this process, there is no need for the participation of the multimodal fusion process, the computational complexity of the segmentation process will not increase significantly, and the segmentation efficiency will not be affected.
[0087] Therefore, the present invention can obtain high-precision segmentation results during tumor tissue segmentation while keeping the segmentation process lightweight.
[0088] The method for obtaining multimodal fusion images includes:
[0089] Step 1: perform edge detection and principal component analysis on the MR image and the ultrasound image respectively to obtain the MR main structure feature map and the ultrasound main structure feature map;
[0090] Step 2: The MR main structure feature map and the ultrasound main structure feature map are used to calculate the MR registration deformation field and the ultrasound registration deformation field respectively using the generator G in the GAN network;
[0091] Step 3: spatially transform the MR main structure feature map and the ultrasound main structure feature map using the MR registration deformation field and the ultrasound registration deformation field to generate an MR registration image and an ultrasound registration image respectively;
[0092] Step 4: Use the discriminator D in the GAN network to compare the MR registered image and the ultrasound registered image with the MR image and the ultrasound image to perform registration discrimination, and identify the MR registered image and the ultrasound registered image with the best registration effect;
[0093] Step 5: Replace the MR image and ultrasound image in step 1 with the MR registered image and ultrasound registered image with the best registration effect identified by the discriminator D, and repeat steps 1 to 5 until the structural feature similarity between the MR registered image and ultrasound registered image with the best registration effect identified by the discriminator D is maximized to obtain the optimal registered image of the MR image and ultrasound image;
[0094] Step 6: Fusing the optimal registration image of the MR image with the ultrasound image to obtain a first multimodal fusion image, and fusing the optimal registration image of the ultrasound image with the MR image to obtain a second multimodal fusion image.
[0095] In the registration of MR images and ultrasound images, the present invention uses structural features as registration processing information, which can align the same physiological structure on the two images after registration. The purpose of registration is to discover the changes in physiological structures in different modalities. The important structure or main structure can reflect most of the characteristics / main features of the disease lesions, while the detailed structure contains the remaining small parts or secondary features. In order to adapt to the real-time registration of MR images and ultrasound images and improve the timeliness or efficiency of registration, the present invention uses the main features in the multimodal image as the registration focus, weakening or even eliminating the registration focus on the detailed features. Compared with registering all structures, the amount of information processing is reduced, the corresponding information processing time is also reduced, and the efficiency is improved. Of course, a certain accuracy advantage of registering all structures is also sacrificed. Therefore, the present invention uses the main features as the registration focus, achieves the sacrifice of some accuracy advantages to improve efficiency advantages, and meets the timeliness requirements of real-time registration of MR images and ultrasound images.
[0096] The present invention first extracts all structural information from MR images and ultrasound images through an edge detection algorithm (such as the Canny operator or other algorithms with equivalent effects), and then extracts the main structural information from all the structural information through a principal component analysis network through a two-layer convolution operation, which is used as the focus of registration. In this way, the main structure is aligned during the registration process, and the changes of the main structure in different modalities are quickly grasped. Correspondingly, most of the features / main features of the disease lesions are quickly grasped, achieving the expectation of real-time registration to assist in rapid diagnosis of the disease.
[0097] The present invention uses a GAN network to achieve two-directional registration of the main structure information, wherein the generator G in the GAN network generates a deformation field for registration in two directions. One direction is to generate a registration deformation field for registering the MR image to the ultrasound image, which is used to realize the spatial transformation of the MR image toward the ultrasound image to complete the main structure registration on the two modalities. The other direction is to generate a registration deformation field for registering the ultrasound image to the MR image, which is used to realize the spatial transformation of the ultrasound image toward the MR image to complete the main structure registration on the two modalities. The two discriminators D in the GAN network are used to discriminate the main structure registration results formed by the registration in the two directions, respectively, to ensure that the main structure registration output in the two registration directions reaches the optimal registration result. Therefore, after the training is completed, the GAN network can obtain the best MR registration image and ultrasound registration image in the two registration directions, that is, to complete the optimal main structure registration of the MR image to the ultrasound image, and to complete the optimal main structure registration of the ultrasound image to the MR image, ensuring high-precision registration independent of each other in the two directions.
[0098] Furthermore, in order to constrain the two independent high-precision registrations to each other and limit the randomness of each in its own independent registration direction, the present invention establishes an optimization goal, repeats the GAN network registration process cyclically, and performs constrained optimization on the optimal principal structure registration results obtained in each of the two registration directions. Specifically, the principal structure registration results in the two registration directions are re-input into the GAN network as two images to be registered, and the two directions are registered repeatedly until the principal structure registration results in the two registration directions achieve the maximum similarity, ensuring that the registration processes in the two registration directions attract each other, that is, the MR image is converted to the ultrasound image. In the process of achieving the optimization goal, the image registration result approaches the registration result of the ultrasound image to the MR image, thereby pulling the registration direction of the MR image to the ultrasound image toward the registration of the ultrasound image to the MR image. Similarly, the ultrasound image to the MR image registration result approaches the registration result of the MR image to the ultrasound image, thereby pulling the registration direction of the ultrasound image to the MR image toward the registration of the MR image. The two registration directions restrict and constrain each other to ensure that the main structure information registration result obtained can fuse the high precision of the two registration directions to obtain higher precision registration results of MR images and ultrasound images.
[0099] The present invention first extracts all structural information from MR images and ultrasound images using an edge detection algorithm (such as the Canny operator or other algorithms with equivalent effects). Then, a principal component analysis network is used to perform a two-layer convolution operation to extract the main structural information from all the structural information, which is used as the registration focus. The details are as follows:
[0100] The method for extracting the MR main structure feature map and the ultrasound main structure feature map includes:
[0101] The Canny operator is used to perform edge detection on the MR image and the ultrasound image respectively to obtain the MR structure feature map and the ultrasound structure feature map;
[0102] The principal component analysis is performed on the MR structural feature map and the ultrasound structural feature map respectively to obtain the MR main structural feature map and the ultrasound main structural feature map.
[0103] The principal component analysis method for the MR structural feature map and the ultrasound structural feature map includes:
[0104] The MR structural feature map is convolved by the first and second convolutional neural networks in the principal component analysis network, and the output results of the first and second convolutional neural networks are fused through an exponential function to obtain the MR main structural feature map PSR_f;
[0105] The first convolutional neural network and the second convolutional neural network in the principal component analysis network are used to perform convolution processing on the ultrasonic structure feature map respectively, and the output results of the first convolutional neural network and the second convolutional neural network are fused through the exponential function to obtain the ultrasonic main structure feature map PSR_r.
[0106] The setting method of the first layer of convolutional neural network and the second layer of convolutional neural network in the principal component analysis network includes:
[0107] An image block centered on each voxel in the MR structural feature map is selected and de-meaned and vectorized. The vectorized result is then formed into a matrix for PCA processing. The eigenvectors corresponding to the obtained eigenvalues are matrixed to obtain the convolution kernel of the first layer of the convolutional neural network in the principal component analysis network for principal component analysis of the MR structural feature map.
[0108] The image blocks centered on each voxel in the output results of the first layer of the convolutional neural network are de-averaged and vectorized, and the vectorized results are then formed into a matrix for PCA processing. The eigenvectors corresponding to the obtained eigenvalues are matrixed to obtain the convolution kernel of the second layer of the convolutional neural network in the principal component analysis network for principal component analysis of the MR structural feature map;
[0109] An image block centered on each voxel in the ultrasound structural feature map is selected and de-meaned and vectorized. The vectorized result is then formed into a matrix for PCA processing. The eigenvectors corresponding to the obtained eigenvalues are matrixed to obtain the convolution kernel of the first layer of the convolutional neural network in the principal component analysis network for principal component analysis of the ultrasound structural feature map.
[0110] The image blocks centered on each voxel in the output results of the first layer of convolutional neural network are de-averaged and vectorized, and the vectorized results are then formed into a matrix for PCA processing. The eigenvectors corresponding to the obtained eigenvalues are matrixed to obtain the convolution kernel of the second layer of convolutional neural network in the principal component analysis network for principal component analysis of the ultrasonic structural feature map.
[0111] The present invention uses the GAN network to realize the registration of the main structure information in two directions, as follows:
[0112] The generation method of the MR registration deformation field and the ultrasound registration deformation field includes:
[0113] The MR main structure feature map PSR_f and the ultrasound main structure feature map PSR_r are input into the generator G to obtain the MR registration deformation field F_f and the ultrasound registration deformation field F_r;
[0114] The generation process of the registration deformation field is:
[0115] (F_f, F_r) = G(PSR_f, PSR_r);
[0116] Where F_f and F_r are the MR registration deformation field and the ultrasound registration deformation field, PSR_f and PSR_r are the MR main structure feature map and the ultrasound main structure feature map, respectively, and G is the generator;
[0117] The generator G is formed by a Transformer model structure combined with graph processing, and the generator G includes two input channels.
[0118] The loss function for training the generator G is:
[0119] L G =λ1L G,adv +λ2L G,MSE +λ3L G,smooth ;
[0120] in,
[0121] Where, L G is the generator loss function, λ1, λ2 and λ3 are L G,adv , L G,MSE and L G,smoothPenalty coefficient, L G,adv To combat the loss, L G,MSE is the registration structure loss, L G,smooth is the deformation field smoothing loss, D is the output discrimination probability of the discriminator D, D(PSR_f_reg) is the discrimination probability of PSR_f_reg, D(PSR_r_reg) is the discrimination probability of PSR_r_reg, PSR_f_reg is the main structural feature map of the MR registration image, PSR_r_reg is the main structural feature map of the ultrasound registration image, PSR_f and PSR_r are the MR main structural feature map and the ultrasound main structural feature map, respectively. represents the gradient of F_f and F_r, E represents the expectation, and F represents the Frobenius norm.
[0122] Using adversarial loss, registration structure loss, and deformation field smoothness loss as the loss of the generator G can ensure that the generator generates smooth and accurate main structure registration results.
[0123] In addition, the training goal of the generator G is to generate registered images that the discriminator D cannot distinguish. Therefore, the optimization of the generator G is transformed into maximization. Therefore, the training of the GAN network of the present invention is a minimax game process between the generator G and the discriminator D, which is an adversarial learning strategy. The present invention accelerates the convergence speed of the generator G by minimizing the loss between the discrimination probability of the registered image and 1. Therefore, L G,adv The learning goal of this adversarial loss term is: for the registered image generated by the generator G, the discrimination value of the discriminator D is close to 1, that is, the discriminator D is misled into misclassifying the registered image as the reference image (i.e., the main structural feature map of the MR image and the main structural feature map of the ultrasound image in the present invention). This adversarial loss term penalizes the difference between the registered image and the input image of the discriminator D, thereby urging the registered image generated by the GAN network to match / consistent with the main structure of the input image, that is, the main structure in the MR registered image matches / consistent with the main structure in the ultrasound image, and the main structure in the ultrasound registered image matches / consistent with the main structure in the MR image, thereby completing the main structure registration in the two registration directions.
[0124] The two discriminators D in the GAN network of the present invention are used to discriminate the main structure registration results formed by the registration in two directions, and are respectively used to ensure that the main structure registration output in the two registration directions achieves the best registration results, as follows:
[0125] Methods for identifying MR registered images and ultrasound registered images with the best registration results include:
[0126] The principal structural feature map PSR_f_reg of the MR registration image and the principal structural feature map PSR_r_reg of the ultrasound registration image are obtained by edge detection and principal component analysis respectively;
[0127] Input the main structural feature map PSR_f_reg of the MR registration image to the first discriminator D, and obtain the MR registration image corresponding to the discrimination result of PSR_f_reg being discriminated as PSR_r as the MR registration image with the best registration effect;
[0128] Input the main structural feature map PSR_r_reg of the ultrasound registration image to the second discriminator D, and obtain the ultrasound registration image corresponding to the discrimination result of PSR_r_reg being discriminated as PSR_f as the ultrasound registration image with the best registration effect;
[0129] The first discriminator D and the second discriminator D are both formed by the Transformer model structure combined with graph processing, and each includes one input channel;
[0130] The loss function for training the first discriminator D and the second discriminator D is:
[0131]
[0132] Where, L D is the loss function of the discriminator D, D(PSR_f_reg) is the discrimination probability of PSR_f_reg, D(PSR_f) is the discrimination probability of PSR_f, D(PSR_r_reg) is the discrimination probability of PSR_r_reg, D(PSR_r) is the discrimination probability of PSR_r, PSR_f_reg is the main structural feature map of the MR registration image, PSR_r_reg is the main structural feature map of the ultrasound registration image, PSR_f and PSR_r are the MR main structural feature map and the ultrasound main structural feature map respectively, E represents the expectation, and F represents the Frobenius norm.
[0133] In the GAN network setting, D() is the output of the discriminator D, which represents the probability that the input image of the discriminator D is discriminated as the reference image (i.e., the main structural feature map of the MR image and the main structural feature map of the ultrasound image in the present invention). That is, the closer the output probability is to 1, the more likely the discriminant network D is to discriminate the input image as the reference image; the closer the output probability is to 0, the more likely the discriminant network D is to discriminate the input image as the generated registration image.
[0134] The edge detection and principal component analysis used in the process of obtaining the principal structural feature map PSR_f_reg of the MR registration image and the principal structural feature map PSR_r_reg of the ultrasound registration image are consistent with the edge detection and principal component analysis used in the process of obtaining the MR principal structural feature map and the ultrasound principal structural feature map.
[0135] Therefore, after training, the GAN network can obtain the best MR registration images and ultrasound registration images in two registration directions, that is, to complete the best principal structure registration of MR images to ultrasound images, and to complete the best principal structure registration of ultrasound images to MR images, ensuring independent high-precision registration in the two directions.
[0136] In order to constrain the two independent high-precision registrations to each other and limit the randomness of each in its own independent registration direction, this paper establishes an optimization goal, repeats the GAN network registration process cyclically, and performs constrained optimization again on the optimal principal structure registration results obtained in each of the two registration directions, as follows:
[0137] The method for generating the optimal registration image of the MR image and the ultrasound image includes:
[0138] The optimization goal is to maximize the similarity of structural features between the MR registration image and the ultrasound registration image with the best registration effect. The optimization objective function is:
[0139]
[0140] Where M is the optimization target identifier, bset(PSR_f_reg) is the MR registered image with the best registration effect, bset(PSR_r_reg) is the ultrasound registered image with the best registration effect, E represents the expectation, and F represents the Frobenius norm;
[0141] Based on the optimization target M, the registration process of the GAN network that has completed the training of the generator G and the first discriminator D and the second discriminator D is repeated multiple times until the optimization target M is achieved. The MR registration image and the ultrasound registration image with the best registration effect generated by the generator G in the GAN network at this time are used as the optimal registration images of the MR image and the ultrasound image.
[0142] The first discriminator D and the second discriminator D have the same structure.
[0143] Its generator G and discriminator D both use the Transformer model combined with graph processing (such as Figure 4The registration network is implemented as shown in Figure 2, where the G network has two input channels and outputs the generated deformation field, while the D network has only one channel, which is used to determine whether the registration is complete. The graph-based Transformer model is the core of the registration network. Its implementation idea is: by dividing the image into blocks, each block corresponds to a node in the graph, and each node searches for the nearest node between them to form an edge, and then uses graph processing and the Transformer model to complete the feature representation. The Transformer model uses the self-attention mechanism to calculate the correlation between blocks, and reduces the size of the key-value feature K and the content feature V through convolution downsampling to reduce the amount of calculation and ensure real-time requirements. The query feature maintains the feature size unchanged, and adds spatial encoding and position encoding to improve the robustness of the algorithm.
[0144] The present invention utilizes multimodal image registration in two directions and performs corresponding fusion in the two registration directions to obtain multimodal fused images in the two registration directions.
[0145] Determination of multimodal target locations includes:
[0146] Perform Unet segmentation in the first multimodal fusion image to obtain the tumor area as the first target location;
[0147] Perform Unet segmentation in the second multimodal fusion image to obtain the tumor area as the second target location;
[0148] The first target position and the second target position are weighted averaged to obtain the multimodal target position P best =w1P m1 +w2P m2 , where P best is the multimodal target position, P m1 is the first target position, P m2 is the second target position, w1 is P m1 The weight of P m2 The weight of
[0149] Where A1=1 / (1+|P m1 ―P MR |+|P m1 ―P B |);
[0150] B1=1 / (1+|P m2 ―P MR |+|P m2 ―P B |);
[0151] w1=A1 / (A1+B1);
[0152] w2=B1 / (A1+B1);
[0153] Where A1 is w1 before normalization, B1 is w2 before normalization, P MR is the target position on the MR image, P B is the target position of the ultrasound image, |P m1 ―P MR | for P m1 and P MR The Euclidean distance between |P m1 ―P B | for P m1 and P B The Euclidean distance between |P m2 ―P MR | for P m2 and P MR The Euclidean distance between |P m2 ―P B | for P m2 and P B The Euclidean distance between .
[0154] After obtaining the multimodal fusion images in two registration directions, the present invention can obtain a high-precision segmentation result on each of the two multimodal fusion images. In order to further improve the segmentation accuracy based on the multimodal images, the segmentation results of each of the two multimodal fusion images are weighted averaged to achieve the fusion of the segmentation results of each of the two multimodal fusion images to obtain a higher-precision segmentation result.
[0155] The weight of each segmentation result on the two multimodal fusion images is determined by the correlation between the tumor segmentation result and the tumor segmentation result on the two single-modality images (MR image and ultrasound image). The higher the total correlation between each segmentation result on the two multimodal fusion images and the tumor segmentation results on the two single-modality images (MR image and ultrasound image), the more comprehensive the extraction of multimodal fusion information during the multimodal segmentation process. The more comprehensive the multimodal feature information, the higher the credibility of the tumor tissue segmented. Therefore, the segmentation results on the multimodal fusion images are given a high weight. In the present invention, the total correlation between each segmentation result on the multimodal fusion image and the tumor segmentation results on the two single-modality images (MR image and ultrasound image) is quantified using Euclidean distance or other correlation calculation methods.
[0156] The present invention maps and binds the segmentation results on a single modality image with the high-precision tumor tissue segmentation results obtained from the multimodal fusion image. This allows the high-precision segmentation results on the multimodal image to be directly obtained based on the low-precision segmentation results on the single modality image. In this process, there is no need for the participation of the multimodal fusion process, the computational complexity of the segmentation process will not be significantly increased, and the segmentation efficiency will not be affected. The details are as follows:
[0157] The target localization model construction method includes:
[0158] The target positions in the MR image and the ultrasound image are used as input items of the convolutional neural network, and the multimodal target positions are used as output items of the convolutional neural network. The convolutional neural network is trained to obtain a target positioning model.
[0159] The target positioning model is P best =CNN(P m1 ,P m2 );
[0160] Where, P best is the multimodal target position, P m1 is the first target position, P m2 is the second target position, and CNN is a convolutional neural network.
[0161] like Figure 2 As shown, the present invention provides a target positioning system based on multimodal image registration and fusion, which is applied to a target positioning method based on multimodal image registration and fusion. The system includes:
[0162] A registration and fusion unit is used to register and fuse the MR image and the ultrasound image to obtain a multimodal fusion image;
[0163] a segmentation and positioning unit, configured to perform Unet segmentation in the multimodal fusion image to obtain a tumor region as the multimodal target location; and to perform Unet segmentation in the MR image and the ultrasound image respectively to obtain a tumor region as the MR image target location and the ultrasound image target location;
[0164] The model construction unit is used to perform nonlinear relationship mapping on the MR image target position, the ultrasound image target position, and the multimodal target position using a convolutional neural network to obtain a target positioning model for determining the multimodal target position based on the MR image target position and the ultrasound image target position.
[0165] The method for determining the position of a multimodal target by segmenting the positioning unit includes:
[0166] Perform Unet segmentation in the first multimodal fusion image to obtain the tumor area as the first target location;
[0167] Perform Unet segmentation in the second multimodal fusion image to obtain the tumor area as the second target location;
[0168] The first target position and the second target position are weighted averaged to obtain the multimodal target position P best =w1P m1 +w2P m2 , where P best is the multimodal target position, P m1 is the first target position, P m2 is the second target position, w1 is P m1 The weight of P m2 The weight of
[0169] Where A1=1 / (1+|P m1 ―P MR |+|P m1 ―P B |);
[0170] B1=1 / (1+|P m2 ―P MR |+|P m2 ―P B |);
[0171] w1=A1 / (A1+B1);
[0172] w2=B1 / (A1+B1);
[0173] Where A1 is w1 before normalization, B1 is w2 before normalization, P MR is the target position on the MR image, P B is the target position of the ultrasound image, |P m1 ―P MR | for P m1 and P MR The Euclidean distance between |P m1 ―P B | for P m1 and P B The Euclidean distance between |P m2 ―P MR | for P m2 and P MR The Euclidean distance between |P m2 ―P B | for P m2 and P B The Euclidean distance between .
[0174] The method of constructing a unit target positioning model includes:
[0175] The target positions in the MR image and the ultrasound image are used as input items of the convolutional neural network, and the multimodal target positions are used as output items of the convolutional neural network. The convolutional neural network is trained to obtain a target positioning model.
[0176] The target positioning model is P best =CNN(P m1 ,P m2 );
[0177] Where, P best is the multimodal target position, P m1 is the first target position, P m2 is the second target position, and CNN is a convolutional neural network.
[0178] The present invention establishes a mapping relationship between the target localization results in the multimodal fusion image of MR images and ultrasound images and the target localization results in MR images and ultrasound images through a convolutional neural network, which can achieve multimodal segmentation results based on a single segmentation result. The target localization accuracy performance is improved while achieving a lightweight model.
[0179] The above embodiments are merely exemplary embodiments of the present application and are not intended to limit the scope of the present application. The scope of protection of the present application is defined by the claims. Those skilled in the art may make various modifications or equivalent substitutions to the present application within the essence and scope of protection of the present application, and such modifications or equivalent substitutions shall also be deemed to fall within the scope of protection of the present application.
Claims
1. A target positioning method based on multimodal image registration and fusion, characterized in that: The following steps are involved: Perform registration and fusion of MR images and ultrasound images to obtain multimodal fusion images; Perform Unet segmentation in the multimodal fusion image to obtain the tumor area as the multimodal target location; Unet segmentation is performed on the MR image and ultrasound image respectively to obtain the tumor area as the target location of the MR image and the target location of the ultrasound image; The MR image target position, the ultrasound image target position, and the multimodal target position are nonlinearly mapped using a convolutional neural network to obtain a target positioning model for determining the multimodal target position based on the MR image target position and the ultrasound image target position.
2. The target positioning method based on multimodal image registration and fusion according to claim 1, characterized in that: The method for acquiring the multimodal fusion image includes: Step 1: perform edge detection and principal component analysis on the MR image and the ultrasound image respectively to obtain the MR main structure feature map and the ultrasound main structure feature map; Step 2: The MR main structure feature map and the ultrasound main structure feature map are used to calculate the MR registration deformation field and the ultrasound registration deformation field respectively using the generator G in the GAN network; Step 3: spatially transform the MR main structure feature map and the ultrasound main structure feature map using the MR registration deformation field and the ultrasound registration deformation field to generate an MR registration image and an ultrasound registration image respectively; Step 4: Use the discriminator D in the GAN network to compare the MR registered image and the ultrasound registered image with the MR image and the ultrasound image to perform registration discrimination, and identify the MR registered image and the ultrasound registered image with the best registration effect; Step 5: Replace the MR image and ultrasound image in step 1 with the MR registered image and ultrasound registered image with the best registration effect identified by the discriminator D, and repeat steps 1 to 5 until the structural feature similarity between the MR registered image and ultrasound registered image with the best registration effect identified by the discriminator D is maximized to obtain the optimal registered image of the MR image and ultrasound image; Step 6: Fusing the optimal registration image of the MR image with the ultrasound image to obtain a first multimodal fusion image, and fusing the optimal registration image of the ultrasound image with the MR image to obtain a second multimodal fusion image.
3. The target positioning method based on multimodal image registration and fusion according to claim 1, characterized in that: The determination of the multimodal target position includes: Perform Unet segmentation in the first multimodal fusion image to obtain the tumor area as the first target location; Perform Unet segmentation in the second multimodal fusion image to obtain the tumor area as the second target location; The first target position and the second target position are weighted averaged to obtain the multimodal target position P best =w1P m1 +w2P m2 , where P best is the multimodal target position, P m1 is the first target position, P m2 is the second target position, w1 is P m1 The weight of P m2 The weight of Where A1=1 / (1+|P m1 ―P MR |+|P m1 ―P B |); B1=1 / (1+|P m2 ―P MR |+|P m2 ―P B |); w1=A1 / (A1+B1); w2=B1 / (A1+B1); Where A1 is w1 before normalization, B1 is w2 before normalization, P MR is the target position on the MR image, P B is the target position of the ultrasound image, |P m1 ―P MR | for P m1 and P MR The Euclidean distance between |P m1 ―P B | for P m1 and P B The Euclidean distance between |P m2 ―P MR | for P m2 and P MR The Euclidean distance between |P m2 ―P B | for P m2 and P B The Euclidean distance between .
4. The target positioning method based on multimodal image registration and fusion according to claim 3, characterized in that: The method for constructing the target location model includes: The target position in the MR image and the target position in the ultrasound image are used as input items of a convolutional neural network, the multimodal target position is used as an output item of the convolutional neural network, and the convolutional neural network is trained to obtain the target positioning model; The target location model is P best =CNN(P m1 ,P m2 ); Where, P best is the multimodal target position, P m1 is the first target position, P m2 is the second target position, and CNN is a convolutional neural network.
5. The target positioning method based on multimodal image registration and fusion according to claim 2, characterized in that: The method for generating the MR registration deformation field and the ultrasound registration deformation field in step 3 includes: The MR main structure feature map PSR_f and the ultrasound main structure feature map PSR_r are input into the generator G to obtain the MR registration deformation field F_f and the ultrasound registration deformation field F_r; The generation process of the registration deformation field is: (F_f, F_r) = G(PSR_f, PSR_r); Where F_f and F_r are the MR registration deformation field and the ultrasound registration deformation field, PSR_f and PSR_r are the MR main structure feature map and the ultrasound main structure feature map, respectively, and G is the generator; The generator G is formed by the Transformer model structure combined with graph processing, and the generator G includes two input channels; The loss function for training the generator G is: L G =λ1L G,adv +λ2L G,MSE +λ3L G,smooth ; in, Where, L G is the generator loss function, λ1, λ2 and λ3 are L G,adv , L G,MSE and L G,smooth Penalty coefficient, L G,adv To combat the loss, L G,MSE is the registration structure loss, L G,smooth is the deformation field smoothing loss, D is the output discrimination probability of the discriminator D, D(PSR_f_reg) is the discrimination probability of PSR_f_reg, D(PSR_r_reg) is the discrimination probability of PSR_r_reg, PSR_f_reg is the main structural feature map of the MR registration image, PSR_r_reg is the main structural feature map of the ultrasound registration image, PSR_f and PSR_r are the MR main structural feature map and the ultrasound main structural feature map, respectively. represents the gradient of F_f and F_r, E represents the expectation, and F represents the Frobenius norm.
6. The target positioning method based on multimodal image registration and fusion according to claim 5, characterized in that: The method for identifying the MR registered image and the ultrasound registered image with the best registration effect in step 4 includes: The principal structural feature map PSR_f_reg of the MR registration image and the principal structural feature map PSR_r_reg of the ultrasound registration image are obtained by edge detection and principal component analysis respectively; Input the main structural feature map PSR_f_reg of the MR registration image to the first discriminator D, and obtain the MR registration image corresponding to the discrimination result of PSR_f_reg being discriminated as PSR_r as the MR registration image with the best registration effect; Input the main structural feature map PSR_r_reg of the ultrasound registration image to the second discriminator D, and obtain the ultrasound registration image corresponding to the discrimination result of PSR_r_reg being discriminated as PSR_f as the ultrasound registration image with the best registration effect; The first discriminator D and the second discriminator D are both formed by the Transformer model structure combined with graph processing, and each includes one input channel; The loss function for training the first discriminator D and the second discriminator D is: Where, L D is the loss function of the discriminator D, D(PSR_f_reg) is the discrimination probability of PSR_f_reg, D(PSR_f) is the discrimination probability of PSR_f, D(PSR_r_reg) is the discrimination probability of PSR_r_reg, D(PSR_r) is the discrimination probability of PSR_r, PSR_f_reg is the main structural feature map of the MR registration image, PSR_r_reg is the main structural feature map of the ultrasound registration image, PSR_f and PSR_r are the MR main structural feature map and the ultrasound main structural feature map respectively, E represents the expectation, and F represents the Frobenius norm.
7. The target positioning method based on multimodal image registration and fusion according to claim 6, characterized in that: The method for generating the optimal registration image of the MR image and the ultrasound image in step 5 includes: The optimization goal is to maximize the similarity of structural features between the MR registration image and the ultrasound registration image with the best registration effect. The optimization objective function is: Where M is the optimization target identifier, best(PSR_f_reg) is the MR registered image with the best registration effect, bset(PSR_r_reg) is the ultrasound registered image with the best registration effect, E represents the expectation, and F represents the Frobenius norm; Based on the optimization target M, the registration process of the GAN network that has completed the training of the generator G and the first discriminator D and the second discriminator D is repeated multiple times until the optimization target M is achieved. The MR registration image and the ultrasound registration image with the best registration effect generated by the generator G in the GAN network at this time are used as the optimal registration images of the MR image and the ultrasound image.
8. A target positioning system based on multimodal image registration and fusion, characterized in that: A target positioning method based on multimodal image registration and fusion as described in any one of claims 1 to 7, the system comprising: A registration and fusion unit is used to register and fuse the MR image and the ultrasound image to obtain a multimodal fusion image; a segmentation and positioning unit, configured to perform Unet segmentation in the multimodal fusion image to obtain a tumor region as the multimodal target location; and to perform Unet segmentation in the MR image and the ultrasound image respectively to obtain a tumor region as the MR image target location and the ultrasound image target location; The model construction unit is used to perform nonlinear relationship mapping on the MR image target position, the ultrasound image target position, and the multimodal target position using a convolutional neural network to obtain a target positioning model for determining the multimodal target position based on the MR image target position and the ultrasound image target position.
9. The target positioning system based on multimodal image registration and fusion according to claim 8, characterized in that: The method for determining the multimodal target position by the segmentation and positioning unit includes: Perform Unet segmentation in the first multimodal fusion image to obtain the tumor area as the first target location; Perform Unet segmentation in the second multimodal fusion image to obtain the tumor area as the second target location; The first target position and the second target position are weighted averaged to obtain the multimodal target position P best =w1P m1 +w2P m2 , where P best is the multimodal target position, P m1 is the first target position, P m2 is the second target position, w1 is P m1 The weight of P m2 The weight of Where A1=1 / (1+|P m1 ―P MR |+|P m1 ―P B |); B1=1 / (1+|P m2 ―P MR |+|P m2 ―P B |); w1=A1 / (A1+B1); w2=B1 / (A1+B1); Where A1 is w1 before normalization, B1 is w2 before normalization, P MR is the target position on the MR image, P B is the target position of the ultrasound image, |P m1 ―P MR | for P m1 and P MR The Euclidean distance between |P m1 ―P B | for P m1 and P B The Euclidean distance between |P m2 ―P MR | for P m2 and P MR The Euclidean distance between |P m2 ―P B | for P m2 and P B The Euclidean distance between .
10. The target positioning system based on multimodal image registration and fusion according to claim 9, characterized in that: The method for constructing a target positioning model of a model positioning unit includes: The target position in the MR image and the target position in the ultrasound image are used as input items of a convolutional neural network, the multimodal target position is used as an output item of the convolutional neural network, and the convolutional neural network is trained to obtain the target positioning model; The target location model is P best =CNN(P m1 ,P m2 ); Where, P best is the multimodal target position, P m1 is the first target position, P m2 is the second target position, and CNN is a convolutional neural network.