Fluorescence-Magnetic Particle Image Fusion Method and Training Method of Multimodal Image Fusion Model
The multimodal image fusion model fusion FMI and MPI solves the problem of insufficient imaging depth and resolution in single-modal imaging, realizes high-precision tumor imaging, and promotes the application of FMI and MPI in the field of biomedical science.
Patent Information
- Application Number
- CN202510135612.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The existing fluorescent molecular imaging (FMI) and magnetic particle imaging (MPI) technologies have problems with insufficient imaging depth or insufficient resolution in single-modal imaging, which limits their clinical applications.
The multimodal image fusion method is adopted, and the trained multimodal image fusion model is used to fuse the two-dimensional near-infrared fluorescence images and three-dimensional magnetic particle tomography images. Through two-dimensional convolution and three-dimensional convolution processing, combined with the adaptive cross attention mechanism, the fusion and filter convolution of multi-scale feature maps are realized to improve imaging quality.
It significantly improves the positioning accuracy and spatial resolution of imaging, makes up for the shortcomings of FMI in imaging depth, improves the resolution of MPI, and achieves more accurate and comprehensive imaging of target areas such as tumors.
Smart Images

Figure CN119991468B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, specifically to the technical field of multimodal medical image fusion, and more specifically to a fluorescence-magnetic particle image fusion method and a training method for a multimodal image fusion model. Background Art
[0002] Fluorescence molecular imaging (FMI) is an emerging molecular imaging technology with the advantages of high sensitivity and high resolution, capable of non-invasive detection of tumors. Fluorescence molecular tomography (FMT) makes up for the deficiency that FMI can only perform two-dimensional imaging and cannot provide three-dimensional spatial information of tumors. It establishes a photon propagation model in the body and inversely solves the three-dimensional spatial distribution of fluorescence molecular probes in the body based on the surface fluorescence obtained by FMI, thereby reconstructing the spatial information of tumors. FMT can reconstruct light sources shallowly distributed in biological tissues with high sensitivity and high resolution. However, due to the strong absorption and scattering of photons in tissues, its imaging depth is limited, and it can only image tumors on the surface. And since the actually collectable data is limited to the surface of the organism, its reconstruction problem has strong ill-posedness and needs to be solved by setting reasonable regularization priors according to experience. The above problems all hinder the clinical application of FMT.
[0003] Magnetic particle imaging (MPI) is an emerging molecular imaging technology used to visualize the spatial distribution of superparamagnetic iron oxide nanoparticles in organisms, with advantages such as no imaging depth limitation, linear quantization, high sensitivity, no background signal interference, and no ionizing radiation hazard, and has broad biomedical application prospects. However, MPI has a low spatial resolution and there is a problem of resolution anisotropy, which seriously affects the imaging accuracy and imaging quality and hinders the clinical application of MPI.
[0004] In view of the technical problems existing in single-modal imaging of FMI, FMT, and MPI, it is necessary to provide a multimodal imaging technical solution to solve technical problems such as insufficient resolution or insufficient imaging depth existing in single-modal imaging technical solutions. Summary of the Invention
[0005] In view of the above problems, the present invention provides a fluorescence-magnetic particle image fusion method and a training method for a multimodal image fusion model that improve the accuracy and quality of multimodal medical image fusion.
[0006] According to the first aspect of the present invention, a fluorescence-magnetic particle image fusion method is provided, including:
[0007] Performing two-dimensional convolution processing on the registered two-dimensional near-infrared fluorescence image by using the trained multi-modal image fusion model to obtain a multi-channel near-infrared fluorescence extended feature map, and initially fusing the multi-channel near-infrared fluorescence extended feature map with the registered three-dimensional magnetic particle tomography image to obtain a three-dimensional near-infrared fluorescence extended feature map;
[0008] Performing multiple rounds of three-dimensional convolution processing on the three-dimensional near-infrared fluorescence extended feature map by using the trained multi-modal image fusion model to obtain a multi-scale near-infrared fluorescence feature map, and performing multiple rounds of three-dimensional convolution processing on the registered three-dimensional magnetic particle tomography image to obtain a multi-scale magnetic particle feature map;
[0009] Based on the adaptive cross-attention mechanism, using the trained multi-modal image fusion model to perform multiple rounds of feature map fusion on the multi-scale near-infrared fluorescence feature map and the multi-scale magnetic particle feature map at the same scale, and performing filtering convolution processing on the result of the multiple rounds of feature map fusion to obtain a three-dimensional fluorescence-magnetic particle fusion image.
[0010] According to an embodiment of the present invention, the above-mentioned initially fusing the multi-channel near-infrared fluorescence extended feature map with the registered three-dimensional magnetic particle tomography image to obtain a three-dimensional near-infrared fluorescence extended feature map includes:
[0011] Using the pre-fusion module of the trained multi-modal image fusion model to perform channel-by-channel multiplication processing on the maximum value of the registered three-dimensional magnetic particle tomography image in each channel and the multi-channel near-infrared fluorescence extended feature map to obtain a three-dimensional near-infrared fluorescence extended feature map.
[0012] According to an embodiment of the present invention, the above-mentioned performing multiple rounds of three-dimensional convolution processing on the three-dimensional near-infrared fluorescence extended feature map by using the trained multi-modal image fusion model to obtain a multi-scale near-infrared fluorescence feature map includes:
[0013] Performing initial three-dimensional convolution processing on the three-dimensional near-infrared fluorescence extended feature map by using the first near-infrared fluorescence convolution module of the trained multi-modal image fusion model, and performing initial activation processing on the result of the initial three-dimensional convolution processing to obtain an initial activation result;
[0014] Performing secondary three-dimensional convolution processing on the initial activation result by using the first near-infrared fluorescence convolution module, and performing secondary activation processing on the result of the secondary three-dimensional convolution processing to obtain a first-scale near-infrared fluorescence feature map.
[0015] According to an embodiment of the present invention, the above-mentioned multi-round three-dimensional convolution processing of the three-dimensional near-infrared fluorescence extended feature map by the trained multi-modal image fusion model to obtain the multi-scale near-infrared fluorescence feature map further includes:
[0016] Using the second near-infrared fluorescence convolution module of the trained multi-modal image fusion model to perform the same processing on the first-scale near-infrared fluorescence feature map as that on the three-dimensional near-infrared fluorescence extended feature map to obtain the second-scale near-infrared fluorescence feature map;
[0017] Using the third near-infrared fluorescence convolution module of the trained multi-modal image fusion model to perform the same processing on the second-scale near-infrared fluorescence feature map as that on the three-dimensional near-infrared fluorescence extended feature map to obtain the third-scale near-infrared fluorescence feature map.
[0018] According to an embodiment of the present invention, the above-mentioned multi-round feature map fusion on the same scale of the multi-scale near-infrared fluorescence feature map and the multi-scale magnetic particle feature map based on the adaptive cross-attention mechanism using the trained multi-modal image fusion model includes:
[0019] Using the first cross-attention mechanism module of the trained multi-modal image fusion model to fuse the first-scale near-infrared fluorescence feature map and the first-scale magnetic particle feature map in the multi-scale magnetic particle feature map to obtain the first fusion feature map;
[0020] Using the second cross-attention mechanism module of the trained multi-modal image fusion model to fuse the second-scale near-infrared fluorescence feature map and the second-scale magnetic particle feature map in the multi-scale magnetic particle feature map to obtain the second fusion feature map;
[0021] Using the third cross-attention mechanism module of the trained multi-modal image fusion model to fuse the third-scale near-infrared fluorescence feature map and the third-scale magnetic particle feature map in the multi-scale magnetic particle feature map to obtain the third fusion feature map;
[0022] Using the trained multi-modal image fusion model to first connect the third fusion feature map with the third-scale magnetic particle feature map, and perform first upsampling on the fused feature map after the first connection to obtain the first upsampled fused feature map;
[0023] Using the trained multi-modal image fusion model to secondarily connect the first upsampled fused feature map with the second fusion feature map, and perform second upsampling on the fused feature map after the secondary connection to obtain the second upsampled fused feature map;
[0024] Using the trained multi-modal image fusion model, the second upsampled fusion image is concatenated with the first fusion feature map three times, and the fused feature map after the three concatenations is upsampled for the third time to obtain the multi-round feature map fusion result.
[0025] According to an embodiment of the present invention, the first cross-attention mechanism module of the above-mentioned trained multi-modal image fusion model fuses the first-scale magnetic particle feature map in the first-scale near-infrared fluorescence feature map and the multi-scale magnetic particle feature map, and the steps for obtaining the first fusion feature map include:
[0026] Data rearrangement is respectively performed on the first-scale near-infrared fluorescence feature map and the first-scale magnetic particle feature map to obtain the rearranged first-scale near-infrared fluorescence feature map and the rearranged first-scale magnetic particle feature map;
[0027] Linear transformation is respectively performed on the rearranged first-scale near-infrared fluorescence feature map and the rearranged first-scale magnetic particle feature map to obtain the first near-infrared fluorescence multi-value matrix and the first magnetic particle multi-value matrix;
[0028] Multi-head attention mechanism partitioning is respectively performed on the rearranged first-scale near-infrared fluorescence feature map and the rearranged first-scale magnetic particle feature map;
[0029] When the multi-head attention mechanism partitioning is completed, the first near-infrared fluorescence multi-value matrix and the first magnetic particle multi-value matrix are subjected to multi-head adaptive cross-attention operation to obtain the first near-infrared fluorescence multi-head attention result and the first magnetic particle multi-head attention result;
[0030] Data rearrangement is respectively performed on the first near-infrared fluorescence multi-head attention result and the first magnetic particle multi-head attention result to obtain the rearranged first near-infrared fluorescence multi-head attention result and the rearranged first magnetic particle multi-head attention result;
[0031] The rearranged first near-infrared fluorescence multi-head attention result is concatenated with the rearranged first-scale near-infrared fluorescence feature map to obtain the first near-infrared fluorescence cross-attention feature;
[0032] The rearranged first magnetic particle multi-head attention result and the rearranged first-scale magnetic particle feature map are concatenated to obtain the first magnetic particle cross-attention feature, and the first near-infrared fluorescence cross-attention feature and the first magnetic particle cross-attention feature are subjected to weighted operation to obtain the first fusion feature map.
[0033] According to an embodiment of the present invention, when the multi-head attention mechanism partitioning is completed, the steps of performing multi-head adaptive cross-attention operation on the first near-infrared fluorescence multi-value matrix and the first magnetic particle multi-value matrix to obtain the first near-infrared fluorescence multi-head attention result and the first magnetic particle multi-head attention result include:
[0034] Transpose the key-value matrix in the first magnetic particle multi-value matrix and multiply it with the query matrix in the first near-infrared fluorescence multi-value matrix, and perform normalization activation processing on the result of the matrix multiplication to obtain the first near-infrared fluorescence attention score;
[0035] Multiply the first near-infrared fluorescence attention score with the value matrix in the first magnetic particle multi-value matrix, and perform layer normalization on the result of the multiplication to obtain the first near-infrared fluorescence multi-head attention result;
[0036] Transpose the key-value matrix in the first near-infrared fluorescence multi-value matrix and multiply it with the query matrix in the first magnetic particle multi-value matrix, and perform normalization activation processing on the result of the matrix multiplication to obtain the first magnetic particle attention score;
[0037] Multiply the first magnetic particle attention score with the value matrix in the first near-infrared fluorescence multi-value matrix, and perform layer normalization on the result of the multiplication to obtain the first magnetic particle multi-head attention result.
[0038] According to the second aspect of the present invention, there is provided a training method for a multi-modal image fusion model, which is applied to the above fluorescence-magnetic particle image fusion method, and is characterized in that it includes:
[0039] Based on the photon propagation model, through linear interpolation operation and Gaussian noise random addition operation, obtain a simulated two-dimensional near-infrared fluorescence image set of a three-dimensional tumor phantom set in a standardized space;
[0040] Perform three-dimensional convolution processing on the point spread function of the three-dimensional magnetic particle tomography image and the image set of the three-dimensional tumor phantom to obtain a simulated three-dimensional magnetic particle tomography image set;
[0041] Use the multi-modal image fusion model to perform multi-modal fusion between the simulated two-dimensional near-infrared fluorescence image set and the simulated three-dimensional magnetic particle tomography image set to obtain a predicted fusion map;
[0042] Use the training loss function to calculate the loss value between the predicted fusion map and the ground truth label corresponding to the predicted fusion map, and use the loss value to update the parameters of the multi-modal image fusion model;
[0043] Iteratively perform multi-modal fusion operation between images, loss value calculation operation, and parameter update operation until all simulated images in the simulated two-dimensional near-infrared fluorescence image set and the simulated three-dimensional magnetic particle tomography image set are processed, and obtain a trained multi-modal image fusion model.
[0044] According to an embodiment of the present invention, the above-mentioned simulation two-dimensional near-infrared fluorescence image set of the three-dimensional tumor phantom set in the standardized space obtained through the linear interpolation operation and the Gaussian noise random addition operation based on the photon propagation model includes:
[0045] A three-dimensional tumor phantom is set in the target tissue mapped into the standardized space, and the pixel values of the three-dimensional tumor phantom are converted into the concentration of the dual-mode probe;
[0046] The concentration distribution of the dual-mode probe is discretized into a concentration vector, and Gaussian noise is randomly added to the concentration vector to obtain a concentration vector with noise;
[0047] Forward calculation is performed on the target tissue using the photon propagation model to obtain the system matrix of fluorescence imaging, and the system matrix of fluorescence imaging is operated with the vector with noise to obtain the fluorescence intensity distribution vector;
[0048] Linear interpolation operation is performed on the fluorescence intensity distribution vector, and Gaussian noise random addition operation is performed on the fluorescence intensity distribution vector after the linear interpolation operation to obtain the simulation two-dimensional near-infrared fluorescence image set.
[0049] According to an embodiment of the present invention, the above-mentioned trained multi-modal image fusion model includes a two-dimensional near-infrared fluorescence image encoder, a three-dimensional magnetic particle tomography image encoder, a pre-fusion module, a first adaptive cross-attention mechanism fusion module, a second adaptive cross-attention mechanism fusion module, a third adaptive cross-attention mechanism fusion module, and an image fusion decoder;
[0050] Among them, the two-dimensional near-infrared fluorescence image encoder includes a two-dimensional convolutional near-infrared fluorescence expansion module, a first near-infrared fluorescence three-dimensional convolutional module, a second near-infrared fluorescence three-dimensional convolutional module, and a third near-infrared fluorescence three-dimensional convolutional module;
[0051] Among them, the three-dimensional magnetic particle tomography image encoder includes a first magnetic particle three-dimensional convolutional module, a second magnetic particle three-dimensional convolutional module, and a third magnetic particle three-dimensional convolutional module;
[0052] Among them, the first adaptive cross-attention mechanism fusion module, the second adaptive cross-attention mechanism fusion module, and the third adaptive cross-attention mechanism fusion module all include an attention mechanism layer, a multi-layer perceptron layer, a residual connection layer, and multiple linear transformation layers;
[0053] Among them, the image fusion decoder includes a filtering convolutional module and multiple up-sampling three-dimensional convolutional modules.
[0054] The third aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0055] The fourth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0056] The present invention fuses FMI and MPI of different modalities through a trained multi-modal image fusion model, making full use of the advantages of high resolution and high sensitivity of FMI in shallow tissues and the characteristics of MPI without imaging depth limitation and linear quantization. This multi-modal fusion method not only makes up for the deficiency of FMI in imaging depth, but also significantly improves the spatial resolution of MPI, thus achieving more accurate and comprehensive imaging of target areas such as tumors. MPI can provide prior information on the location of tumors, while FMI can provide high-quality detailed information. The present invention effectively fuses these two kinds of information through a multi-modal image fusion model based on deep learning methods, significantly improving the positioning accuracy of imaging and promoting the application of FMI technology and MPI technology in the biomedical field. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features and advantages of the present invention will become clearer. In the drawings:
[0058] Figure 1 is a schematic diagram of a multi-modal medical image of a mouse to be measured according to an embodiment of the present invention;
[0059] Figure 2 is a flowchart of a fluorescence-magnetic particle image fusion method according to an embodiment of the present invention;
[0060] Figure 3 is a schematic structural diagram of a CAUnet model according to an embodiment of the present invention;
[0061] Figure 4 is a schematic structural diagram of a pre-fusion module of a CAUnet model according to an embodiment of the present invention;
[0062] Figure 5 is a schematic structural diagram of a downsampling convolution module according to an embodiment of the present invention;
[0063] Figure 6 is a data processing flowchart of an adaptive cross-attention mechanism fusion module according to an embodiment of the present invention;
[0064] Figure 7Schematic diagram of the upsampling module according to an embodiment of the present invention;
[0065] Figure 8 Flowchart of the training method of the multi-modal image fusion model according to an embodiment of the present invention;
[0066] Figure 9 Schematic diagram of the simulated multi-modal image of the three-dimensional tumor phantom according to an embodiment of the present invention;
[0067] Figure 10 Schematic diagram of the three-dimensional optical-magnetic fusion image of the CAUnet model according to an embodiment of the present invention;
[0068] Figure 11 Schematically shows a block diagram of an electronic device suitable for implementing the fluorescence-magnetic particle image fusion method and the training method of the multi-modal image fusion model according to an embodiment of the present invention. Detailed implementation manners
[0069] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0070] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0071] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0072] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).
[0073] Figure 1It is a schematic diagram of multi-modal medical images of a mouse to be measured according to an embodiment of the present invention.
[0074] As Figure 1 shown, the MPI image, near-infrared FMI image, and CT image of the mouse to be measured can display relevant information on the pathological region of the mouse to be measured at different angles. Among them, the three-dimensional computed tomography (CT) image has characteristics such as high density resolution and high image clarity, contains rich position information, and is often used for the registration of other medical images; fluorescence molecular imaging (FMI) has advantages such as high sensitivity and high resolution, and three-dimensional magnetic particle imaging (MPI) has advantages such as strong depth penetration; FMT can provide three-dimensional information on the superficial layer of the body surface; however, the above-mentioned medical images also have their own disadvantages in single-modal imaging. Due to the problem of insufficient imaging depth in single-modal imaging for FMI and FMT, and technical problems such as insufficient resolution and imaging depth in MPI, it is necessary to provide a multi-modal imaging technical solution that integrates at least two of the modal information of FMI, FMT, or MPI. Multi-modal image fusion can make full use of the image information of different modalities, integrate them together, and improve the quality and information content of the images. For example, fusing FMI and MPI, fully combining the advantages of high resolution and high sensitivity of FMI in the shallow layer with the advantage of no imaging depth limitation of MPI, MPI provides prior knowledge of tumor location, and FMI provides high-quality detailed information, and the imaging spatial resolution and positioning accuracy are improved through deep learning methods for fusion.
[0075] In order to solve at least one of the existing technologies, an embodiment of the present invention provides a fluorescence-magnetic particle image fusion method and a training method for a multi-modal image fusion model.
[0076] Figure 2 It is a flowchart of the fluorescence-magnetic particle image fusion method according to an embodiment of the present invention.
[0077] As Figure 2 shown, the fluorescence-magnetic particle image fusion method of this embodiment includes operation S210 to operation S230.
[0078] In operation S210, the registered two-dimensional near-infrared fluorescence image is subjected to two-dimensional convolution processing using the trained multi-modal image fusion model to obtain a multi-channel near-infrared fluorescence extended feature map, and the multi-channel near-infrared fluorescence extended feature map is initially fused with the registered three-dimensional magnetic particle tomography image to obtain a three-dimensional near-infrared fluorescence extended feature map.
[0079] The above-mentioned trained multi-modal image fusion model is constructed based on the Unet neural network with a multiple adaptive cross-attention mechanism, that is, the CAUnet neural network.
[0080] Both the above two-dimensional near-infrared fluorescence image and three-dimensional magnetic particle tomography image are images used to characterize the target pathological region (such as a tumor) of an organism.
[0081] After obtaining the two-dimensional near-infrared fluorescence image and three-dimensional magnetic particle tomography image of the target pathological region, it is necessary to register the above images and perform multimodal image fusion using the registered images.
[0082] The above operation S210 is used to expand the feature map of the two-dimensional near-infrared fluorescence image and convert the two-dimensional image into a three-dimensional image for subsequent fusion with the three-dimensional magnetic particle tomography image.
[0083] In operation S220, the trained multimodal image fusion model performs multiple rounds of three-dimensional convolution processing on the three-dimensional near-infrared fluorescence extended feature map to obtain a multi-scale near-infrared fluorescence feature map, and performs multiple rounds of three-dimensional convolution processing on the registered three-dimensional magnetic particle tomography image to obtain a multi-scale magnetic particle feature map.
[0084] In operation S230, based on the adaptive cross-attention mechanism, the trained multimodal image fusion model performs multiple rounds of feature map fusion on the multi-scale near-infrared fluorescence feature map and the multi-scale magnetic particle feature map at the same scale, and performs filtering convolution processing on the result of the multiple rounds of feature map fusion to obtain a three-dimensional fluorescence-magnetic particle fusion image.
[0085] It should be specifically noted that the purpose of the present invention is to achieve the fusion of multimodal medical images and provide high-quality medical images of the target pathological region, rather than directly obtaining the diagnostic information or health status of the target case region, but only obtaining and processing intermediate information from the target case region; at the same time, the above operations S210~S230 and other operations of the embodiments of the present invention all belong to information processing operations implemented by devices such as computers.
[0086] The present invention fuses different modalities of FMI and MPI through a trained multimodal image fusion model, making full use of the advantages of high resolution and high sensitivity of FMI in shallow tissues and the characteristics of MPI without imaging depth limitation and linear quantization. This multimodal fusion method not only makes up for the deficiency of FMI in imaging depth but also significantly improves the spatial resolution of MPI, thus achieving more accurate and comprehensive imaging of target regions such as tumors. MPI can provide prior information on the location of tumors, while FMI can provide high-quality detailed information. The present invention effectively fuses these two types of information through a multimodal image fusion model based on deep learning methods, significantly improving the positioning accuracy of imaging and promoting the application of FMI technology and MPI technology in the field of biomedicine.
[0087] Before performing the above operations S210 to S230, it is necessary to obtain a two-dimensional near-infrared fluorescence image, a three-dimensional magnetic particle tomography image, and a three-dimensional computed tomography (CT) image of the target pathological region of the organism (such as a tumor region), and register the above medical images.
[0088] The following further details the acquisition process of the above multi-modal medical images and the registration process of the multi-modal medical images through specific embodiments.
[0089] It should be specifically noted that before obtaining the above medical images of the target pathological region of the organism, authorization from the organism (such as the patient or the patient's close relative) or the owner of the organism was obtained to collect the above medical images of the target pathological region; and with the permission of the organism or the owner of the organism, the above medical images were processed. The relevant processes strictly comply with the requirements of laws and regulations, strict confidentiality measures were taken, and public order and good customs were not violated. A corresponding operation entry was provided for the organism or the owner of the organism to choose to authorize or reject.
[0090] Taking the target pathological region of the organism as the tumor and its surrounding tissue region as an example, a two-dimensional near-infrared fluorescence image on the body surface containing tumor information, a three-dimensional MPI tomography image, and a CT image containing anatomical structure information of the surrounding tissues and organs of the tumor are obtained by injecting a fluorescence / magnetic particle dual-modal probe into the tested organism.
[0091] Taking the tumor and its surrounding tissue region as the region of interest, map its CT image and MPI image into the internal standardized imaging space (SIS). The SIS is constructed in advance according to the task requirements and should ensure that it can accommodate the region of interest.
[0092] Map the two-dimensional near-infrared fluorescence image on the body surface to the surface of the SIS, match the surface fluorescence image with the surface of the CT volume data, and crop the fluorescence image to retain the region mapped to the surface of the SIS.
[0093] In some preferred embodiments, map the CT image into the internal discrete SIS, that is: use the central coordinates of the discrete SIS as the center of the imaging space of the CT image; take each pixel of the preprocessed CT image as a voxel point, obtain the grid node closest to the current voxel point in the discrete SIS, and assign the organ attribute corresponding to the current voxel point to the grid node; traverse the voxel points corresponding to each pixel and map the preprocessed CT image into the internal SIS.
[0094] In some preferred embodiments, the three-dimensional MPI tomographic image is mapped inside the discretized SIS, that is: registration reference points are set, and the imaging space coordinate system of the three-dimensional MPI tomographic image is adjusted to be consistent with the imaging space coordinate system of the CT image; the resolution of the MPI three-dimensional tomographic image and the CT image is adjusted to be the same by linear interpolation.
[0095] The cropped surface fluorescence image and the three-dimensional MPI image mapped into the SIS are input into the trained fluorescence-magnetic particle image fusion model CAUnet (i.e., the trained multi-modal image fusion model, the same below) for image fusion to obtain a three-dimensional optical-magnetic fusion image, that is, the fluorescence-magnetic particle multi-modal image fusion method shown in the above operations S210 to S230.
[0096] The multi-modal image fusion model of the present invention will be described below through specific embodiments in conjunction with the drawings, so that those skilled in the art can clearly understand how the technical solution provided by the present invention realizes multi-modal image fusion.
[0097] Figure 3 It is a schematic structural diagram of the CAUnet model according to an embodiment of the present invention.
[0098] Taking the trained fluorescence-magnetic particle image fusion model CAUnet model as the above-mentioned trained multi-modal image fusion model, as Figure 3 shown, the above CAUnet includes an FMI image encoder, an MPI image encoder, a pre-fusion module, a first adaptive cross-attention mechanism fusion module, a second adaptive cross-attention mechanism fusion module, a third adaptive cross-attention mechanism fusion module, and an image fusion decoder.
[0099] The FMI image encoder is constructed by connecting a two-dimensional convolutional block and three three-dimensional convolutional blocks in sequence, which are respectively used as an FMI expansion module, a first FMI convolutional module, a second FMI convolutional module, and a third FMI convolutional module, and their outputs are an FMI expansion feature map, a first FMI feature map, a second FMI feature map, and a third FMI feature map.
[0100] Among them, the FMI expansion module is used to obtain a multi-channel near-infrared fluorescence expansion feature map; the FMI expansion module is constructed based on a 5×5 two-dimensional convolutional layer
[0101] According to an embodiment of the present invention, the initial fusion of the multi-channel near-infrared fluorescence expansion feature map and the registered three-dimensional magnetic particle tomographic image to obtain a three-dimensional near-infrared fluorescence expansion feature map includes: using the pre-fusion module of the trained multi-modal image fusion model to perform a channel-by-channel multiplication process on the maximum value of the registered three-dimensional magnetic particle tomographic image in each channel and the multi-channel near-infrared fluorescence expansion feature map to obtain a three-dimensional near-infrared fluorescence expansion feature map.
[0102] Figure 4 It is a schematic structural diagram of the pre-fusion module of the CAUnet model according to an embodiment of the present invention.
[0103] The multi-channel near-infrared fluorescence extended feature map and the registered three-dimensional magnetic particle tomography image are realized by the pre-fusion module of CAUnet. Among them, the pre-fusion module is used to fuse the FMI extended feature map and the MPI three-dimensional image input into the model. As Figure 4 shown, the input of the pre-fusion module is the FMI extended feature map (i.e., the multi-channel near-infrared fluorescence extended feature map) and the three-dimensional MPI image. This module is configured to take the maximum value of each slice (channel) of the input three-dimensional MPI image and multiply it with the FMI extended feature map channel by channel.
[0104] According to an embodiment of the present invention, the above-mentioned multi-round three-dimensional convolution processing of the three-dimensional near-infrared fluorescence extended feature map by the trained multi-modal image fusion model to obtain the multi-scale near-infrared fluorescence feature map includes: using the first near-infrared fluorescence convolution module of the trained multi-modal image fusion model to perform initial three-dimensional convolution processing on the three-dimensional near-infrared fluorescence extended feature map, and performing initial activation processing on the result of the initial three-dimensional convolution processing to obtain an initial activation result; using the first near-infrared fluorescence convolution module to perform secondary three-dimensional convolution processing on the initial activation result, and performing secondary activation processing on the result of the secondary three-dimensional convolution processing to obtain the first-scale near-infrared fluorescence feature map. Using the second near-infrared fluorescence convolution module of the trained multi-modal image fusion model to perform the same processing on the first-scale near-infrared fluorescence feature map as on the three-dimensional near-infrared fluorescence extended feature map to obtain the second-scale near-infrared fluorescence feature map; using the third near-infrared fluorescence convolution module of the trained multi-modal image fusion model to perform the same processing on the second-scale near-infrared fluorescence feature map as on the three-dimensional near-infrared fluorescence extended feature map to obtain the third-scale near-infrared fluorescence feature map.
[0105] Figure 5 It is a schematic structural diagram of the downsampling convolution module according to an embodiment of the present invention.
[0106] The above-mentioned acquisition of the multi-scale near-infrared fluorescence feature map is completed by the first FMI convolution module (i.e., the first near-infrared fluorescence convolution module, the same below), the second FMI convolution module (the second near-infrared fluorescence convolution module, the same below), and the third FMI convolution module (the third near-infrared fluorescence convolution module, the same below). Since the above-mentioned first FMI convolution module, second FMI convolution module, and third FMI convolution module are all constructed based on a 5×5×5 three-dimensional convolution layer with a stride of 1, a LeakyReLU activation layer, a 3×2×2 three-dimensional convolution layer with a channel stride of 1 and a spatial stride of 2 (stride of (1, 2, 2)), and a LeakyReLU activation layer connected in sequence, asFigure 5 As shown. Therefore, obtaining the multi-scale near-infrared fluorescence feature map requires two three-dimensional convolutions and two activation processes.
[0107] The MPI image encoder is constructed based on three sequentially connected three-dimensional convolution blocks, namely the first MPI convolution module, the second MPI convolution module, and the third MPI convolution module. Their outputs are the first MPI feature map (i.e., the first-scale magnetic particle feature map, the same below), the second MPI feature map (the second-scale magnetic particle feature map, the same below), and the third MPI feature map (the third-scale magnetic particle feature map, the same below);
[0108] The above-mentioned first MPI convolution module, second MPI convolution module, and third MPI convolution module are all constructed based on a sequentially connected 5×5×5 three-dimensional convolution layer with a stride of 1, a LeakyReLU activation layer, a 3×2×2 three-dimensional convolution layer with a channel stride of 1 and a spatial stride of 2 (stride of (1, 2, 2)), and a LeakyReLU activation layer. The structure is as Figure 5 shown. Therefore, the process of obtaining the first, second, and third-scale magnetic particle feature maps is similar to the process of obtaining the first, second, and third-scale near-infrared fluorescence feature maps.
[0109] According to an embodiment of the present invention, the above-mentioned multi-round feature map fusion of the multi-scale near-infrared fluorescence feature map and the multi-scale magnetic particle feature map at the same scale based on the adaptive cross-attention mechanism includes: using the first cross-attention mechanism module of the trained multi-modal image fusion model to fuse the first-scale near-infrared fluorescence feature map and the first-scale magnetic particle feature map in the multi-scale magnetic particle feature map to obtain a first fusion feature map; using the second cross-attention mechanism module of the trained multi-modal image fusion model to fuse the second-scale near-infrared fluorescence feature map and the second-scale magnetic particle feature map in the multi-scale magnetic particle feature map to obtain a second fusion feature map; using the third cross-attention mechanism module of the trained multi-modal image fusion model to fuse the third-scale near-infrared fluorescence feature map and the third-scale magnetic particle feature map in the multi-scale magnetic particle feature map to obtain a third fusion feature map; using the trained multi-modal image fusion model to first connect the third fusion feature map with the third-scale magnetic particle feature map, and perform first upsampling on the fused feature map after the first connection to obtain a first upsampled fused feature map; using the trained multi-modal image fusion model to secondarily connect the first upsampled fused feature map with the second fusion feature map, and perform second upsampling on the fused feature map after the secondary connection to obtain a second upsampled fused feature map; using the trained multi-modal image fusion model to thirdly connect the second upsampled fused image with the first fusion feature map, and perform third upsampling on the fused feature map after the third connection to obtain the multi-round feature map fusion result.
[0110] The above-mentioned obtaining of the multi-round feature map fusion result utilizes the first adaptive cross-attention mechanism fusion module, the second adaptive cross-attention mechanism fusion module, and the third adaptive cross-attention mechanism fusion module of the CAUnet model, that is, three times of multi-modal feature map fusion are performed, and those skilled in the art can set the number of times of feature map fusion according to actual needs.
[0111] The above-mentioned adaptive cross-attention mechanism fusion module is used to fuse the outputs of the FMI convolution module and the MPI convolution module. For the Nth adaptive cross-attention mechanism module, the inputs are the Nth FMI feature map and the Nth MPI feature map. The adaptive cross-attention mechanism fusion module includes the first adaptive cross-attention mechanism module, the second adaptive cross-attention mechanism module, and the adaptive weighting module; after the inputs are parallelly calculated by the first and second cross-attention mechanism modules to obtain the first attention feature and the second attention feature, the adaptive cross-attention feature of the adaptive cross-attention mechanism fusion module is calculated through the adaptive weighting module, denoted as the Nth adaptive cross-attention feature.
[0112] According to an embodiment of the present invention, the first cross-attention mechanism module of the above-mentioned trained multi-modal image fusion model fuses the first-scale magnetic particle feature map in the first-scale near-infrared fluorescence feature map and the multi-scale magnetic particle feature maps, and obtaining the first fusion feature map includes: respectively rearranging the data of the first-scale near-infrared fluorescence feature map and the first-scale magnetic particle feature map to obtain the rearranged first-scale near-infrared fluorescence feature map and the rearranged first-scale magnetic particle feature map; respectively performing linear transformation on the rearranged first-scale near-infrared fluorescence feature map and the rearranged first-scale magnetic particle feature map to obtain a first near-infrared fluorescence multi-value matrix and a first magnetic particle multi-value matrix; respectively performing multi-head attention mechanism division on the rearranged first-scale near-infrared fluorescence feature map and the rearranged first-scale magnetic particle feature map; in the case where the multi-head attention mechanism division is completed, performing multi-head adaptive cross-attention operation on the first near-infrared fluorescence multi-value matrix and the first magnetic particle multi-value matrix to obtain a first near-infrared fluorescence multi-head attention result and a first magnetic particle multi-head attention result; respectively rearranging the data of the first near-infrared fluorescence multi-head attention result and the first magnetic particle multi-head attention result to obtain the rearranged first near-infrared fluorescence multi-head attention result and the rearranged first magnetic particle multi-head attention result; connecting the rearranged first near-infrared fluorescence multi-head attention result with the rearranged first-scale near-infrared fluorescence feature map to obtain a first near-infrared fluorescence cross-attention feature; connecting the rearranged first magnetic particle multi-head attention result and the rearranged first-scale magnetic particle feature map to obtain a first magnetic particle cross-attention feature, and performing weighted operation on the first near-infrared fluorescence cross-attention feature and the first magnetic particle cross-attention feature to obtain the first fusion feature map.
[0113] According to an embodiment of the present invention, in the case where the multi-head attention mechanism division is completed, performing multi-head adaptive cross-attention operation on the first near-infrared fluorescence multi-value matrix and the first magnetic particle multi-value matrix to obtain a first near-infrared fluorescence multi-head attention result and a first magnetic particle multi-head attention result includes: transposing the key-value matrix in the first magnetic particle multi-value matrix and multiplying it with the query matrix in the first near-infrared fluorescence multi-value matrix, and performing normalization activation processing on the result of the matrix multiplication to obtain a first near-infrared fluorescence attention score; multiplying the first near-infrared fluorescence attention score with the value matrix in the first magnetic particle multi-value matrix, and performing layer normalization on the result of the multiplication to obtain a first near-infrared fluorescence multi-head attention result; transposing the key-value matrix in the first near-infrared fluorescence multi-value matrix and multiplying it with the query matrix in the first magnetic particle multi-value matrix, and performing normalization activation processing on the result of the matrix multiplication to obtain a first magnetic particle attention score; multiplying the first magnetic particle attention score with the value matrix in the first near-infrared fluorescence multi-value matrix, and performing layer normalization on the result of the multiplication to obtain a first magnetic particle multi-head attention result.
[0114] The above embodiments illustrate the process of obtaining the first fused feature map, and the process of obtaining other fused feature maps is similar to that of the first fused feature map.
[0115] The following further elaborates on the process of obtaining the above-mentioned fused feature maps through specific embodiments.
[0116] Figure 6 It is a data processing flowchart of the adaptive cross-attention mechanism fusion module according to an embodiment of the present invention.
[0117] Taking the first cross-attention mechanism fusion module as an example, as Figure 5 shown, it includes three linear transformation layers W q , W k , W v , an attention mechanism layer, a multi-layer perceptron layer, and a residual connection; after the input FMI features are rearranged, they are mapped to Q (Query) through the linear transformation W q , and the MPI features are rearranged and then mapped to K (key) and V (Value) through the linear transformations W k , W v ; and Q, K, and V are divided into multiple heads and input into the attention mechanism layer for the calculation of the multi-head attention mechanism; the attention mechanism layer multiplies the transpose of the matrices Q and K and normalizes them using the Softmax function to obtain attention scores, multiplies the attention scores by the matrix V, and performs layer normalization; the multi-layer perceptron layer consists of two linearly connected layers and a LeakyReLU activation function connected in sequence; the output of the attention mechanism layer is rearranged, and a residual connection is made with the rearranged FMI features, and then passed through the multi-layer perceptron layer and the residual connection, and the data is rearranged to restore the shape of the input data to obtain the first cross-attention feature.
[0118] The second cross-attention mechanism module has the same structure as the first cross-attention module, but its input order is opposite to that of the first cross-attention mechanism module, and its output is the second cross-attention feature.
[0119] The adaptive weighting module is used to calculate the weights for adding the first cross-attention mechanism and the second cross-attention mechanism, and is composed of an adaptive average pooling layer, a fully connected layer, and a ReLU activation function layer connected in sequence; the input of the adaptive weighting module is a three-dimensional magnetic particle image, the first cross-attention feature, and the second cross-attention feature; the three-dimensional magnetic particle image undergoes adaptive weighted pooling in the tomographic (channel) dimension, and after passing through the fully connected layer and the ReLU activation function layer, the weight of the first cross-attention feature is output , and the weight of the second cross-attention mechanism feature is set to , the output of the adaptive weighting module is obtained by performing weighted summation on them.
[0120] Figure 7 It is a schematic structural diagram of the upsampling module according to an embodiment of the present invention.
[0121] The image fusion decoder module is constructed based on three sequentially connected three-dimensional upsampling convolutional modules and a filtering convolutional module, which are respectively used as the first upsampling convolutional module, the second upsampling convolutional module, the third upsampling convolutional module, and the filtering convolutional module; wherein, the structure of the upsampling convolutional module is as Figure 7 shown.
[0122] The first upsampling convolutional module, the second upsampling convolutional module, and the third upsampling convolutional module are all based on a 3×2×2 three-dimensional transposed convolutional layer with a channel stride of 1 and a spatial stride of 2 (stride of (1, 2, 2)), a LeakyReLU activation layer, a 5×5×5 three-dimensional convolutional layer with a stride of 1, and a LeakyReLU activation layer, as Figure 7 shown.
[0123] The input of the third upsampling module is the concatenation of the third MPI feature map and the third adaptive cross-attention feature in the channel dimension; the input of the second upsampling module is the concatenation of the output of the third upsampling module and the second adaptive cross-attention feature in the channel dimension; the input of the first upsampling module is the concatenation of the output of the second upsampling module and the first adaptive cross-attention feature in the channel dimension.
[0124] The filtering convolutional module is based on a 5×5×5 three-dimensional convolutional layer and a ReLU activation layer; the input of the filtering convolutional module is the output of the first upsampling module, and the output after passing through the three-dimensional convolutional layer and the ReLU activation layer is used as the output of the entire CAUnet.
[0125] Figure 8 It is a flowchart of the training method of the multi-modal image fusion model according to an embodiment of the present invention.
[0126] As Figure 8 shown, the above-mentioned training method of the multi-modal image fusion model is applied to the above-mentioned fluorescence-magnetic particle image fusion method, including operations S810 to S850.
[0127] In operation S810, based on the photon propagation model, through linear interpolation operation and Gaussian noise random addition operation, a simulated two-dimensional near-infrared fluorescence image set of the three-dimensional tumor phantom set in the standardized space is obtained.
[0128] In operation S820, the point spread function of the three-dimensional magnetic particle tomography image is subjected to three-dimensional convolution processing with the image set of the three-dimensional tumor phantom to obtain a simulated three-dimensional magnetic particle tomography image set.
[0129] In operation S830, a multi-modal image fusion model is used to perform multi-modal fusion between the simulated two-dimensional near-infrared fluorescence image set and the simulated three-dimensional magnetic particle tomography image set to obtain a predicted fusion map.
[0130] In operation S840, a training loss function is used to calculate the loss value between the predicted fusion map and the ground truth label corresponding to the predicted fusion map, and the loss value is used to update the parameters of the multi-modal image fusion model.
[0131] In operation S850, the multi-modal fusion operation between images, the loss value calculation operation, and the parameter update operation are iteratively performed until all the simulated images in the simulated two-dimensional near-infrared fluorescence image set and the simulated three-dimensional magnetic particle tomography image set are processed, and a trained multi-modal image fusion model is obtained.
[0132] The multi-modal image fusion model obtained by training through the above operations S810 to S850 has good generalization and can be applied to, but is not limited to, the multi-modal image fusion of tumor pathological regions; at the same time, the multi-modal image fusion model obtained by training through the above operations has good fusion accuracy and high image resolution, and can be used to assist relevant personnel in accurately judging the state of the pathological region.
[0133] According to an embodiment of the present invention, the above-mentioned trained multi-modal image fusion model includes a two-dimensional near-infrared fluorescence image encoder, a three-dimensional magnetic particle tomography image encoder, a pre-fusion module, a first adaptive cross-attention mechanism fusion module, a second adaptive cross-attention mechanism fusion module, a third adaptive cross-attention mechanism fusion module, and an image fusion decoder; wherein, the two-dimensional near-infrared fluorescence image encoder includes a two-dimensional convolutional near-infrared fluorescence expansion module, a first near-infrared fluorescence three-dimensional convolutional module, a second near-infrared fluorescence three-dimensional convolutional module, and a third near-infrared fluorescence three-dimensional convolutional module; wherein, the three-dimensional magnetic particle tomography image encoder includes a first magnetic particle three-dimensional convolutional module, a second magnetic particle three-dimensional convolutional module, and a third magnetic particle three-dimensional convolutional module; wherein, the first adaptive cross-attention mechanism fusion module, the second adaptive cross-attention mechanism fusion module, and the third adaptive cross-attention mechanism fusion module all include an attention mechanism layer, a multi-layer perceptron layer, a residual connection layer, and a plurality of linear transformation layers; wherein, the image fusion decoder includes a filtering convolutional module and a plurality of upsampling three-dimensional convolutional modules.
[0134] According to an embodiment of the present invention, the above-mentioned simulation two-dimensional near-infrared fluorescence image set of the three-dimensional tumor phantom set in the standardized space obtained through the linear interpolation operation and the Gaussian noise random addition operation based on the photon propagation model includes: setting a three-dimensional tumor phantom in the target tissue mapped into the standardized space, and converting the pixel values of the three-dimensional tumor phantom into the concentration of the dual-mode probe; discretizing the concentration distribution of the dual-mode probe into a concentration vector, and randomly adding Gaussian noise to the concentration vector to obtain a concentration vector with noise; performing forward calculation on the target tissue using the photon propagation model to obtain the system matrix of fluorescence imaging, and performing an operation on the system matrix of fluorescence imaging and the vector with noise to obtain a fluorescence intensity distribution vector; performing a linear interpolation operation on the fluorescence intensity distribution vector, and performing a Gaussian noise random addition operation on the fluorescence intensity distribution vector after the linear interpolation operation to obtain a simulation two-dimensional near-infrared fluorescence image set.
[0135] For the acquisition processes of the simulation two-dimensional near-infrared fluorescence image set and the simulation three-dimensional magnetic particle tomography image set, the present invention will be described in detail through the following specific embodiments.
[0136] Figure 9 It is a schematic diagram of the simulation multimodal image of the three-dimensional tumor phantom according to an embodiment of the present invention.
[0137] Set a three-dimensional tumor phantom in the tissue mapped into the SIS, convert its pixel values into the concentration of the dual-mode probe, and discretize the in-vivo probe concentration distribution into a vector and add Gaussian noise.
[0138] In this specific embodiment, the set three-dimensional tumor phantoms include a single tumor phantom and a double tumor phantom; the single tumor phantom includes a spherical phantom, a cluster phantom, and an MNIST phantom; the diameter of the single tumor phantom is randomly set to 0.9 - 4.5 mm; the cluster phantom is formed by superimposing 3 - 5 spherical phantoms with overlapping positions; the MNIST phantom is obtained by randomly magnifying the MNIST dataset image by 1 - 3 times and mapping the coordinates to the SIS after expanding to a random number between 6 - 20 in the channel dimension; the double tumor phantom is obtained by superimposing two single tumor phantoms.
[0139] In this specific embodiment, converting the pixel values of the tumor phantom into the concentration of the dual-mode probe is performed by the method shown in formula (1):
[0140] (1),
[0141] where, is the concentration signal of the dual-mode probe at pixel point r, is the pixel gray value at pixel point r, and is the set maximum concentration of the dual-mode probe and the maximum pixel gray value, where is preferably set to 5×10 7 mmol / L, is preferably set to 255, and the tumor phantom image is as Figure 9 shown.
[0142] The system matrix is obtained by performing a forward model calculation on the tissue according to the photon propagation model ; the fluorescence intensity distribution vector of the tissue surface nodes is calculated using the linear equation ; linear interpolation is performed on to obtain a simulated surface fluorescence image, and Gaussian noise is added. The simulated fluorescence image is as Figure 9 shown.
[0143] In this embodiment, the method for obtaining the system matrix by performing a forward model calculation on the tissue according to the photon propagation model is as follows: Assume that the excitation light source is located on the upper surface of the SIS; the propagation process of fluorescence photons in the imaging object tissue is described by the diffusion approximation equation described by formula (2) and the refractive index deviation between the object surface and air is described by the Robin boundary condition shown in formula (3):
[0144] (2),
[0145] (3),
[0146] where x and m represent the excitation light and the emission light, respectively. represents the photon flux density at position . represents the diffusion coefficient, where , is the absorption coefficient, , is the scattering coefficient, is the anisotropy coefficient; is the intensity of the excitation light, is the position of the excitation light. represents the fluorescence yield at position . represents the boundary of the biological tissue, is the normal vector of the biological tissue surface, is the optical refractive index deviation between the imaging object boundary and air. The system matrix is obtained by solving equations (2) and (3).
[0147] The point spread function of the three-dimensional MPI is calculated and three-dimensionally convolved with the tumor phantom image to obtain a simulated three-dimensional MPI image.
[0148] In this embodiment, the method for calculating the point spread function of 3D MPI is as follows: Set the diameter of the magnetic particles (preferably 20 nm), temperature (preferably 20 °C), saturation magnetization intensity (preferably ), maximum magnetic particle concentration (preferably 5×107 mmol / L), select the field magnetic field gradient (preferably randomly from 2.6 T / m×1.3 T / m×1.3 T / m to 4 T / m×2 T / m×2 T / m, where the field gradient magnitudes in the y and z directions are the same, and the field gradient magnitude in the x direction is half of that in the x direction), set the coil sensitivity (preferably 1.0), imaging field of view size (preferably set to 20 mm×20 mm×15 mm), and calculate the point spread function of 3D MPI according to formula (4), as shown in formula (4):
[0149] (4),
[0150] Wherein, , , , are the selected field gradients in the x, y, and z directions, , where is is the Boltzmann constant, is the vacuum permeability, is the temperature, is the particle magnetic moment, is the Lagrangian function, is the derivative of the Lagrangian function; The simulated 3D MPI image is as Figure 9 shown.
[0151] Using the simulated surface FMI image and the simulated 3D MPI image as training samples, and using the set tumor phantom as the ground truth label.
[0152] During the model training process, calculate the loss value between the output result of the multimodal image fusion model and the corresponding ground truth label, and update the model parameters of the CAUnet model.
[0153] In the embodiment of the present invention, preferably, the loss function is composed of the Dice coefficient loss function and the mean square error (MSE) loss function weighted, where the Dice coefficient loss weight is 0.9 and the MSE loss function weight is 0.1. Calculate the loss value of the model, perform backpropagation, adjust the parameters, and update the parameters of the CAUnet model. The optimization algorithm used is Adaptive Moment Estimation (Adam).
[0154] In this embodiment, the model is trained iteratively until the set number of pre-training times is reached or the set accuracy is achieved, at which point the training ends and the trained CAUnet model is obtained.
[0155] Figure 10 It is a schematic diagram of a three-dimensional optical-magnetic fusion image output by the CAUnet model according to an embodiment of the present invention.
[0156] In this embodiment, the collected two-dimensional near-infrared fluorescence images, three-dimensional magnetic particle images, and CT image data are preprocessed, and the preprocessed data is input into the trained CAUnet model for multi-modal image fusion processing to obtain a Figure 10 three-dimensional optical-magnetic fusion image as shown.
[0157] Figure 11 It schematically shows a block diagram of an electronic device suitable for implementing the fluorescence-magnetic particle image fusion method and the training method of the multi-modal image fusion model according to an embodiment of the present invention.
[0158] As Figure 11 shown, the electronic device 1100 according to an embodiment of the present invention includes a processor 1101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1102 or a program loaded from a storage section 1108 into a random access memory (RAM) 1103. The processor 1101 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 1101 can also include on-board memory for caching purposes. The processor 1101 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0159] In the RAM 1103, various programs and data required for the operation of the electronic device 1100 are stored. The processor 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. The processor 1101 executes various operations of the method flow according to an embodiment of the present invention by executing the program in the ROM 1102 and / or the RAM 1103. It should be noted that the program can also be stored in one or more memories other than the ROM 1102 and the RAM 1103. The processor 1101 can also execute various operations of the method flow according to an embodiment of the present invention by executing the program stored in the one or more memories.
[0160] According to an embodiment of the present invention, the electronic device 1100 may further include an input / output (I / O) interface 1105, and the input / output (I / O) interface 1105 is also connected to the bus 1104. The electronic device 1100 may further include one or more of the following components connected to the input / output (I / O) interface 1105: an input part 1106 including a keyboard, a mouse, etc.; an output part 1107 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage part 1108 including a hard disk, etc.; and a communication part 1109 including a network interface card such as a LAN card, a modem, etc. The communication part 1109 performs communication processing via a network such as the Internet. The drive 1110 is also connected to the input / output (I / O) interface 1105 as needed. A removable medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1110 as needed so that a computer program read from it can be installed into the storage part 1108 as needed.
[0161] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the methods according to the embodiments of the present invention are implemented.
[0162] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, device, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM 1102 and / or the RAM 1103 described above and / or one or more memories other than the ROM 1102 and the RAM 1103.
[0163] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0164] Those skilled in the art will appreciate that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0165] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A fluorescence-magnetic particle image fusion method, characterized in that The method includes: Performing two-dimensional convolution processing on the registered two-dimensional near-infrared fluorescence image by using the trained multi-modal image fusion model to obtain a multi-channel near-infrared fluorescence extended feature map, and initially fusing the multi-channel near-infrared fluorescence extended feature map with the registered three-dimensional magnetic particle tomography image to obtain a three-dimensional near-infrared fluorescence extended feature map; Performing initial three-dimensional convolution processing on the three-dimensional near-infrared fluorescence extended feature map by using the first near-infrared fluorescence convolution module of the trained multi-modal image fusion model, and performing initial activation processing on the result of the initial three-dimensional convolution processing to obtain an initial activation result; Performing secondary three-dimensional convolution processing on the initial activation result by using the first near-infrared fluorescence convolution module, and performing secondary activation processing on the result of the secondary three-dimensional convolution processing to obtain a first-scale near-infrared fluorescence feature map; Performing the same processing on the first-scale near-infrared fluorescence feature map as that on the three-dimensional near-infrared fluorescence extended feature map by using the second near-infrared fluorescence convolution module of the trained multi-modal image fusion model to obtain a second-scale near-infrared fluorescence feature map; Performing the same processing on the second-scale near-infrared fluorescence feature map as that on the three-dimensional near-infrared fluorescence extended feature map by using the third near-infrared fluorescence convolution module of the trained multi-modal image fusion model to obtain a third-scale near-infrared fluorescence feature map, and performing multi-round three-dimensional convolution processing on the registered three-dimensional magnetic particle tomography image to obtain a multi-scale magnetic particle feature map; Fusing the first-scale near-infrared fluorescence feature map and the first-scale magnetic particle feature map in the multi-scale magnetic particle feature map by using the first cross-attention mechanism module of the trained multi-modal image fusion model to obtain a first fusion feature map; Fusing the second-scale near-infrared fluorescence feature map and the second-scale magnetic particle feature map in the multi-scale magnetic particle feature map by using the second cross-attention mechanism module of the trained multi-modal image fusion model to obtain a second fusion feature map; Fusing the third-scale near-infrared fluorescence feature map and the third-scale magnetic particle feature map in the multi-scale magnetic particle feature map by using the third cross-attention mechanism module of the trained multi-modal image fusion model to obtain a third fusion feature map; First connecting the third fusion feature map and the third-scale magnetic particle feature map by using the trained multi-modal image fusion model, and performing first upsampling on the fused feature map after the first connection to obtain a first upsampled fused feature map; Second connecting the first upsampled fused feature map and the second fusion feature map by using the trained multi-modal image fusion model, and performing second upsampling on the fused feature map after the second connection to obtain a second upsampled fused feature map; Using the trained multi-modal image fusion model, the second upsampled fusion image is concatenated with the first fusion feature map three times, and the fused feature map after the three concatenations is upsampled for the third time to obtain the multi-round feature map fusion result, and the multi-round feature map fusion result is subjected to filtering convolution processing to obtain the three-dimensional fluorescence-magnetic particle fusion image.
2. The method according to claim 1, characterized in that, The initial fusion of the multi-channel near-infrared fluorescence extended feature map and the registered three-dimensional magnetic particle tomography image to obtain the three-dimensional near-infrared fluorescence extended feature map includes: Using the pre-fusion module of the trained multi-modal image fusion model, the maximum value of the registered three-dimensional magnetic particle tomography image on each channel is multiplied by the multi-channel near-infrared fluorescence extended feature map channel by channel to obtain the three-dimensional near-infrared fluorescence extended feature map.
3. The method according to claim 1, wherein Using the first cross-attention mechanism module of the trained multi-modal image fusion model to fuse the first-scale near-infrared fluorescence feature map and the first-scale magnetic particle feature map in the multi-scale magnetic particle feature map to obtain the first fusion feature map, including: Respectively rearrange the data of the first-scale near-infrared fluorescence feature map and the first-scale magnetic particle feature map to obtain the rearranged first-scale near-infrared fluorescence feature map and the rearranged first-scale magnetic particle feature map; Respectively perform linear transformation on the rearranged first-scale near-infrared fluorescence feature map and the rearranged first-scale magnetic particle feature map to obtain the first near-infrared fluorescence multi-value matrix and the first magnetic particle multi-value matrix; Respectively perform multi-head attention mechanism partitioning on the rearranged first-scale near-infrared fluorescence feature map and the rearranged first-scale magnetic particle feature map; After the multi-head attention mechanism partitioning is completed, perform multi-head adaptive cross-attention operation on the first near-infrared fluorescence multi-value matrix and the first magnetic particle multi-value matrix to obtain the first near-infrared fluorescence multi-head attention result and the first magnetic particle multi-head attention result; Respectively rearrange the data of the first near-infrared fluorescence multi-head attention result and the first magnetic particle multi-head attention result to obtain the rearranged first near-infrared fluorescence multi-head attention result and the rearranged first magnetic particle multi-head attention result; Connect the rearranged first near-infrared fluorescence multi-head attention result with the rearranged first-scale near-infrared fluorescence feature map to obtain the first near-infrared fluorescence cross-attention feature; Connect the rearranged first magnetic particle multi-head attention result and the rearranged first-scale magnetic particle feature map to obtain the first magnetic particle cross-attention feature, and perform weighted operation on the first near-infrared fluorescence cross-attention feature and the first magnetic particle cross-attention feature to obtain the first fusion feature map.
4. The method according to claim 3, wherein After the multi-head attention mechanism partitioning is completed, perform multi-head adaptive cross-attention operation on the first near-infrared fluorescence multi-value matrix and the first magnetic particle multi-value matrix to obtain the first near-infrared fluorescence multi-head attention result and the first magnetic particle multi-head attention result, including: Transpose the key-value matrix in the first magnetic particle multi-value matrix and multiply it with the query matrix in the first near-infrared fluorescence multi-value matrix, and perform normalization activation processing on the result of the matrix multiplication to obtain the first near-infrared fluorescence attention score; Multiply the first near-infrared fluorescence attention score with the value matrix in the first magnetic particle multi-value matrix, and perform layer normalization on the result of the multiplication to obtain the first near-infrared fluorescence multi-head attention result; Transpose the key-value matrix in the first near-infrared fluorescence multi-value matrix and multiply it with the query matrix in the first magnetic particle multi-value matrix, and perform normalization activation processing on the result of the matrix multiplication to obtain the first magnetic particle attention score; Multiply the first magnetic particle attention score with the value matrix in the first near-infrared fluorescence multi-value matrix, and perform layer normalization on the result of the multiplication to obtain the first magnetic particle multi-head attention result.
5. A training method for a multi-modal image fusion model, applied to the method according to any one of claims 1 to 4, characterized in that, The method includes: Based on the photon propagation model, through linear interpolation operation and Gaussian noise random addition operation, obtain a simulated two-dimensional near-infrared fluorescence image set of a three-dimensional tumor phantom set in a standardized space; Perform three-dimensional convolution processing on the point spread function of the three-dimensional magnetic particle tomography image and the image set of the three-dimensional tumor phantom to obtain a simulated three-dimensional magnetic particle tomography image set; Use the multi-modal image fusion model to perform multi-modal fusion between the simulated two-dimensional near-infrared fluorescence image set and the simulated three-dimensional magnetic particle tomography image set to obtain a predicted fusion map; Use the training loss function to calculate the loss value between the predicted fusion map and the ground truth label corresponding to the predicted fusion map, and use the loss value to update the parameters of the multi-modal image fusion model; Iteratively perform multi-modal fusion operation between images, loss value calculation operation, and parameter update operation until all simulated images in the simulated two-dimensional near-infrared fluorescence image set and the simulated three-dimensional magnetic particle tomography image set are processed to obtain the trained multi-modal image fusion model.
6. The method according to claim 5, characterized in that, Based on the photon propagation model, through linear interpolation operation and Gaussian noise random addition operation, obtaining a simulated two-dimensional near-infrared fluorescence image set of a three-dimensional tumor phantom set in a standardized space includes: Set the three-dimensional tumor phantom in the target tissue mapped to the standardized space, and convert the pixel value of the three-dimensional tumor phantom into the concentration of the dual-mode probe; Discretize the concentration distribution of the dual-mode probe into a concentration vector, and randomly add Gaussian noise to the concentration vector to obtain a concentration vector with noise; Use the photon propagation model to perform forward calculation on the target tissue to obtain the system matrix of fluorescence imaging, and perform an operation on the system matrix of fluorescence imaging and the vector with noise to obtain a fluorescence intensity distribution vector; Perform a linear interpolation operation on the fluorescence intensity distribution vector, and perform a Gaussian noise random addition operation on the fluorescence intensity distribution vector after the linear interpolation operation to obtain the simulated two-dimensional near-infrared fluorescence image set.
7. The method according to claim 5, wherein The trained multi-modal image fusion model includes a two-dimensional near-infrared fluorescence image encoder, a three-dimensional magnetic particle tomography image encoder, a pre-fusion module, a first adaptive cross-attention mechanism fusion module, a second adaptive cross-attention mechanism fusion module, a third adaptive cross-attention mechanism fusion module, and an image fusion decoder; Among them, the two-dimensional near-infrared fluorescence image encoder includes a two-dimensional convolutional near-infrared fluorescence expansion module, a first near-infrared fluorescence three-dimensional convolutional module, a second near-infrared fluorescence three-dimensional convolutional module, and a third near-infrared fluorescence three-dimensional convolutional module; Among them, the three-dimensional magnetic particle tomography image encoder includes a first magnetic particle three-dimensional convolutional module, a second magnetic particle three-dimensional convolutional module, and a third magnetic particle three-dimensional convolutional module; Among them, the first adaptive cross-attention mechanism fusion module, the second adaptive cross-attention mechanism fusion module, and the third adaptive cross-attention mechanism fusion module all include an attention mechanism layer, a multi-layer perceptron layer, a residual connection layer, and multiple linear transformation layers; Among them, the image fusion decoder includes a filtering convolutional module and multiple upsampling three-dimensional convolutional modules.
Citation Information
Patent Citations
Multi-mode imaging system for small animal magnetic particle imaging and fluorescence molecular tomography
CN115844365A
Magnetic particle imaging (MPI) and fluorescence molecular tomography (FMT)-fused multimodal imaging system for small animal
US11940508B1