Fluorescence-magnetic particle image fusion method and multi-modal image fusion model training method

Through multimodal image fusion technology, combined with fluorescence and magnetic particle imaging, the problem of insufficient imaging depth and resolution in single-modal imaging is solved, and high-precision and high-quality medical image fusion is achieved.

CN119991468AActive Publication Date: 2025-05-13INST OF AUTOMATION CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510135612.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-13
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

Fluorescent molecular imaging and magnetic particle imaging have problems such as insufficient imaging depth and low spatial resolution in single-modal imaging, which hinders its widespread use in clinical applications.

Method used

A fluorescence-magnetic particle image fusion method is adopted to carry out multiple rounds of convolution processing and feature map fusion of two-dimensional near-infrared fluorescence images and three-dimensional magnetic particle tomography images through the trained multi-modal image fusion model. Combined with the adaptive cross attention mechanism, high-quality three-dimensional fluorescence-magnetic particle fusion images are generated.

Benefits of technology

It significantly improves the fusion accuracy and quality of multimodal medical images, makes up for the insufficient imaging depth and resolution of single-modal imaging technology, and achieves more accurate and comprehensive imaging of target areas such as tumors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991468A_ABST
    Figure CN119991468A_ABST
Patent Text Reader

Abstract

The invention provides a fluorescence-magnetic particle image fusion method which can be applied to the technical field of medical image processing. The method comprises the following steps: performing two-dimensional convolution processing on a registered two-dimensional near-infrared fluorescence image to obtain a multi-channel near-infrared fluorescence extended feature map, and performing initial fusion on the multi-channel near-infrared fluorescence extended feature map and a registered three-dimensional magnetic particle cross-sectional image to obtain a three-dimensional near-infrared fluorescence extended feature map; performing multi-round three-dimensional convolution processing on the three-dimensional near-infrared fluorescence extended feature map to obtain a multi-scale near-infrared fluorescence feature map, and performing multi-round three-dimensional convolution processing on the registered three-dimensional magnetic particle cross-sectional image to obtain a multi-scale magnetic particle feature map; and based on an adaptive cross attention mechanism, performing multi-round feature map fusion on the multi-scale near-infrared fluorescence feature map and the multi-scale magnetic particle feature map on the same scale, and performing filtering convolution processing on a multi-round feature map fusion result to obtain a three-dimensional fluorescence-magnetic particle fusion image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, specifically to the technical field of multimodal medical image fusion, and more specifically to a fluorescence-magnetic particle image fusion method and a training method for a multimodal image fusion model. Background Art

[0002] Fluorescence Molecular Imaging (FMI) is an emerging molecular imaging technology with the advantages of high sensitivity and high resolution, and can detect tumors non-invasively. Fluorescence Molecular Tomography (FMT) makes up for the deficiency of FMI that it can only perform two-dimensional imaging and cannot provide three-dimensional spatial information of tumors. It establishes a propagation model of photons in the body and reversely solves the surface fluorescence obtained by FMI to obtain the three-dimensional spatial distribution of fluorescent molecular probes in the body, thereby reconstructing the spatial information of the tumor. FMT can reconstruct shallow light sources in biological tissues with high sensitivity and high resolution. However, due to the strong absorption and scattering of photons in tissues, its imaging depth is limited, and it can only image superficial tumors. And because the data that can actually be collected is limited to the surface of the organism, the reconstruction problem has a strong ill-posedness, and it needs to set a reasonable regularization prior based on experience to solve it. The above problems have hindered the clinical application of FMT.

[0003] Magnetic Particle Imaging (MPI) is an emerging molecular imaging technology used to visualize the spatial distribution of superparamagnetic iron oxide nanoparticles in vivo. It has the advantages of no imaging depth limitation, linear quantification, high sensitivity, no background signal interference, and no ionizing radiation hazards, and has broad prospects for biomedical applications. However, MPI has low spatial resolution and resolution anisotropy, which seriously affects imaging accuracy and quality, and hinders the clinical application of MPI.

[0004] In view of the technical problems existing in single-modality imaging of FMI, FMT and MPI, it is necessary to provide a multi-modality imaging technical solution to solve the technical problems existing in single-modality imaging technical solutions such as insufficient resolution or insufficient imaging depth. Summary of the invention

[0005] In view of the above problems, the present invention provides a fluorescence-magnetic particle image fusion method and a training method for a multimodal image fusion model for improving the accuracy and quality of multimodal medical image fusion.

[0006] According to a first aspect of the present invention, there is provided a fluorescence-magnetic particle image fusion method, comprising:

[0007] The trained multimodal image fusion model is used to perform two-dimensional convolution processing on the registered two-dimensional near-infrared fluorescence image to obtain a multi-channel near-infrared fluorescence extended feature map, and the multi-channel near-infrared fluorescence extended feature map is initially fused with the registered three-dimensional magnetic particle tomography image to obtain a three-dimensional near-infrared fluorescence extended feature map;

[0008] The trained multimodal image fusion model is used to perform multiple rounds of three-dimensional convolution processing on the three-dimensional near-infrared fluorescence extension feature map to obtain a multi-scale near-infrared fluorescence feature map, and multiple rounds of three-dimensional convolution processing is performed on the registered three-dimensional magnetic particle tomography image to obtain a multi-scale magnetic particle feature map;

[0009] Based on the adaptive cross-attention mechanism, the trained multimodal image fusion model is used to perform multi-round feature map fusion on the same scale of multi-scale near-infrared fluorescence feature maps and multi-scale magnetic particle feature maps. The fusion results of the multi-round feature maps are then subjected to filtered convolution processing to obtain a three-dimensional fluorescence-magnetic particle fusion image.

[0010] According to an embodiment of the present invention, the above-mentioned initial fusion of the multi-channel near-infrared fluorescence extended feature map with the registered three-dimensional magnetic particle tomography image to obtain the three-dimensional near-infrared fluorescence extended feature map includes:

[0011] The pre-fusion module of the trained multimodal image fusion model is used to multiply the maximum value of the registered three-dimensional magnetic particle tomography image in each channel with the multi-channel near-infrared fluorescence extended feature map channel by channel to obtain the three-dimensional near-infrared fluorescence extended feature map.

[0012] According to an embodiment of the present invention, the above-mentioned multi-scale near-infrared fluorescence feature map obtained by performing multiple rounds of three-dimensional convolution processing on the three-dimensional near-infrared fluorescence extended feature map using the trained multi-modal image fusion model includes:

[0013] Using the first near-infrared fluorescence convolution module of the trained multimodal image fusion model, an initial three-dimensional convolution process is performed on the three-dimensional near-infrared fluorescence extended feature map, and an initial activation process is performed on the result of the initial three-dimensional convolution process to obtain an initial activation result;

[0014] The first near-infrared fluorescence convolution module is used to perform a secondary three-dimensional convolution process on the initial activation result, and the result of the secondary three-dimensional convolution process is subjected to a secondary activation process to obtain a first-scale near-infrared fluorescence feature map.

[0015] According to an embodiment of the present invention, the above-mentioned method of performing multiple rounds of three-dimensional convolution processing on the three-dimensional near-infrared fluorescence extended feature map using the trained multimodal image fusion model to obtain a multi-scale near-infrared fluorescence feature map also includes:

[0016] The second near-infrared fluorescence convolution module of the trained multimodal image fusion model is used to perform the same processing on the first-scale near-infrared fluorescence feature map as the three-dimensional near-infrared fluorescence extension feature map to obtain a second-scale near-infrared fluorescence feature map;

[0017] The third near-infrared fluorescence convolution module of the trained multimodal image fusion model is used to perform the same processing on the second-scale near-infrared fluorescence feature map as the three-dimensional near-infrared fluorescence extension feature map to obtain the third-scale near-infrared fluorescence feature map.

[0018] According to an embodiment of the present invention, the above-mentioned method of performing multi-round feature map fusion on the same scale of the multi-scale near-infrared fluorescence feature map and the multi-scale magnetic particle feature map based on the adaptive cross-attention mechanism using the trained multimodal image fusion model includes:

[0019] The first scale near-infrared fluorescence feature map and the first scale magnetic particle feature map in the multi-scale magnetic particle feature map are fused by using the first cross attention mechanism module of the trained multi-modal image fusion model to obtain a first fused feature map;

[0020] The second scale near-infrared fluorescence feature map and the second scale magnetic particle feature map in the multi-scale magnetic particle feature map are fused by using the second cross attention mechanism module of the trained multi-modal image fusion model to obtain a second fused feature map;

[0021] The third-scale near-infrared fluorescence feature map and the third-scale magnetic particle feature map in the multi-scale magnetic particle feature map are fused using the third cross-attention mechanism module of the trained multi-modal image fusion model to obtain a third fused feature map;

[0022] The third fused feature map is first connected with the third-scale magnetic particle feature map by using the trained multimodal image fusion model, and the fused feature map after the first connection is first upsampled to obtain the fused feature map after the first upsampling;

[0023] Using the trained multimodal image fusion model, the first upsampled fusion feature map is connected to the second fusion feature map twice, and the secondary connected fusion feature map is upsampled twice to obtain the second upsampled fusion feature map;

[0024] The trained multimodal image fusion model is used to connect the second upsampled fused image with the first fused feature map three times, and the three-connected fused feature map is upsampled for the third time to obtain the multi-round feature map fusion result.

[0025] According to an embodiment of the present invention, the first cross attention mechanism module of the multimodal image fusion model completed by training fuses the first-scale near-infrared fluorescence feature map with the first-scale magnetic particle feature map in the multi-scale magnetic particle feature map to obtain a first fused feature map including:

[0026] Rearranging the data of the first-scale near-infrared fluorescence characteristic map and the first-scale magnetic particle characteristic map respectively to obtain a rearranged first-scale near-infrared fluorescence characteristic map and a rearranged first-scale magnetic particle characteristic map;

[0027] The rearranged first-scale near-infrared fluorescence characteristic map and the rearranged first-scale magnetic particle characteristic map are linearly transformed to obtain a first near-infrared fluorescence multi-value matrix and a first magnetic particle multi-value matrix;

[0028] The multi-head attention mechanism is respectively used to divide the rearranged first-scale near-infrared fluorescence feature map and the rearranged first-scale magnetic particle feature map;

[0029] When the multi-head attention mechanism division is completed, a multi-head adaptive cross-attention operation is performed on the first near-infrared fluorescence multi-value matrix and the first magnetic particle multi-value matrix to obtain a first near-infrared fluorescence multi-head attention result and a first magnetic particle multi-head attention result;

[0030] Rearranging the data of the first near-infrared fluorescence multi-head attention result and the first magnetic particle multi-head attention result respectively to obtain the rearranged first near-infrared fluorescence multi-head attention result and the rearranged first magnetic particle multi-head attention result;

[0031] The rearranged first near-infrared fluorescence multi-head attention result is connected with the rearranged first-scale near-infrared fluorescence feature map to obtain the first near-infrared fluorescence cross-attention feature;

[0032] The rearranged first magnetic particle multi-head attention result and the rearranged first scale magnetic particle feature map are connected to obtain the first magnetic particle cross-attention feature, and the first near-infrared fluorescence cross-attention feature and the first magnetic particle cross-attention feature are weighted to obtain the first fusion feature map.

[0033] According to an embodiment of the present invention, when the multi-head attention mechanism is divided, the first near-infrared fluorescence multi-value matrix and the first magnetic particle multi-value matrix are subjected to multi-head adaptive cross attention operation to obtain the first near-infrared fluorescence multi-head attention result and the first magnetic particle multi-head attention result, including:

[0034] The key value matrix in the first magnetic particle multi-value matrix is ​​transposed and then multiplied with the query matrix in the first near-infrared fluorescence multi-value matrix, and the result of the matrix multiplication is normalized and activated to obtain a first near-infrared fluorescence attention score;

[0035] Multiplying the first near-infrared fluorescence attention score with the value matrix in the first magnetic particle multi-value matrix, and performing layer-normalization on the multiplied result to obtain the first near-infrared fluorescence multi-head attention result;

[0036] The key value matrix in the first near-infrared fluorescence multi-value matrix is ​​transposed and then multiplied with the query matrix in the first magnetic particle multi-value matrix, and the result of the matrix multiplication is normalized and activated to obtain the first magnetic particle attention score;

[0037] The first magnetic particle attention score is multiplied by the value matrix in the first near-infrared fluorescence multi-value matrix, and the multiplied result is layer-normalized to obtain the first magnetic particle multi-head attention result.

[0038] According to a second aspect of the present invention, a method for training a multimodal image fusion model is provided, which is applied to the above-mentioned fluorescence-magnetic particle image fusion method, and is characterized by comprising:

[0039] Based on the photon propagation model, a set of simulated two-dimensional near-infrared fluorescence images of a three-dimensional tumor phantom set in a standardized space is obtained through linear difference operations and Gaussian noise random addition operations;

[0040] Performing three-dimensional convolution processing on the point spread function of the three-dimensional magnetic particle tomographic image and the image set of the three-dimensional tumor phantom to obtain a simulated three-dimensional magnetic particle tomographic image set;

[0041] The multimodal image fusion model is used to perform multimodal fusion between simulated two-dimensional near-infrared fluorescence image sets and simulated three-dimensional magnetic particle tomography image sets to obtain a predicted fusion image;

[0042] The training loss function is used to calculate the loss value between the predicted fusion image and the true value label corresponding to the predicted fusion image, and the loss value is used to update the parameters of the multimodal image fusion model;

[0043] The multimodal fusion operation, loss value calculation operation and parameter update operation between images are iterated until all the simulated images in the simulated two-dimensional near-infrared fluorescence image set and the simulated three-dimensional magnetic particle tomography image set are processed to obtain a trained multimodal image fusion model.

[0044] According to an embodiment of the present invention, the above-mentioned simulated two-dimensional near-infrared fluorescence image set of a three-dimensional tumor phantom set in a standardized space is obtained by linear difference operation and Gaussian noise random addition operation based on the photon propagation model, including:

[0045] Setting a three-dimensional tumor phantom within a target tissue mapped into a standardized space, and converting pixel values ​​of the three-dimensional tumor phantom into concentrations of a dual-mode probe;

[0046] The concentration distribution of the dual-mode probe is discretized into a concentration vector, and Gaussian noise is randomly added to the concentration vector to obtain a concentration vector with noise;

[0047] The photon propagation model target tissue is used for forward calculation to obtain the system matrix of fluorescence imaging, and the system matrix of fluorescence imaging is operated with the vector with noise to obtain the fluorescence intensity distribution vector;

[0048] A linear interpolation operation is performed on the fluorescence intensity distribution vector, and a Gaussian noise random addition operation is performed on the fluorescence intensity distribution vector after the linear difference operation to obtain a simulated two-dimensional near-infrared fluorescence image set.

[0049] According to an embodiment of the present invention, the multimodal image fusion model trained above includes a two-dimensional near-infrared fluorescence image encoder, a three-dimensional magnetic particle tomography image encoder, a front fusion module, a first adaptive cross-attention mechanism fusion module, a second adaptive cross-attention mechanism fusion module, a third adaptive cross-attention mechanism fusion module and an image fusion decoder;

[0050] The two-dimensional near-infrared fluorescence image encoder includes a two-dimensional convolution near-infrared fluorescence expansion module, a first near-infrared fluorescence three-dimensional convolution module, a second near-infrared fluorescence three-dimensional convolution module and a third near-infrared fluorescence three-dimensional convolution module;

[0051] Wherein, the three-dimensional magnetic particle tomographic image encoder includes a first magnetic particle three-dimensional convolution module, a second magnetic particle three-dimensional convolution module and a third magnetic particle three-dimensional convolution module;

[0052] The first adaptive cross-attention mechanism fusion module, the second adaptive cross-attention mechanism fusion module and the third adaptive cross-attention mechanism fusion module each include an attention mechanism layer, a multi-layer perceptron layer, a residual connection layer and a plurality of linear transformation layers;

[0053] Among them, the image fusion decoder includes a filtering convolution module and multiple upsampling three-dimensional convolution modules.

[0054] A third aspect of the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0055] The fourth aspect of the present invention further provides a computer-readable storage medium on which a computer program or instruction is stored, and the steps of the above method are implemented when the above computer program or instruction is executed by a processor.

[0056] The present invention fuses FMI and MPI of different modalities through a trained multimodal image fusion model, making full use of the advantages of high resolution and high sensitivity of FMI in shallow tissues and the characteristics of MPI without imaging depth limitation and linear quantification. This multimodal fusion method not only makes up for the shortcomings of FMI in imaging depth, but also significantly improves the spatial resolution of MPI, thereby achieving more accurate and comprehensive imaging of target areas such as tumors. MPI can provide prior information on the location of the tumor, while FMI can provide high-quality detail information. The present invention effectively fuses these two types of information through a multimodal image fusion model based on a deep learning method, significantly improving the positioning accuracy of imaging and promoting the application of FMI technology and MPI technology in the biomedical field. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:

[0058] Figure 1 is a schematic diagram of a multimodal medical image of a tested mouse according to an embodiment of the present invention;

[0059] Figure 2 is a flow chart of a fluorescence-magnetic particle image fusion method according to an embodiment of the present invention;

[0060] Figure 3 is a schematic diagram of the structure of a CAUnet model according to an embodiment of the present invention;

[0061] Figure 4 is a schematic diagram of the structure of a pre-fusion module of a CAUnet model according to an embodiment of the present invention;

[0062] Figure 5 is a schematic diagram of the structure of a downsampling convolution module according to an embodiment of the present invention;

[0063] Figure 6 is a data processing flow chart of an adaptive cross-attention mechanism fusion module according to an embodiment of the present invention;

[0064] Figure 7is a schematic structural diagram of an upsampling module according to an embodiment of the present invention;

[0065] Figure 8 is a flowchart of a training method for a multimodal image fusion model according to an embodiment of the present invention;

[0066] Fig. 9 is a schematic diagram of a simulated multimodal image of a three-dimensional tumor phantom according to an embodiment of the present invention;

[0067] Fig.10 is a schematic diagram of a three-dimensional optical-magnetic fusion image of a CAUnet model according to an embodiment of the present invention;

[0068] Fig.11 A block diagram of an electronic device suitable for implementing a fluorescence-magnetic particle image fusion method and a multimodal image fusion model training method according to an embodiment of the present invention is schematically shown. DETAILED DESCRIPTION

[0069] Below, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of concepts of the present invention.

[0070] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.

[0071] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0072] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0073] Figure 1is a schematic diagram of a multimodal medical image of a tested mouse according to an embodiment of the present invention.

[0074] like Figure 1 As shown in the figure, the MPI image, near-infrared FMI image and CT image of the tested mouse can display the relevant information of the pathological area of ​​the tested mouse at different angles. Among them, the three-dimensional computed tomography (CT) image has the characteristics of high density resolution and high image clarity, contains rich position information, and is often used for the registration of other medical images; fluorescent molecular imaging (FMI) has the advantages of high sensitivity and high resolution, and three-dimensional magnetic particle imaging (MPI) has the advantages of strong depth penetration; FMT can provide three-dimensional information of the superficial layer of the body surface; however, the above medical images also have their own shortcomings in single-modality imaging. Since FMI and FMT have the problem of insufficient imaging depth in single-modality imaging, and MPI has technical problems such as insufficient resolution and imaging depth, it is necessary to provide a multi-modal imaging technology solution that integrates at least two modal information of FMI, FMT or MPI. Multi-modal image fusion can make full use of image information of different modalities, integrate them together, and improve the quality and information content of the image. For example, FMI and MPI can be integrated to fully combine the advantages of FMI's high resolution and high sensitivity in the shallow layer with the advantage of MPI without imaging depth limitation. MPI provides tumor location priors, and FMI provides high-quality detail information. Through deep learning methods, the spatial resolution and positioning accuracy of imaging can be improved through fusion.

[0075] In order to solve at least one of the problems of the prior art, an embodiment of the present invention provides a fluorescence-magnetic particle image fusion method and a training method for a multimodal image fusion model.

[0076] Figure 2 is a flow chart of a fluorescence-magnetic particle image fusion method according to an embodiment of the present invention.

[0077] like Figure 2 As shown, the fluorescence-magnetic particle image fusion method of this embodiment includes operations S210 to S230.

[0078] In operation S210, the trained multimodal image fusion model is used to perform two-dimensional convolution processing on the aligned two-dimensional near-infrared fluorescence image to obtain a multi-channel near-infrared fluorescence extended feature map, and the multi-channel near-infrared fluorescence extended feature map is initially fused with the aligned three-dimensional magnetic particle tomography image to obtain a three-dimensional near-infrared fluorescence extended feature map.

[0079] The multimodal image fusion model trained above is constructed based on a Unet neural network with multiple adaptive cross-attention mechanisms, namely, a CAUnet neural network.

[0080] The above-mentioned two-dimensional near-infrared fluorescence image and three-dimensional magnetic particle tomography image are both images used to characterize a target pathological region (such as a tumor) of a biological body.

[0081] After acquiring the two-dimensional near-infrared fluorescence image and the three-dimensional magnetic particle tomography image of the target pathological area, the above images need to be registered and the registered images are used for multimodal image fusion.

[0082] The above operation S210 is used to expand the feature map of the two-dimensional near-infrared fluorescence image and convert the two-dimensional image into a three-dimensional image so as to be subsequently fused with the three-dimensional magnetic particle tomography image.

[0083] In operation S220, the trained multimodal image fusion model is used to perform multiple rounds of three-dimensional convolution processing on the three-dimensional near-infrared fluorescence extended feature map to obtain a multi-scale near-infrared fluorescence feature map, and the aligned three-dimensional magnetic particle tomography image is subjected to multiple rounds of three-dimensional convolution processing to obtain a multi-scale magnetic particle feature map.

[0084] In operation S230, based on the adaptive cross-attention mechanism, the trained multimodal image fusion model is used to perform multi-round feature map fusion on the same scale for the multi-scale near-infrared fluorescence feature map and the multi-scale magnetic particle feature map, and the multi-round feature map fusion results are subjected to filtering convolution processing to obtain a three-dimensional fluorescence-magnetic particle fusion image.

[0085] It should be particularly noted that the purpose of the present invention is to achieve the fusion of multimodal medical images and provide high-quality medical images of the target pathological area, rather than to directly obtain the diagnostic information or health status of the target case area, but only to obtain and process intermediate information from the target case area; at the same time, the above operations S210~S230 and other operations of the embodiments of the present invention are all information processing operations implemented by computers and other devices.

[0086] The present invention fuses FMI and MPI of different modalities through a trained multimodal image fusion model, making full use of the advantages of high resolution and high sensitivity of FMI in shallow tissues and the characteristics of MPI without imaging depth limitation and linear quantification. This multimodal fusion method not only makes up for the shortcomings of FMI in imaging depth, but also significantly improves the spatial resolution of MPI, thereby achieving more accurate and comprehensive imaging of target areas such as tumors. MPI can provide prior information on the location of the tumor, while FMI can provide high-quality detail information. The present invention effectively fuses these two types of information through a multimodal image fusion model based on a deep learning method, significantly improving the positioning accuracy of imaging and promoting the application of FMI technology and MPI technology in the biomedical field.

[0087] Before performing the above operations S210 to S230, it is necessary to obtain a two-dimensional near-infrared fluorescence image, a three-dimensional magnetic particle tomography image, and a three-dimensional computed tomography (CT) image of the target case area (such as a tumor area) of the biological body, and align the above medical images.

[0088] The above-mentioned multimodal medical image acquisition process and multimodal medical image registration process are further described in detail below through specific embodiments.

[0089] It should be noted that before obtaining the above-mentioned medical images of the target pathological area of ​​the organism, authorization is obtained from the organism (such as the patient or the patient's close relatives) or the owner of the organism to collect the above-mentioned medical images of the target pathological area; and the above-mentioned medical images are processed with the permission of the organism or the owner of the organism. The relevant process strictly abides by the requirements of laws and regulations, adopts strict confidentiality measures, does not violate public order and good morals, and provides corresponding operation entrances for the organism or the owner of the organism to choose to authorize or refuse.

[0090] Taking the case where the target area of ​​the organism is the tumor and its surrounding tissue area as an example, by injecting a fluorescent / magnetic particle dual-modality probe into the organism under test, a two-dimensional near-infrared fluorescence image of the organism surface containing tumor information, a three-dimensional MPI tomographic image, and a CT image containing the anatomical structure information of the tissues and organs around the tumor are obtained.

[0091] The tumor and its surrounding tissue area are taken as the region of interest, and its CT image and MPI image are mapped into the standardized imaging space (SIS). The SIS is constructed in advance according to the task requirements and should ensure that the region of interest can be accommodated.

[0092] The body surface two-dimensional near-infrared fluorescence image is mapped to the SIS surface, the surface fluorescence image is matched with the surface of the CT volume data, and the fluorescence image is cropped to retain the area mapped to the SIS surface.

[0093] In some preferred embodiments, the CT image is mapped to the interior of the discretized SIS, that is: the central coordinates of the discretized SIS are used as the imaging space center of the CT image; each pixel of the preprocessed CT image is used as a voxel point, the grid node closest to the current voxel point is obtained in the discretized SIS, and the organ attribute corresponding to the current voxel point is assigned to the grid node; the voxel point corresponding to each pixel is traversed to map the preprocessed CT image to the interior of the SIS.

[0094] In some preferred embodiments, the three-dimensional MPI tomographic image is mapped into the discretized SIS, that is, a registration reference point is set, and the imaging space coordinate system of the three-dimensional MPI tomographic image is adjusted to be consistent with the imaging space coordinate system of the CT image; the resolution of the MPI three-dimensional tomographic image and the CT image are adjusted to be the same through linear interpolation.

[0095] The cropped surface fluorescence image and the three-dimensional MPI image mapped to the SIS are input into the trained fluorescence-magnetic particle image fusion model CAUnet (i.e., the trained multimodal image fusion model, the same below) for image fusion to obtain a three-dimensional optical-magnetic fusion image, i.e., the fluorescence-magnetic particle multimodal image fusion method shown in the above operations S210 to S230.

[0096] The multimodal image fusion model of the present invention is described below through specific embodiments and in conjunction with the accompanying drawings, so that those skilled in the art can clearly understand how the technical solution provided by the present invention realizes multimodal image fusion.

[0097] Figure 3 is a schematic diagram of the structure of a CAUnet model according to an embodiment of the present invention.

[0098] The trained fluorescence-magnetic particle image fusion model CAUnet is used as the multimodal image fusion model completed by the above training, such as Figure 3 As shown, the above CAUnet includes an FMI image encoder, an MPI image encoder, a front fusion module, a first adaptive cross-attention mechanism fusion module, a second adaptive cross-attention mechanism fusion module, a third adaptive cross-attention mechanism fusion module, and an image fusion decoder.

[0099] The FMI image encoder includes a two-dimensional convolution block and three three-dimensional convolution blocks connected in sequence, which serve as an FMI extension module, a first FMI convolution module, a second FMI convolution module, and a third FMI convolution module, respectively, and their outputs are an FMI extension feature map, a first FMI feature map, a second FMI feature map, and a third FMI feature map, respectively.

[0100] Among them, the FMI extension module is used to obtain the multi-channel near-infrared fluorescence extended feature map; the FMI extension module is built based on a 5×5 two-dimensional convolution layer

[0101] According to an embodiment of the present invention, the above-mentioned initial fusion of the multi-channel near-infrared fluorescence extended feature map with the aligned three-dimensional magnetic particle tomography image to obtain the three-dimensional near-infrared fluorescence extended feature map includes: using the pre-fusion module of the trained multimodal image fusion model to multiply the maximum value of the aligned three-dimensional magnetic particle tomography image on each channel with the multi-channel near-infrared fluorescence extended feature map channel by channel to obtain the three-dimensional near-infrared fluorescence extended feature map.

[0102] Figure 4 is a schematic diagram of the structure of a pre-fusion module of a CAUnet model according to an embodiment of the present invention.

[0103] The multi-channel near-infrared fluorescence extended feature map and the registered three-dimensional magnetic particle tomography image are realized by the pre-fusion module of CAUnet, where the pre-fusion module is used to fuse the FMI extended feature map with the MPI three-dimensional image of the input model. Figure 4 As shown, the input of the front fusion module is the FMI extended feature map (i.e., the multi-channel near-infrared fluorescence extended feature map) and the three-dimensional MPI image. The module is configured to take the maximum value of each slice (channel) of the input three-dimensional MPI image and multiply it channel by channel with the FMI extended feature map.

[0104] According to an embodiment of the present invention, the above-mentioned use of the trained multimodal image fusion model to perform multiple rounds of three-dimensional convolution processing on the three-dimensional near-infrared fluorescence extension feature map to obtain a multi-scale near-infrared fluorescence feature map includes: using the first near-infrared fluorescence convolution module of the trained multimodal image fusion model to perform initial three-dimensional convolution processing on the three-dimensional near-infrared fluorescence extension feature map, and performing initial activation processing on the result of the initial three-dimensional convolution processing to obtain an initial activation result; using the first near-infrared fluorescence convolution module to perform secondary three-dimensional convolution processing on the initial activation result, and performing secondary activation processing on the result of the secondary three-dimensional convolution processing to obtain a first-scale near-infrared fluorescence feature map. Using the second near-infrared fluorescence convolution module of the trained multimodal image fusion model to perform the same processing as the three-dimensional near-infrared fluorescence extension feature map on the first-scale near-infrared fluorescence feature map to obtain a second-scale near-infrared fluorescence feature map; using the third near-infrared fluorescence convolution module of the trained multimodal image fusion model to perform the same processing as the three-dimensional near-infrared fluorescence extension feature map on the second-scale near-infrared fluorescence feature map to obtain a third-scale near-infrared fluorescence feature map.

[0105] Figure 5 is a schematic diagram of the structure of a downsampling convolution module according to an embodiment of the present invention.

[0106] The above-mentioned multi-scale near-infrared fluorescence feature map is obtained by the first FMI convolution module (i.e., the first near-infrared fluorescence convolution module, the same below), the second FMI convolution module (the second near-infrared fluorescence convolution module, the same below), and the third FMI convolution module (the third near-infrared fluorescence convolution module, the same below). Since the above-mentioned first FMI convolution module, the second FMI convolution module, and the third FMI convolution module are all based on a 5×5×5 three-dimensional convolution layer with a step size of 1, a LeakyReLU activation layer, a channel step size of 1, a spatial step size of 2 (a step size of (1,2,2)), and a 3×2×2 three-dimensional convolution layer, and a LeakyReLU activation layer, which are connected in sequence, as shown in FIG. Figure 5 Therefore, obtaining the multi-scale near-infrared fluorescence feature map requires two three-dimensional convolutions and two activation processes.

[0107] The MPI image encoder is constructed based on three three-dimensional convolution blocks connected in sequence, which are respectively used as the first MPI convolution module, the second MPI convolution module, and the third MPI convolution module, and their outputs are the first MPI feature map (i.e., the first-scale magnetic particle feature map, the same below), the second MPI feature map (the second-scale magnetic particle feature map, the same below), and the third MPI feature map (the third-scale magnetic particle feature map, the same below);

[0108] The first MPI convolution module, the second MPI convolution module, and the third MPI convolution module are all constructed based on a 5×5×5 three-dimensional convolution layer with a step size of 1, a LeakyReLU activation layer, a 3×2×2 three-dimensional convolution layer with a channel step size of 1 and a spatial step size of 2 (a step size of (1,2,2)), and a LeakyReLU activation layer, which are connected in sequence. The structure is as follows: Figure 5 Therefore, the process of obtaining the first, second and third scale magnetic particle characteristic images is similar to the process of obtaining the first, second and third scale near-infrared fluorescence characteristic images.

[0109] According to an embodiment of the present invention, the above-mentioned method of performing multi-round feature map fusion on the same scale of multi-scale near-infrared fluorescence feature maps and multi-scale magnetic particle feature maps using a trained multi-modal image fusion model based on an adaptive cross-attention mechanism includes: using a first cross-attention mechanism module of the trained multi-modal image fusion model to fuse the first-scale near-infrared fluorescence feature map with the first-scale magnetic particle feature map in the multi-scale magnetic particle feature map to obtain a first fused feature map; using a second cross-attention mechanism module of the trained multi-modal image fusion model to fuse the second-scale near-infrared fluorescence feature map with the second-scale magnetic particle feature map in the multi-scale magnetic particle feature map to obtain a second fused feature map; using a third cross-attention mechanism module of the trained multi-modal image fusion model to fuse the third-scale near-infrared fluorescence feature map with the second-scale magnetic particle feature map in the multi-scale magnetic particle feature map to obtain a second fused feature map. The feature map is fused with the third-scale magnetic particle feature map in the multi-scale magnetic particle feature map to obtain a third fused feature map; the third fused feature map is connected to the third-scale magnetic particle feature map for the first time using the trained multimodal image fusion model, and the fused feature map after the first connection is upsampled for the first time to obtain the fused feature map after the first upsampling; the first upsampled fused feature map is connected to the second fused feature map for the second time using the trained multimodal image fusion model, and the fused feature map after the second connection is upsampled for the second time to obtain the fused feature map after the second upsampling; the second upsampled fused image is connected to the first fused feature map for the third time using the trained multimodal image fusion model, and the fused feature map after the third connection is upsampled for the third time to obtain the multi-round feature map fusion result.

[0110] The above-mentioned acquisition of multi-round feature map fusion results utilizes the first adaptive cross-attention mechanism fusion module, the second adaptive cross-attention mechanism fusion module, and the third adaptive cross-attention mechanism fusion module of the CAUnet model, that is, three multimodal feature map fusions are performed. Technical personnel in this field can set the number of feature map fusions according to actual needs.

[0111] The above-mentioned adaptive cross-attention mechanism fusion module is used to fuse the outputs of the FMI convolution module and the MPI convolution module. For the Nth adaptive cross-attention mechanism module, the input is the Nth FMI feature map and the Nth MPI feature map. The adaptive cross-attention mechanism fusion module includes a first adaptive cross-attention mechanism module, a second adaptive cross-attention mechanism module, and an adaptive weighting module. After the input is calculated in parallel by the first and second cross-attention mechanism modules to obtain the first attention feature and the second attention feature, the feature of the adaptive cross-attention mechanism fusion module is calculated by the adaptive weighting module, which is recorded as the Nth adaptive cross-attention feature.

[0112] According to an embodiment of the present invention, the first cross-attention mechanism module of the multimodal image fusion model that has been trained is used to fuse the first-scale near-infrared fluorescence feature map with the first-scale magnetic particle feature map in the multi-scale magnetic particle feature map to obtain the first fused feature map, including: rearranging the data of the first-scale near-infrared fluorescence feature map and the first-scale magnetic particle feature map respectively to obtain the rearranged first-scale near-infrared fluorescence feature map and the rearranged first-scale magnetic particle feature map; linearly transforming the rearranged first-scale near-infrared fluorescence feature map and the rearranged first-scale magnetic particle feature map respectively to obtain the first near-infrared fluorescence multi-value matrix and the first magnetic particle multi-value matrix; performing multi-head attention mechanism division on the rearranged first-scale near-infrared fluorescence feature map and the rearranged first-scale magnetic particle feature map respectively; when the multi-head attention mechanism division is completed, the first near-infrared fluorescence feature map is transformed into the first-scale near-infrared fluorescence feature map. A multi-head adaptive cross-attention operation is performed on the optical multi-valued matrix and the first magnetic particle multi-valued matrix to obtain a first near-infrared fluorescence multi-head attention result and a first magnetic particle multi-head attention result; data of the first near-infrared fluorescence multi-head attention result and the first magnetic particle multi-head attention result are rearranged respectively to obtain the rearranged first near-infrared fluorescence multi-head attention result and the rearranged first magnetic particle multi-head attention result; the rearranged first near-infrared fluorescence multi-head attention result is connected with the rearranged first-scale near-infrared fluorescence feature map to obtain the first near-infrared fluorescence cross-attention feature; the rearranged first magnetic particle multi-head attention result is connected with the rearranged first-scale magnetic particle feature map to obtain the first magnetic particle cross-attention feature, and a weighted operation is performed on the first near-infrared fluorescence cross-attention feature and the first magnetic particle cross-attention feature to obtain a first fusion feature map.

[0113] According to an embodiment of the present invention, when the multi-head attention mechanism is divided, the first near-infrared fluorescence multi-value matrix and the first magnetic particle multi-value matrix are subjected to multi-head adaptive cross-attention operation to obtain the first near-infrared fluorescence multi-head attention result and the first magnetic particle multi-head attention result, including: transposing the key value matrix in the first magnetic particle multi-value matrix and multiplying it with the query matrix in the first near-infrared fluorescence multi-value matrix, and performing normalization and activation processing on the result after matrix multiplication to obtain the first near-infrared fluorescence attention score; multiplying the first near-infrared fluorescence attention score with the value matrix in the first magnetic particle multi-value matrix, and performing layer normalization on the result after multiplication to obtain the first near-infrared fluorescence multi-head attention result; transposing the key value matrix in the first near-infrared fluorescence multi-value matrix and multiplying it with the query matrix in the first magnetic particle multi-value matrix, and performing normalization and activation processing on the result after matrix multiplication to obtain the first magnetic particle attention score; multiplying the first magnetic particle attention score with the value matrix in the first near-infrared fluorescence multi-value matrix, and performing layer normalization on the result after multiplication to obtain the first magnetic particle multi-head attention result.

[0114] The above embodiment illustrates the process of acquiring the first fused feature map, and the process of acquiring other fused feature maps is similar to that of the first fused feature map.

[0115] The following is a further detailed description of the acquisition process of the above fusion feature map through a specific implementation method.

[0116] Figure 6 It is a data processing flow chart of the adaptive cross-attention mechanism fusion module according to an embodiment of the present invention.

[0117] Take the first cross-attention mechanism fusion module as an example, Figure 5 As shown, it includes three linear transformation layers W q , W k , W v , attention mechanism layer, multi-layer perceptron layer, and residual connection; the input FMI features are rearranged after data is transformed by linear transformation W q Mapped to Q (Query), MPI features are rearranged and transformed into W through linear transformation k , W v Mapped to K (key) and V (Value); and divide Q, K, V into multiple heads and input into the attention mechanism layer for the calculation of the multi-head attention mechanism; the attention mechanism layer performs matrix multiplication on the transpose of matrices Q and K and uses the Softmax function to normalize to obtain the attention score, multiplies the attention score with the matrix V, and performs layer standardization; the multi-layer perceptron layer is composed of two linear layers connected in sequence and a LeakyReLU activation function; the output of the attention mechanism layer is rearranged, and residually connected with the rearranged FMI feature, and then passed through the multi-layer perceptron layer and the residual connection, and the data is rearranged to restore it to the shape of the input data to obtain the first cross attention feature.

[0118] The second cross-attention mechanism module has the same structure as the first cross-attention module, but its input order is opposite to that of the first cross-attention mechanism module, and its output is the second cross-attention feature.

[0119] The adaptive weighting module is used to calculate the sum of the weights of the first cross-attention mechanism and the second cross-attention mechanism, and is composed of an adaptive average pooling layer, a fully connected layer, and a ReLU activation function layer connected in sequence; the input of the adaptive weighting module is a three-dimensional magnetic particle image, a first cross-attention feature, and a second cross-attention feature; the three-dimensional magnetic particle image is adaptively weighted pooled in the fault (channel) dimension, and the weight of the first cross-attention feature is output after passing through the fully connected layer and the ReLU activation function layer. , the weight of the second cross-attention mechanism feature is set to , and the output of the adaptive weighted module is obtained by weighted summing them.

[0120] Figure 7 is a schematic diagram of the structure of an upsampling module according to an embodiment of the present invention.

[0121] The image fusion decoder module is constructed based on three sequentially connected three-dimensional upsampling convolution modules and a filtering convolution module, which are respectively used as the first upsampling convolution module, the second upsampling convolution module, the third upsampling convolution module, and the filtering convolution module; wherein the structure of the upsampling convolution module is as follows Figure 7 shown.

[0122] The first upsampling convolution module, the second upsampling convolution module, and the third upsampling convolution module are all constructed based on a 3×2×2 three-dimensional deconvolution layer with a channel step of 1 and a spatial step of 2 (step (1,2,2)), a LeakyReLU activation layer, a 5×5×5 three-dimensional convolution layer with a step of 1, and a LeakyReLU activation layer, as shown in Figure 1. Figure 7 shown.

[0123] A third upsampling module, whose input is the connection of the third MPI feature map and the third adaptive cross-attention feature in the channel dimension; a second upsampling module, whose input is the connection of the output of the third upsampling module and the second adaptive cross-attention feature in the channel dimension; a first upsampling module, whose input is the connection of the output of the second upsampling module and the first adaptive cross-attention feature in the channel dimension.

[0124] The filter convolution module is constructed based on a 5×5×5 three-dimensional convolution layer and a ReLU activation layer; the input of the filter convolution module is the output of the first upsampling module, and the output after the three-dimensional convolution layer and the ReLU activation layer is used as the output of the entire CAUnet.

[0125] Figure 8 4 is a flowchart of a method for training a multimodal image fusion model according to an embodiment of the present invention.

[0126] like Figure 8 As shown, the training method of the multimodal image fusion model is applied to the fluorescence-magnetic particle image fusion method, including operations S810 to S850.

[0127] In operation S810, based on a photon propagation model, a set of simulated two-dimensional near-infrared fluorescence images of a three-dimensional tumor phantom set in a standardized space is obtained through a linear difference operation and a Gaussian noise random addition operation.

[0128] In operation S820, a three-dimensional convolution process is performed on the point spread function of the three-dimensional magnetic particle tomographic image and the image set of the three-dimensional tumor phantom to obtain a simulated three-dimensional magnetic particle tomographic image set.

[0129] In operation S830, a multimodal image fusion model is used to perform multimodal fusion between the simulated two-dimensional near-infrared fluorescence image set and the simulated three-dimensional magnetic particle tomography image set to obtain a predicted fusion image.

[0130] In operation S840, a loss value between the predicted fusion image and a true value label corresponding to the predicted fusion image is calculated using a training loss function, and a parameter of the multimodal image fusion model is updated using the loss value.

[0131] In operation S850, the multimodal fusion operation, loss value calculation operation and parameter update operation between images are iterated until all simulated images in the simulated two-dimensional near-infrared fluorescence image set and the simulated three-dimensional magnetic particle tomography image set are processed to obtain a trained multimodal image fusion model.

[0132] The multimodal image fusion model obtained through the training of the above operations S810 to S850 has good generalization, and can be applied to and is not limited to the multimodal image fusion of tumor pathological areas; at the same time, the multimodal image fusion model obtained through the training of the above operations has good fusion accuracy and high image resolution, and can be used to assist relevant personnel to accurately judge the status of the pathological area.

[0133] According to an embodiment of the present invention, the multimodal image fusion model trained above includes a two-dimensional near-infrared fluorescence image encoder, a three-dimensional magnetic particle tomography image encoder, a front fusion module, a first adaptive cross-attention mechanism fusion module, a second adaptive cross-attention mechanism fusion module, a third adaptive cross-attention mechanism fusion module and an image fusion decoder; wherein, the two-dimensional near-infrared fluorescence image encoder includes a two-dimensional convolution near-infrared fluorescence expansion module, a first near-infrared fluorescence three-dimensional convolution module, a second near-infrared fluorescence three-dimensional convolution module and a third near-infrared fluorescence three-dimensional convolution module; wherein, the three-dimensional magnetic particle tomography image encoder includes a first magnetic particle three-dimensional convolution module, a second magnetic particle three-dimensional convolution module and a third magnetic particle three-dimensional convolution module; wherein, the first adaptive cross-attention mechanism fusion module, the second adaptive cross-attention mechanism fusion module and the third adaptive cross-attention mechanism fusion module all include an attention mechanism layer, a multi-layer perceptron layer, a residual connection layer and multiple linear transformation layers; wherein, the image fusion decoder includes a filtering convolution module and multiple upsampling three-dimensional convolution modules.

[0134] According to an embodiment of the present invention, the above-mentioned method based on the photon propagation model, through linear difference operation and Gaussian noise random addition operation, obtains the simulated two-dimensional near-infrared fluorescence image set of the three-dimensional tumor phantom set in the standardized space, including: setting the three-dimensional tumor phantom in the target tissue mapped to the standardized space, and converting the pixel value of the three-dimensional tumor phantom into the concentration of the dual-mode probe; discretizing the concentration distribution of the dual-mode probe into a concentration vector, and randomly adding Gaussian noise to the concentration vector to obtain a noisy concentration vector; using the photon propagation model target tissue to perform forward calculation to obtain a system matrix of fluorescence imaging, and operating the system matrix of fluorescence imaging with the vector with noise to obtain a fluorescence intensity distribution vector; performing a linear interpolation operation on the fluorescence intensity distribution vector, and performing a Gaussian noise random addition operation on the fluorescence intensity distribution vector after the linear difference operation to obtain a simulated two-dimensional near-infrared fluorescence image set.

[0135] The present invention describes in detail the process of acquiring a simulated two-dimensional near-infrared fluorescence image set and a simulated three-dimensional magnetic particle tomography image set through the following specific implementation methods.

[0136] Fig. 9 is a schematic diagram of a simulated multimodal image of a three-dimensional tumor phantom according to an embodiment of the present invention.

[0137] A 3D tumor phantom is set up within the tissue mapped into the SIS, its pixel values ​​are converted into the concentration of the dual-mode probe, and the in vivo probe concentration distribution is discretized into a vector , and add Gaussian noise.

[0138] In this specific embodiment, the three-dimensional tumor phantoms provided include single tumor phantoms and double tumor phantoms; single tumor phantoms include spherical phantoms, clustered phantoms, and MNIST phantoms; the diameter of the single tumor phantom is randomly set to 0.9-4.5 mm; the clustered phantom is set by superimposing spherical phantoms that overlap each other in 3-5 positions; the MNIST phantom is obtained by randomly enlarging the MNIST data set image by 1-3 times, and mapping the coordinates to SIS after expanding the channel dimension to a random number of 6-20; the double tumor phantom is obtained by superimposing two single tumor phantoms.

[0139] In this specific embodiment, the pixel value of the tumor phantom is converted into the concentration of the dual-mode probe, and the method is shown in formula (1):

[0140] (1),

[0141] in, is the dual-mode probe concentration signal at pixel point r, is the gray value of the pixel at pixel point r, and is the set maximum dual-mode probe concentration and maximum pixel grayscale value, where The preferred setting is 5×10 7 mmol / L, The preferred setting is 255. The tumor phantom image is as follows Fig. 9 shown.

[0142] The system matrix is ​​obtained by forward model calculation of tissue according to the photon propagation model ; Using linear equations Calculate the fluorescence intensity distribution vector of the tissue surface nodes ;right Linear interpolation is performed to obtain a simulated surface fluorescence image, and Gaussian noise is added to obtain a simulated fluorescence image such as Fig. 9 shown.

[0143] In this embodiment, the system matrix is ​​obtained by performing forward model calculation on the tissue according to the photon propagation model. The method is as follows: assume that the excitation light source is located on the upper surface of the SIS; describe the propagation process of the fluorescent photons in the imaging object tissue by the diffusion approximation equation described by formula (2) and describe the refractive index deviation between the object surface and the air by the Robin boundary condition shown in formula (3):

[0144] (2),

[0145] (3),

[0146] Wherein, x and m represent the excitation light and the emission light, respectively. Indicates location The photon flux density at . represents the diffusion coefficient, where , is the absorption coefficient, , is the scattering coefficient, is the anisotropy coefficient; is the intensity of the excitation light, is the position of the excitation light. Indicates location The fluorescence yield. Represents the boundary of biological tissue, is the normal vector of the biological tissue surface, is the optical refractive index deviation between the imaging object boundary and the air. Solving equations (2) and (3) yields the system matrix.

[0147] The point spread function of the three-dimensional MPI is calculated and three-dimensionally convolved with the tumor phantom image to obtain a simulated three-dimensional MPI image.

[0148] In this embodiment, the method for calculating the point spread function of the three-dimensional MPI is: setting the diameter of the magnetic particles (preferably 20 nm), the temperature (preferably 20° C.), the saturation magnetization (preferably ), maximum magnetic particle concentration (preferably 5×107mmol / L), select magnetic field gradient (preferably randomly 2.6T / m×1.3T / m×1.3T / m to 4T / m×2T / m×2T / m, where the field gradients in the y and z directions are the same and half of the field gradients in the x direction), set coil sensitivity (preferably 1.0), imaging field size (preferably set to 20mm×20mm×15mm), calculate the point spread function of the three-dimensional MPI according to formula (4), as shown in formula (4):

[0149] (4),

[0150] in, , , , is the selected field gradient in the x, y, and z directions, , where is the Boltzmann constant, is the vacuum permeability, is the temperature, is the particle magnetic moment, is the Lagrangian function, is the derivative of the Lagrangian function; the simulated three-dimensional MPI image is as follows Fig. 9 shown.

[0151] The simulated surface FMI image and the simulated three-dimensional MPI image are used as training samples, and the set tumor phantom is used as a true value label.

[0152] During the model training process, the loss value of the output result of the multimodal image fusion model and the corresponding true value label is calculated, and the model parameters of the CAUnet model are updated.

[0153] In the embodiment of the present invention, the loss function is preferably formed by weighting the Dice coefficient loss function and the mean square error (MSE) loss function, wherein the Dice coefficient loss weight is 0.9 and the MSE loss function weight is 0.1, the loss value of the model is calculated, back propagation is performed, the parameters are adjusted, and the parameters of the CAUnet model are updated. The optimization algorithm used is Adaptive Moment Estimation (Adam).

[0154] In this embodiment, the model is trained in a cycle until a set number of pre-training times or a set accuracy is reached, then the training is terminated to obtain a trained CAUnet model.

[0155] Fig.10 is a schematic diagram of a three-dimensional optical-magnetic fusion image output by a CAUnet model according to an embodiment of the present invention.

[0156] In this embodiment, the collected two-dimensional near-infrared fluorescence image, three-dimensional magnetic particle image, and CT image data are preprocessed, and the preprocessed data are input into the trained CAUnet model for multimodal image fusion processing to obtain the following: Fig.10 The three-dimensional optical-magnetic fusion image shown.

[0157] Fig.11 A block diagram of an electronic device suitable for implementing a fluorescence-magnetic particle image fusion method and a multimodal image fusion model training method according to an embodiment of the present invention is schematically shown.

[0158] like Fig.11 As shown, the electronic device 1100 according to an embodiment of the present invention includes a processor 1101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1102 or a program loaded from a storage part 1108 to a random access memory (RAM) 1103. The processor 1101 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1101 may also include an onboard memory for caching purposes. The processor 1101 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0159] In RAM 1103, various programs and data required for the operation of electronic device 1100 are stored. Processor 1101, ROM 1102 and RAM 1103 are connected to each other via bus 1104. Processor 1101 performs various operations of the method flow according to the embodiment of the present invention by executing the program in ROM 1102 and / or RAM 1103. It should be noted that the program can also be stored in one or more memories other than ROM 1102 and RAM 1103. Processor 1101 can also perform various operations of the method flow according to the embodiment of the present invention by executing the program stored in the one or more memories.

[0160] According to an embodiment of the present invention, the electronic device 1100 may further include an input / output (I / O) interface 1105, which is also connected to the bus 1104. The electronic device 1100 may further include one or more of the following components connected to the input / output (I / O) interface 1105: an input portion 1106 including a keyboard, a mouse, etc.; an output portion 1107 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 1108 including a hard disk, etc.; and a communication portion 1109 including a network interface card such as a LAN card, a modem, etc. The communication portion 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the input / output (I / O) interface 1105 as needed. A removable medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1110 as needed, so that a computer program read therefrom is installed into the storage portion 1108 as needed.

[0161] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiment; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present invention is implemented.

[0162] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM 1102 and / or RAM 1103 described above and / or one or more memories other than ROM 1102 and RAM 1103.

[0163] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0164] It will be appreciated by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention may be combined and / or combined in various ways. All of these combinations and / or combinations fall within the scope of the present invention.

[0165] The embodiments of the present invention are described above. However, these embodiments are only for the purpose of illustration, and are not intended to limit the scope of the present invention. Although each embodiment is described above, it does not mean that the measures in each embodiment cannot be used in combination advantageously. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.

Claims

1. A fluorescence-magnetic particle image fusion method, characterized in that: The method comprises: Using the trained multimodal image fusion model, a two-dimensional convolution process is performed on the registered two-dimensional near-infrared fluorescence image to obtain a multi-channel near-infrared fluorescence extended feature map, and the multi-channel near-infrared fluorescence extended feature map is initially fused with the registered three-dimensional magnetic particle tomography image to obtain a three-dimensional near-infrared fluorescence extended feature map; Using the trained multimodal image fusion model, the three-dimensional near-infrared fluorescence extended feature map is subjected to multiple rounds of three-dimensional convolution processing to obtain a multi-scale near-infrared fluorescence feature map, and the registered three-dimensional magnetic particle tomography image is subjected to multiple rounds of three-dimensional convolution processing to obtain a multi-scale magnetic particle feature map; Based on the adaptive cross-attention mechanism, the trained multimodal image fusion model is used to perform multi-round feature map fusion on the same scale for the multi-scale near-infrared fluorescence feature map and the multi-scale magnetic particle feature map, and the multi-round feature map fusion results are subjected to filtering convolution processing to obtain a three-dimensional fluorescence-magnetic particle fusion image.

2. The method according to claim 1, characterized in that Initially fusing the multi-channel near-infrared fluorescence extended feature map with the registered three-dimensional magnetic particle tomography image to obtain a three-dimensional near-infrared fluorescence extended feature map includes: The pre-fusion module of the trained multimodal image fusion model is used to multiply the maximum value of the registered three-dimensional magnetic particle tomography image on each channel with the multi-channel near-infrared fluorescence extension feature map channel by channel to obtain the three-dimensional near-infrared fluorescence extension feature map.

3. The method according to claim 1, characterized in that The trained multimodal image fusion model is used to perform multiple rounds of three-dimensional convolution processing on the three-dimensional near-infrared fluorescence extended feature map to obtain a multi-scale near-infrared fluorescence feature map including: Using the first near-infrared fluorescence convolution module of the trained multimodal image fusion model to perform initial three-dimensional convolution processing on the three-dimensional near-infrared fluorescence extended feature map, and performing initial activation processing on the result of the initial three-dimensional convolution processing to obtain an initial activation result; The first near-infrared fluorescence convolution module is used to perform a secondary three-dimensional convolution process on the initial activation result, and a secondary activation process is performed on the result of the secondary three-dimensional convolution process to obtain a first-scale near-infrared fluorescence feature map.

4. The method according to claim 3, characterized in that Also includes: Using the second near-infrared fluorescence convolution module of the trained multimodal image fusion model, the first-scale near-infrared fluorescence feature map is processed in the same way as the three-dimensional near-infrared fluorescence extension feature map to obtain a second-scale near-infrared fluorescence feature map; The third near-infrared fluorescence convolution module of the trained multimodal image fusion model is used to perform the same processing on the second-scale near-infrared fluorescence feature map as the three-dimensional near-infrared fluorescence extension feature map to obtain a third-scale near-infrared fluorescence feature map.

5. The method according to claim 4, characterized in that Based on the adaptive cross attention mechanism, using the trained multimodal image fusion model to perform multi-round feature map fusion on the same scale on the multi-scale near-infrared fluorescence feature map and the multi-scale magnetic particle feature map includes: Using the first cross attention mechanism module of the trained multimodal image fusion model, the first-scale near-infrared fluorescence feature map and the first-scale magnetic particle feature map in the multi-scale magnetic particle feature map are fused to obtain a first fused feature map; Using the second cross attention mechanism module of the trained multimodal image fusion model, the second-scale near-infrared fluorescence feature map and the second-scale magnetic particle feature map in the multi-scale magnetic particle feature map are fused to obtain a second fused feature map; Using the third cross attention mechanism module of the trained multimodal image fusion model, the third-scale near-infrared fluorescence feature map and the third-scale magnetic particle feature map in the multi-scale magnetic particle feature map are fused to obtain a third fused feature map; The third fusion feature map is first connected with the third-scale magnetic particle feature map by using the trained multimodal image fusion model, and the fusion feature map after the first connection is first upsampled to obtain the fusion feature map after the first upsampling; Using the trained multimodal image fusion model, the first upsampled fusion feature map is connected twice with the second fusion feature map, and the second upsampled fusion feature map is upsampled twice to obtain a second upsampled fusion feature map; The trained multimodal image fusion model is used to connect the second upsampled fused image with the first fused feature map three times, and the fused feature map after the three connections is upsampled for a third time to obtain the multi-round feature map fusion result.

6. The method according to claim 5, characterized in that The first-scale near-infrared fluorescence feature map and the first-scale magnetic particle feature map in the multi-scale magnetic particle feature map are fused using the first cross-attention mechanism module of the trained multi-modal image fusion model to obtain a first fused feature map including: Rearranging the data of the first-scale near-infrared fluorescence characteristic map and the first-scale magnetic particle characteristic map respectively to obtain a rearranged first-scale near-infrared fluorescence characteristic map and a rearranged first-scale magnetic particle characteristic map; Linearly transforming the rearranged first-scale near-infrared fluorescence characteristic map and the rearranged first-scale magnetic particle characteristic map to obtain a first near-infrared fluorescence multi-value matrix and a first magnetic particle multi-value matrix; Performing multi-head attention mechanism division on the rearranged first-scale near-infrared fluorescence feature map and the rearranged first-scale magnetic particle feature map respectively; When the multi-head attention mechanism division is completed, the first near-infrared fluorescence multi-value matrix and the first magnetic particle multi-value matrix are subjected to a multi-head adaptive cross-attention operation to obtain a first near-infrared fluorescence multi-head attention result and a first magnetic particle multi-head attention result; Rearranging the data of the first near-infrared fluorescence multi-head attention result and the first magnetic particle multi-head attention result respectively to obtain a rearranged first near-infrared fluorescence multi-head attention result and a rearranged first magnetic particle multi-head attention result; Connecting the rearranged first near-infrared fluorescence multi-head attention result with the rearranged first-scale near-infrared fluorescence feature map to obtain a first near-infrared fluorescence cross-attention feature; The rearranged first magnetic particle multi-head attention result and the rearranged first-scale magnetic particle feature map are connected to obtain the first magnetic particle cross-attention feature, and the first near-infrared fluorescence cross-attention feature and the first magnetic particle cross-attention feature are weighted to obtain the first fusion feature map.

7. The method according to claim 6, characterized in that When the multi-head attention mechanism is divided, the first near-infrared fluorescence multi-value matrix and the first magnetic particle multi-value matrix are subjected to multi-head adaptive cross attention operation to obtain the first near-infrared fluorescence multi-head attention result and the first magnetic particle multi-head attention result, including: Transposing the key value matrix in the first magnetic particle multi-value matrix and multiplying it with the query matrix in the first near-infrared fluorescence multi-value matrix, and performing normalization activation processing on the result of the matrix multiplication to obtain a first near-infrared fluorescence attention score; Multiplying the first near-infrared fluorescence attention score by the value matrix in the first magnetic particle multi-value matrix, and performing layer normalization on the multiplied result to obtain the first near-infrared fluorescence multi-head attention result; Transposing the key value matrix in the first near-infrared fluorescence multi-value matrix and multiplying it with the query matrix in the first magnetic particle multi-value matrix, and performing normalization activation processing on the result of the matrix multiplication to obtain a first magnetic particle attention score; The first magnetic particle attention score is multiplied by the value matrix in the first near-infrared fluorescence multi-value matrix, and the multiplied result is layer-normalized to obtain the first magnetic particle multi-head attention result.

8. A method for training a multimodal image fusion model, applied to any of the methods of claims 1 to 7, characterized in that: The method comprises: Based on the photon propagation model, a set of simulated two-dimensional near-infrared fluorescence images of a three-dimensional tumor phantom set in a standardized space is obtained through linear difference operations and Gaussian noise random addition operations; Performing three-dimensional convolution processing on the point spread function of the three-dimensional magnetic particle tomographic image and the image set of the three-dimensional tumor phantom to obtain a simulated three-dimensional magnetic particle tomographic image set; Using the multimodal image fusion model, multimodal fusion is performed on the simulated two-dimensional near-infrared fluorescence image set and the simulated three-dimensional magnetic particle tomography image set to obtain a predicted fusion image; Calculating a loss value between the predicted fusion image and a true value label corresponding to the predicted fusion image by using a training loss function, and updating parameters of the multimodal image fusion model by using the loss value; The multimodal fusion operation, the loss value calculation operation and the parameter update operation between images are iterated until all the simulated images in the simulated two-dimensional near-infrared fluorescence image set and the simulated three-dimensional magnetic particle tomography image set are processed to obtain the trained multimodal image fusion model.

9. The method according to claim 8, characterized in that: Based on the photon propagation model, through linear difference operation and Gaussian noise random addition operation, the simulated two-dimensional near-infrared fluorescence image set of the three-dimensional tumor phantom set in the standardized space is obtained, including: Setting the three-dimensional tumor phantom in the target tissue mapped into the standardized space, and converting the pixel values ​​of the three-dimensional tumor phantom into the concentration of the dual-mode probe; discretizing the concentration distribution of the dual-mode probe into a concentration vector, and randomly adding Gaussian noise to the concentration vector to obtain a concentration vector with noise; Performing forward calculation on the target tissue using the photon propagation model to obtain a system matrix of fluorescence imaging, and performing operation on the system matrix of fluorescence imaging and the vector with noise to obtain a fluorescence intensity distribution vector; A linear interpolation operation is performed on the fluorescence intensity distribution vector, and a Gaussian noise random addition operation is performed on the fluorescence intensity distribution vector after the linear difference operation, so as to obtain the simulated two-dimensional near-infrared fluorescence image set.

10. The method according to claim 8, characterized in that The trained multimodal image fusion model includes a two-dimensional near-infrared fluorescence image encoder, a three-dimensional magnetic particle tomography image encoder, a front fusion module, a first adaptive cross-attention mechanism fusion module, a second adaptive cross-attention mechanism fusion module, a third adaptive cross-attention mechanism fusion module and an image fusion decoder; Wherein, the two-dimensional near-infrared fluorescence image encoder includes a two-dimensional convolution near-infrared fluorescence expansion module, a first near-infrared fluorescence three-dimensional convolution module, a second near-infrared fluorescence three-dimensional convolution module and a third near-infrared fluorescence three-dimensional convolution module; Wherein, the three-dimensional magnetic particle tomographic image encoder comprises a first magnetic particle three-dimensional convolution module, a second magnetic particle three-dimensional convolution module and a third magnetic particle three-dimensional convolution module; The first adaptive cross-attention mechanism fusion module, the second adaptive cross-attention mechanism fusion module and the third adaptive cross-attention mechanism fusion module each include an attention mechanism layer, a multi-layer perceptron layer, a residual connection layer and a plurality of linear transformation layers; Wherein, the image fusion decoder includes a filtering convolution module and a plurality of upsampling three-dimensional convolution modules.

Citation Information

Patent Citations

  • Fluorescent molecular tomography reconstruction method based on magnetic particle imaging prior guidance

    CN114581553A

  • Multi-mode imaging system for small animal magnetic particle imaging and fluorescence molecular tomography

    CN115844365A

  • Video human body behavior recognition method based on spatial-temporal feature enhancement network

    CN116912727A

  • MRI image segmentation system and method based on mixed attention supervision U-shaped network

    CN118587442A

  • Magnetic particle imaging (MPI) and fluorescence molecular tomography (FMT)-fused multimodal imaging system for small animal

    US11940508B1

Cited By

  • Near-infrared two-region fluorescence-magnetic particle-CT (Computed Tomography) three-mode fusion imaging method

    CN120876566A

  • A near-infrared two-zone fluorescence-magnetic particle-CT three-modality fusion imaging method

    CN120876566B

  • Living body detection method and system based on multi-modal imaging integration analysis

    CN121059098A

  • Self-supervised structured light microscopy reconstruction method and system based on pixel rearrangement

    JP2026509603A

  • Self-supervised structured light microscopy reconstruction method and system based on pixel rearrangement

    JP7866807B2