PET-CT (positron emission tomography-computed tomography) multi-modal image fusion method, system, equipment and medium
By performing spatial alignment of PET-CT images and feature extraction and frequency domain decomposition processing in the FEMFM model of multimodal image fusion, the image fusion quality problem under low contrast or non-obvious morphological features in the prior art is solved, and high-quality PET-CT image fusion is achieved.
Patent Information
- Application Number
- CN202510140817.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-30
AI Technical Summary
The existing PET-CT image fusion method is difficult to obtain high-quality fusion images when the source image contrast is low or the morphological characteristics are not obvious.
After obtaining the PET image and CT image for spatial alignment, it is input into the multimodal image fusion FEMFM model, and local and global features are extracted using the position convolution PCB module, combined with the even complex wavelet transform DTCWT and the Gaussian-Laplace operator LoG decomposition to extract low-frequency and high-frequency information, and fused through the local gradient energy fusion LGE module.
The quality of PET-CT image fusion is improved, the accuracy of lesion recognition is enhanced, and the fusion image is clearer and more diagnostic value while maintaining the key information of the source image.
Smart Images

Figure CN120071065A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and particularly relates to a PET-CT multimodal image fusion method, system, device and medium. Background Art
[0002] Computed tomography (CT) is widely used to analyze high-resolution anatomical structures. However, the information provided by CT cannot directly reflect the metabolic activities of tissues. On the other hand, positron emission tomography (PET) provides detailed metabolic information that is helpful for diagnosing and analyzing tumors. However, PET is a low-resolution modality and lacks the anatomical structure of the structure. Multimodal medical images (such as PET / CT) play an important role in modern clinical diagnosis, because they combine the metabolic information of PET images and the anatomical structure information of CT images, providing more comprehensive support for early disease diagnosis, tumor staging and treatment evaluation.
[0003] Image fusion refers to fusing the image information of different modalities into a single-modal image. In the field of clinical diagnosis, medical imaging plays a crucial role in tumor diagnosis. Due to the differences in imaging parameters and imaging mechanisms of different imaging technologies, different medical images produce different imaging results on the same cells, tissues, bones, organs and systems. With the development of image fusion technology, multimodal medical image fusion has become increasingly important. By fusing different-modal medical images, information complementarity can be achieved, and redundant information between different-modal medical images can be reduced.
[0004] Existing medical image fusion methods mainly fall into two categories: traditional fusion methods and deep learning-based methods. Traditional fusion methods can be divided into multi-scale transformation methods, sparse representation methods, fuzzy logic methods, hybrid methods, subspace methods, etc. Traditional image fusion algorithms usually use digital transformation algorithms to convert images into the spatial domain or transform domain, then perform activity level measurements in the spatial domain or transform domain, and manually design fusion rules to achieve image fusion. Finally, inverse transformation operations are used to reconstruct the fused image. Traditional medical image fusion methods have achieved good fusion effects. However, traditional fusion methods are not robust, have weak generalization ability, and are difficult to optimize. When dealing with large-scale image data, manually designed algorithms require a large amount of computing resources and time.
[0005] In recent years, due to the powerful data mining ability of neural networks, deep learning (DL) has attracted great attention in the field of medical image processing. The CNN-based fusion framework mainly designs the CNN network structure and loss function, extracts features from multiple input images, fuses the features, and finally reconstructs the fused image. This framework has the ability to automatically learn the fusion rules, so there is no need to manually design complex fusion rules. However, in existing medical image fusion methods, the quality of the fused image depends to a large extent on the quality of the source images. When the contrast of the source images is low and the morphological features are not obvious, it is difficult to obtain high-quality fused images. Summary of the Invention
[0006] To solve the deficiencies in existing PET-CT image fusion, the present invention provides a PET-CT multimodal image fusion method, system, device, and medium.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] A PET-CT multimodal image fusion method, comprising the following steps:
[0009] Obtain the PET image and CT image data of the object to be measured, and spatially align the PET image and CT image;
[0010] Input the aligned CT image and PET image into the multimodal image fusion FEMFM model. The position convolution PCB module in the FEMFM model extracts the local and global features of the CT image and PET image; extract the low-frequency information of the aligned PET image through the dual-tree complex wavelet transform DTCWT, decompose the high-frequency information of the aligned CT image through the Laplacian of Gaussian operator LoG, and input the low-frequency information of the PET image and the high-frequency information of the CT image into the convolutional pooling module CPB for processing to obtain PET low-frequency features and CT high-frequency features; input the local and global features, as well as the PET low-frequency features and CT high-frequency features, into the local gradient energy fusion LGE module for fusion, and output the PET-CT fused image after hierarchical processing.
[0011] Preferably, the position convolution PCB module performs feature extraction, specifically including the following steps:
[0012] First perform 1×1×1 convolution processing on the input aligned CT image and PET image, and generate two separate outputs including a first branch and a second branch through the pointwise convolution module PWConv. The branch includes the GCC-H module for extracting horizontal feature information and the GCC-V module for extracting vertical feature information;
[0013] The first output of the first branch passes through GCC-H that captures horizontal feature information, and the second output passes through GCC-V that extracts vertical information; the first output of the second branch passes through GCC-V that extracts vertical information, and the second output passes through GCC-H that captures horizontal feature information. The results of the four outputs are concatenated together, and Pre-Norm is used to normalize the concatenated features;
[0014] The normalized features are passed through a feed-forward network FFN. After the channel-level attention mechanism is processed in parallel with the FFN, the image after feature extraction is output through the residual.
[0015] Preferably, the local gradient energy fusion LGE module performs fusion, which specifically includes the following steps:
[0016] For the input image I 1 (x, y), I 2 (x, y), I 3 (x, y) and I 4 (x, y), the local gradient is calculated using a gradient operator. The calculation formula for the gradient magnitude G(x, y) is:
[0017]
[0018] where (x, y) is the pixel position; the local gradient energy is the result of summing the gradient magnitude G(x, y) within a certain window W, and the local gradient energy E(x, y) is expressed as: W is an n×n window used to calculate the gradient energy around the current pixel;
[0019] Different weights are assigned to images of different modalities, and the weights are calculated according to the local gradient energy of each image; E 1 (x, y), E 2 (x, y), E 3 (x, y) and E 4 (x, y) respectively represent the local gradient energy of the images I 1 , I 2 , I 3 and I 4 at the position (x, y). The weights are specifically calculated through the following formula:
[0020]
[0021]
[0022] After obtaining the weights, the images are fused, and the fused image F(x, y) is specifically:
[0023]
[0024] Preferably, before spatially aligning the PET image and the CT image, normalizing the PET image and the CT image respectively is also included;
[0025] Converting the pixel values of the CT image into HU values and normalizing them includes the following steps:
[0026] Calculating the HU value of each pixel in the CT image through the following formula:
[0027]
[0028] Using the maximum-minimum normalization method to normalize the HU values of the CT image, specifically through the following formula:
[0029]
[0030] where CT is the pixel value of the original CT image, CT max and CT min are the maximum and minimum values of the CT image respectively, and CT norm is the pixel value of the normalized CT image;
[0031] Mapping the pixel values of the normalized CT image to the original image to obtain the normalized CT image;
[0032] Converting the radioactivity uptake values of the PET image into SUV values and normalizing them includes the following steps:
[0033] Calculating the SUV value of each pixel in the PET image, specifically through the following formula:
[0034]
[0035] Using the Z-score normalization method to normalize the SUV values of the PET image, specifically through the following formula:
[0036]
[0037] where PET is the pixel value of the original PET image, μ is the pixel mean of the PET image, σ is the pixel standard deviation of the PET image, and PET norm is the pixel value of the normalized PET image;
[0038] Mapping the pixel values of the normalized PET image to the original image to obtain the normalized PET image;
[0039] Spatially aligning the normalized CT image and the normalized PET image.
[0040] Preferably, the normalized CT image and the normalized PET image are spatially aligned through the following steps:
[0041] Obtain the affine matrices of the CT image and the PET image after normalization processing;
[0042] Convert the coordinates of the original PET and CT images to the standard reference coordinates, calculate the affine matrix between the standard reference coordinates and the target image coordinates; through matrix multiplication, convert the images to the coordinates of the target image;
[0043]
[0044] Calculate the position of each voxel in the new coordinate system through resampling, and recalculate the values of these positions using the cubic interpolation algorithm. The size of the resampled image is calculated through the following formula:
[0045] x = (a 1 × b 1 ) / c 1 ;
[0046] y = (a 2 × b 2 ) / c 2 ;
[0047] z = (a 3 × b 3 ) / c 3 ;
[0048] where x and y respectively represent the number of columns and rows of the resampled image; a 1 , a 2 and a 3 respectively represent the pixel spacings of the x-axis, y-axis and z-axis of the original data; b 1 , b 2 and b 3 respectively represent the length, width and height of the original data; c 1 , c 2 and c 3 respectively represent the pixel spacings of the x-axis, y-axis and z-axis after resampling required.
[0049] The present invention also provides a PET-CT multimodal image fusion system, specifically including:
[0050] A data processing module, configured to obtain the PET image and CT image data of the object to be measured; and spatially align the PET image and the CT image;
[0051] An image fusion module is used to input the aligned CT image and PET image into a multimodal image fusion FEMFM model. The local and global features of the CT image and PET image are extracted by the position convolution PCB module in the FEMFM model; the low-frequency information of the aligned PET image is extracted by the dual-tree complex wavelet transform DTCWT, and the high-frequency information of the aligned CT image is decomposed by the Laplacian of Gaussian operator LoG. The low-frequency information of the PET image and the high-frequency information of the CT image are respectively input into the convolutional pooling module CPB for processing to obtain PET low-frequency features and CT high-frequency features; the local and global features, as well as the PET low-frequency features and CT high-frequency features, are input into the local gradient energy fusion LGE module for fusion, and after hierarchical processing, a PET-CT fusion image is output.
[0052] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps in the above-mentioned PET-CT multimodal image fusion method.
[0053] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is loaded by a processor, it can execute the steps in the above-mentioned PET-CT multimodal image fusion method.
[0054] The PET-CT multimodal image fusion method provided by the present invention has the following beneficial effects:
[0055] In the present invention, the obtained PET image and CT image are aligned. The quality of the original images is improved through the aligned images to facilitate feature extraction. The aligned images are input into the multimodal image fusion FEMFM model, and feature extraction is performed by the position convolution PCB module in the FEMFM model to extract richer features from the input images. The DTCWT transform and LoG operator decomposition are combined to extract the low-frequency information of PET and the high-frequency information of CT respectively, highlighting the metabolic information in the PET image and retaining the anatomical structure in the CT image at the same time. The PET low-frequency features and CT high-frequency features are input into the LGE module for fusion to enhance the complementary information between PET and CT and improve the fusion quality. After hierarchical processing, a PET-CT fusion image is output. The accuracy of lesion recognition is improved, and while the fusion image retains the key information of the source images, the fusion quality is further enhanced, making it clearer and more diagnostically valuable. Description of the Drawings
[0056] To more clearly illustrate the embodiments of the present invention and their design solutions, the accompanying drawings required for this embodiment will be briefly introduced below. The accompanying drawings in the following description are only partial embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0057] Figure 1 It is a flowchart of a PET-CT multimodal image fusion method of the present invention.
[0058] Figure 2 It is a preprocessing flowchart in the embodiment of the present invention. Among them, Figure 2 (a) is the preprocessing process of the PET image; Figure 2 (b) is the preprocessing process of the CT image.
[0059] Figure 3 It is a structural diagram of the FEMFM model of the present invention.
[0060] Figure 4 It is a structural diagram of the PCB module of the present invention.
[0061] Figure 5 It is a structural diagram of the GCC-V module in the embodiment of the present invention.
[0062] Figure 6 It is a structural diagram of the GCC-H module in the embodiment of the present invention. Detailed implementation manners
[0063] To enable those skilled in the art to better understand the technical solutions of the present invention and be able to implement them, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0064] Embodiment
[0065] The present invention provides a PET-CT multimodal image fusion method, as Figure 1 shown, which specifically includes the following steps:
[0066] Step 1: Parse the DICOM file storing PET / CT image data to obtain PET image and CT image data.
[0067] Step 2: Preprocess the PET image and CT image data, as Figure 2 shown, Figure 2 (b) is to convert the pixel values of the CT image to HU values and perform maximum-minimum normalization processing.
[0068] Calculate the HU value (Hounsfield Units) of each pixel in the CT image to characterize the density information of different tissue structures in the CT image, specifically through the following formula:
[0069]
[0070] Adopt the maximum-minimum normalization method to normalize the HU values of the CT image, specifically through the following formula:
[0071]
[0072] where CT is the pixel value of the original CT image, CT max and CT min are the maximum and minimum values of the CT image respectively, and CT norm is the pixel value of the normalized CT image.
[0073] Figure 2 (a) of is to convert the radioactive uptake value of the PET image into the SUV value and perform Z-score normalization. Calculate the SUV value (Standard Uptake Value) of each pixel in the PET image to characterize the uptake degree of the tissue in the PET image for the radioactive tracer, specifically through the following formula:
[0074]
[0075] Adopt the Z-score normalization method to normalize the SUV values of the PET image, specifically through the following formula:
[0076]
[0077] where PET is the pixel value of the original PET image, μ is the pixel mean value of the PET image, σ is the pixel standard deviation of the PET image, and PET nor m is the pixel value of the normalized PET image.
[0078] Step 3: Unify the voxel spacings of the normalized CT image and the PET image to align the spaces of the images to be fused. Specifically through the following steps:
[0079] S31. To calculate the affine transformation and determine the target space, obtain the affine matrices of the normalized CT image and the PET image respectively.
[0080] S32. Since the voxel spacings of PET and CT are inconsistent, to spatially align the PET image and the CT image, the coordinates of the original PET and CT images are transformed into the standard reference coordinates through the following formula, and then the affine matrix between the standard reference coordinates and the target image coordinates is calculated. The image is transformed to the target image through matrix multiplication.
[0081]
[0082] S33. Since the affine transformation determines a new coordinate system, the actual image data needs to be re-interpolated in the new coordinate system. Resampling calculates the positions of each voxel in the new coordinate system and recalculates the values at these positions using the cubic interpolation algorithm. Finally, the size of the resampled image is calculated through the following formula:
[0083] x = (a 1 ×b 1 ) / c 1 , y = (a 2 ×b 2 ) / c 2 , z = (a 3 ×b 3 ) / c 3 ;
[0084] where x and y represent the number of columns and rows of the resampled image respectively; a 1 , a 2 and a 3 represent the pixel spacings of the x-axis, y-axis, and z-axis of the original data respectively; b 1 , b 2 and b 3 represent the length, width, and height of the original data respectively; c 1 , c 2 and c 3 represent the pixel spacings of the x-axis, y-axis, and z-axis after resampling respectively.
[0085] Step 4: Use the frequency feature enhanced multi-modal image fusion model (Frequency feature Enhanced Multmodel image Fusion Model, FEMFM), such as Figure 3As shown, the fused image is output. PET and CT are two separate branches. After feature extraction by the PCB module, the low-frequency information of the PET image is extracted using DTCWT, and the high-frequency information of the CT image is decomposed using LoG. Both pass through a Convolution and Pooling Block (CPB) to adjust the corresponding number of channels and dimensions. When fusing in the LEG module, the complementary information of PET and CT is enhanced to improve the fusion quality. The CPB module consists of 1×1×1 convolution and pooling operations. After the last fusion, the image size is restored, and finally, a fused image with both high-resolution structural information and accurate functional information is output.
[0086] S41. Input the normalized and spatially aligned PET and CT images into the FEMFM model for fusion. This model adopts an encoder-decoder architecture and introduces a Position Convolution Block (PCB) to extract richer local and global features, thereby improving the quality of the fused image.
[0087] As Figure 4 shown, the input data is first processed through 1×1×1 convolution. Pointwise convolution (PWConv) is applied to the features, generating two separate outputs, which will be processed differently along the horizontal and vertical axes. The first output of PWConv passes through GCC-H, which captures horizontal feature information. The second output is passed through GCC-V to extract vertical information. Similarly, the second branch passes the features in the reverse order. The results of these four branches (GCC-H and GCC-V in both directions) are concatenated together to combine the extracted horizontal and vertical features. The concatenated features are normalized using Pre-Norm to standardize the feature distribution before further processing. Then, the normalized features are passed through a feed-forward network (FFN), which typically involves fully connected layers to learn higher-level representations. A channel-level attention mechanism is applied in parallel with the FFN to focus on the most important channels of the features. Throughout the module, there are residual connections where the input at different stages is added back to the output, helping to maintain the gradient flow during training and improve learning. Final output: After the final residual addition, the processed features are passed to the next layer for further processing.
[0088] S42. Utilize the metabolic information of the PET image and the structural information of CT to extract the low-frequency features of the PET image through the Dual-Tree Complex Wavelet Transform (DTCWT), and extract the high-frequency features of the CT image through the Laplacian of Gaussian (LoG) to enhance the anatomical structure details and metabolic information of the fused image.
[0089] As Figure 5 shown, GCC-V is a vertical convolution method. The dual-tree complex wavelet transform is a signal processing method based on wavelet transform, which can effectively extract the amplitude and phase information of the signal, and at the same time has good direction selectivity and frequency band separation characteristics. It is an implementation of the complex wavelet transform, which overcomes the disadvantages of the discrete wavelet transform such as insufficient directionality and frequency band aliasing.
[0090] The DTCWT implements the complex wavelet transform through two parallel wavelet transform trees (one for the real part and the other for the imaginary part). Specifically, by performing two sets of real wavelet transforms on the input signal, one as the main filter bank and the other as the auxiliary filter bank, the real and imaginary parts of the complex wavelet coefficients are generated respectively, so as to obtain approximate directionality and the properties of the complex coefficients.
[0091] Assume that ψ(t) is a wavelet basis function, and the input signal x(t) is decomposed through two wavelet transform trees to obtain the coefficients in complex form:
[0092] W DTCWT (x)(a,b) = W A (x)(a,b) + jW B (x)(a,b);
[0093] where: W A (x)(a,b) represents the real part wavelet coefficient generated by TreeA, and W B (x)(a,b) represents the imaginary part wavelet coefficient generated by TreeB, and j is the imaginary unit.
[0094] For scale a and b translations, TreeA and TreeB calculate the approximation coefficients and detail coefficients respectively through filters:
[0095]
[0096] where ψ(A) and ψ(B) are the wavelet basis functions in TreeA and TreeB respectively.
[0097] As Figure 6 shown, GCC-H is a horizontal convolution. First, the Gaussian filter is used to smooth the image to remove noise, and then the Laplacian operator is applied to detect the edges in the image. Performing the smoothing process first can reduce the interference of noise on edge detection and make the edge detection more stable.
[0098] The Laplacian of Gaussian operator can be defined through convolution operation, and its mathematical expression is:
[0099]
[0100] Where: I (x,y) represents the input image, and G (x,y) represents the Gaussian kernel function, which is defined as:
[0101]
[0102] where σ is the standard deviation of the Gaussian kernel, controlling the smoothness degree, is the Laplace operator, which is usually expressed as:
[0103] Combining the Gaussian kernel and the Laplace operator, the Gaussian-Laplace kernel function is obtained:
[0104]
[0105] The receptive fields of GCC-V and GCC-H cover all the pixels in the same column and the same row respectively. Using them jointly can extract global features from all input pixels.
[0106] S43. Respectively perform convolution and pooling operations on the extracted PET low-frequency feature map and CT high-frequency feature map, and then fuse the PET and CT image features, as well as the low-frequency feature (LF) of the PET image and the high-frequency feature (HF) of the CT image, through the Local Gradient Energy (LGE) fusion module. The same processing is performed on the last three layers in the encoding stage.
[0107] For the input image I 1 (x, y), I 2 (x, y), I 3 (x, y) and I 4 (x, y), the gradient operator is used to calculate the local gradient. The calculation formula for the gradient magnitude G(x, y) is:
[0108]
[0109] where (x, y) is the pixel position; the local gradient energy is the result of summing the gradient magnitude G(x, y) within a certain window W, and the local gradient energy E(x, y) is expressed as: W is an n×n window used to calculate the gradient energy around the current pixel.
[0110] Different weights are assigned to images of different modalities, and the weights are calculated according to the local gradient energy of each image; E 1 (x, y), E 2 (x, y), E 3 (x, y) and E 4 (x, y) respectively represent the images I 1 、I2 , I 3 and I 4 The local gradient energy at position (x, y), and the weights are specifically calculated by the following formula:
[0111]
[0112] After obtaining the weights, the images are fused, and the fused image F(x, y) is specifically:
[0113]
[0114] S44. The feature map after encoder fusion is upsampled using the decoder while further extracting features, and finally the fused feature map is restored to the same size as the original image, thereby realizing the fusion of the two modalities of PET-CT and obtaining the fused image.
[0115] Advantages of the present invention: 1. Enhanced detail information: Through the effective fusion of high and low frequency information, the contrast and details of the lesion area are significantly enhanced, which helps to improve the diagnostic accuracy of doctors. 2. Noise suppression and robustness: Frequency domain decomposition and LGE fusion suppress imaging noise and improve image quality. 3. Automated processing to improve efficiency: Automated deep learning and frequency domain fusion methods improve the efficiency of image processing and reduce human intervention. 4. High adaptability and versatility: This method is not only applicable to the fusion of PET and CT images, but can also be extended to other medical imaging modalities, with high versatility.
[0116] The present invention also provides a PET-CT multimodal image fusion system, which specifically includes:
[0117] A data processing module, configured to obtain PET image and CT image data of a to-be-detected object; and perform spatial alignment on the PET image and the CT image.
[0118] An image fusion module, configured to input the aligned CT image and PET image into the multimodal image fusion FEMFM model. The position convolution PCB module in the FEMFM model extracts local and global features of the CT image and the PET image; the low-frequency information of the aligned PET image is extracted by the dual-tree complex wavelet transform DTCWT, and the high-frequency information of the aligned CT image is decomposed by the Laplacian of Gaussian operator LoG. The low-frequency information of the PET image and the high-frequency information of the CT image are respectively input into the convolutional pooling module CPB for processing to obtain PET low-frequency features and CT high-frequency features; the local and global features, as well as the PET low-frequency features and CT high-frequency features, are input into the local gradient energy fusion LGE module for fusion, and after hierarchical processing, a PET-CT fused image is output.
[0119] Each module in the above PET-CT multimodal image fusion system can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0120] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps in an embodiment of a PET-CT multimodal image fusion method. For the specific implementation method, reference can be made to the method embodiment, which will not be elaborated here.
[0121] Furthermore, the present invention also provides a non-transitory computer-readable storage medium containing instructions, on which a computer program is stored. For example, a memory containing instructions. The above instructions can be executed by a processor of a computer device to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc. When the computer program is executed by the processor, it can implement the steps in an embodiment of a PET-CT multimodal image fusion method. For the specific implementation method, reference can be made to the method embodiment, which will not be elaborated here.
[0122] Those skilled in the art should understand that the embodiments of the present invention can provide methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0123] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0124] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the block or blocks.
[0125] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the block or blocks.
[0126] It should be noted that the above-described specific embodiments can enable those skilled in the art to more comprehensively understand the present invention, but do not limit the present invention in any way. Therefore, although the present specification and embodiments have described the present invention in detail, those skilled in the art should understand that the present invention can still be modified or equivalently replaced; and all technical solutions and improvements that do not depart from the spirit and scope of the present invention are covered by the protection scope of the patent of the present invention. Any reference numeral in the claims should not be construed as limiting the claimed claim. Any simple variation or equivalent replacement of the technical solutions that can be obviously obtained by any person skilled in the art within the technical scope disclosed by the present invention falls within the protection scope of the present invention.
Claims
1. A PET-CT multimodal image fusion method, characterized in that: The following steps are involved: Acquire PET image and CT image data of the object to be tested, and spatially align the PET image and the CT image; Input the aligned CT image and PET image into the multimodal image fusion FEMFM model, and extract the local and global features of the CT image and PET image by the position convolution PCB module in the FEMFM model; The low-frequency information of the aligned PET image is extracted by dual complex wavelet transform DTCWT, and the high-frequency information of the aligned CT image is decomposed by Gaussian-Laplacian operator LoG. The low-frequency information of the PET image and the high-frequency information of the CT image are respectively input into the convolution pooling module CPB for processing to obtain PET low-frequency features and CT high-frequency features; the local and global features as well as the PET low-frequency features and CT high-frequency features are input into the local gradient energy fusion LGE module for fusion, and the PET-CT fusion image is output after step-by-step processing.
2. A PET-CT multimodal image fusion method according to claim 1, characterized in that: The position convolution PCB module performs feature extraction, specifically including the following steps: The aligned CT image and PET image are first subjected to 1×1×1 convolution processing, and two separate outputs are generated through a point-by-point convolution module PWConv, including a first branch and a second branch, wherein the branches include a GCC-H module for extracting horizontal feature information and a GCC-V module for extracting vertical feature information; The first output of the first branch passes through GCC-H that captures horizontal feature information, and the second output passes through GCC-V that extracts vertical information; the first output of the second branch passes through GCC-V that extracts vertical information, and the second output passes through GCC-H that captures horizontal feature information. The results of the four outputs are connected together, and the concatenated features are normalized using Pre-Norm. The normalized features are transmitted through the feed-forward network FFN, and after the channel-level attention mechanism is processed in parallel with the FFN, the feature-extracted image is output through the residual.
3. The PET-CT multimodal image fusion method according to claim 1, characterized in that: The local gradient energy fusion LGE module performs fusion, specifically including the following steps: The gradient operator is used to calculate the local gradient of the input images I1(x,y), I2(x,y), I3(x,y) and I4(x,y). The calculation formula of the gradient amplitude G(x,y) is: Where (x, y) is the pixel position; the local gradient energy is the result of summing the gradient amplitude G(x, y) within a certain window W. The local gradient energy E(x, y) is expressed as: W is an n×n window used to calculate the gradient energy around the current pixel; Different weights are assigned to images of different modalities, and the weights are calculated based on the local gradient energy of each image; E1(x, y), E2(x, y), E3(x, y), and E4(x, y) represent the local gradient energies of images I1, I2, I3, and I4 at position (x, y), respectively. The weights are calculated using the following formula: After obtaining the weights, the images are fused to obtain the fused image F(x,y):
4. The PET-CT multimodal image fusion method according to claim 1, characterized in that: Before spatially aligning the PET image and the CT image, the method further includes normalizing the PET image and the CT image respectively; The pixel values of the CT image are converted into HU values and normalized, including the following steps: The HU value of each pixel in the CT image is calculated using the following formula: The maximum-minimum normalization method is used to normalize the HU value of the CT image, specifically through the following formula: Among them, CT is the pixel value of the original CT image, CT max and CT min are the maximum and minimum values of the CT image, respectively. norm is the normalized pixel value of the CT image; Mapping the normalized CT image pixel values to the original image to obtain a normalized CT image; The conversion of the radioactive uptake value of the PET image into the SUV value and normalization processing include the following steps: The SUV value of each pixel in the PET image is calculated using the following formula: The Z-score normalization method is used to normalize the SUV value of the PET image, specifically through the following formula: Where PET is the pixel value of the original PET image, μ is the pixel mean of the PET image, σ is the pixel standard deviation of the PET image, and PET norm is the normalized PET image pixel value; Mapping the normalized PET image pixel values to the original image to obtain a normalized PET image; The normalized CT image and the normalized PET image are spatially aligned.
5. A PET-CT multimodal image fusion method according to claim 4, characterized in that: The normalized CT image and the normalized PET image are spatially aligned by the following steps: Obtaining the affine matrix of the normalized CT image and the normalized PET image; Convert the original PET and CT image coordinates to the standard reference coordinates, calculate the affine matrix between the standard reference coordinates and the target image coordinates; convert the image to the target image coordinates through matrix multiplication; The position of each voxel in the new coordinate system is calculated by resampling, and the values of these positions are recalculated using the cubic interpolation algorithm. The resampled image size is calculated using the following formula: x=(a1×b1) / c1; y = (a2 × b2) / c2; z=(a3×b3) / c3; Where x and y represent the number of columns and rows of the resampled image, respectively; a1, a2, and a3 represent the pixel spacing of the x-axis, y-axis, and z-axis of the original data, respectively; b1, b2, and b3 represent the length, width, and height of the original data, respectively; c1, c2, and c3 represent the pixel spacing of the x-axis, y-axis, and z-axis after resampling, respectively.
6. A PET-CT multimodal image fusion system, characterized in that: include: A data processing module, used for acquiring PET image and CT image data of the object to be tested; and performing spatial alignment on the PET image and the CT image; An image fusion module, used for inputting the aligned CT image and PET image into a multimodal image fusion FEMFM model, and extracting local and global features of the CT image and PET image by a position convolution PCB module in the FEMFM model; The low-frequency information of the aligned PET image is extracted by dual complex wavelet transform DTCWT, and the high-frequency information of the aligned CT image is decomposed by Gaussian-Laplacian operator LoG. The low-frequency information of the PET image and the high-frequency information of the CT image are respectively input into the convolution pooling module CPB for processing to obtain PET low-frequency features and CT high-frequency features; the local and global features as well as the PET low-frequency features and CT high-frequency features are input into the local gradient energy fusion LGE module for fusion, and the PET-CT fusion image is output after step-by-step processing.
7. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is loaded into a processor, it can execute the steps of the method according to any one of claims 1 to 5.