Method and system for generating PET (positron emission tomography) image based on MRI (magnetic resonance imaging) image of multi-scale spectrum alignment
By employing a multi-scale spectral alignment method, low-frequency structures and high-frequency textures of MRI and PET are explicitly distinguished. Combined with Transformer structure and adaptive wavelet modulation, the artifact and structural distortion problems in cross-modal generation are solved, achieving high-quality PET image generation suitable for the auxiliary diagnosis of neurodegenerative diseases and tumors.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG UNIV
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-21
AI Technical Summary
Existing diffusion models fail to explicitly distinguish the different roles of low-frequency structures and high-frequency textures in cross-modal generation tasks from MRI to PET, resulting in false edges and structural distortion in the generated images. Furthermore, the lack of differentiation in time-step modulation methods affects the stability and consistency of the generation.
A multi-scale spectral alignment method is adopted, which explicitly distinguishes the low-frequency backbone and high-frequency sub-band functions through a wavelet pyramid framework. It combines the Transformer structure to perform cross-scale frequency domain feature fusion and introduces an adaptive wavelet modulation mechanism to dynamically balance low-frequency stability and high-frequency detail recovery during the diffusion process.
It effectively reduces high-frequency artifacts and structural distortion, improves the consistency of metabolic distribution and physical rationality of generated images, and enhances the stability and naturalness of the generation process. It is suitable for auxiliary diagnosis and disease assessment of neurodegenerative diseases and tumors.
Smart Images

Figure CN121904233A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing technology, and in particular to a method and system for generating PET images from MRI images based on multi-scale spectral alignment. Background Technology
[0002] Accurate diagnosis and assessment of major diseases such as neurodegenerative diseases and tumors heavily rely on the comprehensive analysis of multimodal medical imaging. Magnetic Resonance Imaging (MRI) provides high-resolution anatomical information, offering significant advantages in characterizing tissue morphology, boundaries, and structural changes. Positron Emission Tomography (PET) reflects the metabolic and functional state at the tissue or cellular level through tracer distribution, playing a crucial role in early disease detection, staging assessment, and efficacy monitoring. However, PET imaging typically suffers from high costs, the need for radioactive tracer injections, and stringent equipment deployment and maintenance requirements, limiting its application in high-frequency follow-up, long-term monitoring, and large-scale population screening. In contrast, MRI offers advantages such as no ionizing radiation, high repeatability, high imaging resolution, and widespread availability, making it one of the most widely used imaging techniques in clinical practice. Therefore, if PET images can be generated from widely available MRI images, it will not only help reduce the cost of clinical diagnosis and follow-up, but also provide an effective alternative for the auxiliary diagnosis and evaluation of neurodegenerative diseases and tumors in scenarios where PET imaging is limited or difficult to obtain, and has important clinical application value.
[0003] In the past, cross-modal medical image generation methods employed Generative Adversarial Networks (GANs) or their variants to directly learn the mapping relationship from MRI to PET in pixel space. However, these methods rely on adversarial training mechanisms, resulting in an unstable overall optimization process prone to problems such as modality collapse, making it difficult to meet the stability and consistency requirements of medical image generation.
[0004] In recent years, diffusion models (DM) have gradually become the mainstream technology in the field of medical image generation due to their advantages in probabilistic modeling and generation stability. Diffusion models model the target distribution through a progressive noise addition and denoising process, and have significant advantages over GANs in terms of generation quality and training stability. They have been widely used in cross-modal medical image synthesis tasks.
[0005] However, most existing diffusion modeling methods model directly in pixel space, lacking explicit characterization of the physical semantic differences between different frequency bands. In the MRI-to-PET cross-modal generation task, MRI and PET exhibit significant asymmetry in imaging mechanisms and spectral distribution: MRI contains a large amount of high-frequency texture information related to anatomical structures, while PET images are mainly characterized by smooth metabolic distribution dominated by low frequencies. Existing diffusion models typically process different frequency components in a uniform manner, failing to distinguish the different roles of low-frequency structural consistency and high-frequency texture suppression in the generation mechanism. This leads to the indiscriminate preservation of high-frequency structural information during denoising, introducing false edges or structural distortions, affecting the physical rationality and diagnostic value of the generated PET images. Furthermore, the generation process of diffusion models is essentially a dynamic process that evolves gradually over time, but the time-step conditionalization in existing methods usually acts on the feature representation in the form of global additive or scale-offset modulation, lacking explicit binding with spectral structure. This makes it difficult to characterize the differentiated requirements of low-frequency metabolic morphological stability and high-frequency detail control at different diffusion stages, further limiting the effectiveness of diffusion models in cross-modal MRI-to-PET generation tasks.
[0006] In research addressing the aforementioned issues, some studies have attempted to shift the modeling focus from the spatial domain to the frequency domain, aiming to improve cross-modal generation by constraining information from different frequency bands. These methods largely rely on frequency domain analysis based on Fourier transforms, which only characterizes the global frequency distribution of the signal and struggles to preserve spatial location information and directional structure. In medical image generation tasks, this can easily weaken the structural representation of local anatomical boundaries and lesion regions.
[0007] To overcome the aforementioned shortcomings, recent medical image generation methods have incorporated discrete wavelet transform (DWT) to represent images as low-frequency structural components and high-frequency detail components in multiple directions, thus balancing frequency information and spatial locality to some extent. Representative works such as WDM (Wavelet Diffusion Models) and cWDM (Conditional Wavelet Diffusion Models) stitch together the sub-bands (LL, LH, HL, HH) after wavelet decomposition along the channel dimension and input them as a whole variable into a diffusion model for joint modeling, then reconstruct the output result through inverse wavelet transform. However, these methods still treat different frequency bands as equal variables for uniform prediction in their modeling strategy, failing to explicitly distinguish the different physical semantics and mechanisms of action of low-frequency structural information and high-frequency details in cross-modal generation. Therefore, it is difficult to perform targeted modeling and constraints on the essential differences in low-frequency metabolic distribution and high-frequency structural texture between MRI and PET.
[0008] Furthermore, although WDM and cWDM have migrated the diffusion process to the wavelet domain, their injection of time-step information still follows the global modulation strategy in traditional diffusion models. This involves encoding the diffusion time step into a one-dimensional embedding vector and uniformly applying it to all wavelet sub-band features using channel-level additive bias or scale-shift. This time-step modulation method essentially imposes the same evolutionary constraints on low-frequency and high-frequency components, implicitly assuming that different frequency bands have consistent response characteristics to the time step during diffusion. However, in the MRI-to-PET cross-modal generation task, the low-frequency sub-band mainly carries stable anatomical structures and metabolic distribution trends, while the high-frequency sub-band corresponds more to the detailed textures in MRI that are not entirely consistent with those in PET. Their evolutionary requirements during diffusion differ significantly. Existing methods fail to apply differentiated modulation to different frequency bands in the time dimension, preventing the multi-scale characteristics of wavelet representation from being fully utilized in diffusion modeling, thus further limiting its potential in cross-modal spectral alignment and generation stability. Summary of the Invention
[0009] To address the shortcomings of existing technologies, this invention proposes a method and system for generating PET images from MRI images based on multi-scale spectral alignment. Given the different imaging mechanisms of MRI and PET, their corresponding spectral distribution characteristics differ significantly. This invention reconstructs a cross-modal generation paradigm from two levels: model structure and generation mechanism. Unlike existing methods that use wavelet subbands as equivalent variables for joint prediction, this invention introduces wavelet decomposition into the network as a structural inductive constraint. Within the wavelet pyramid framework, it explicitly distinguishes the functional roles of low-frequency backbone components and directional high-frequency subbands. This allows low-frequency components to undertake the main modeling tasks of cross-modal metabolic morphology and global consistency, while high-frequency components participate only in detail compensation and structural constraints in a controlled manner, thus forming a spectral redistribution path of "low-frequency dominance and high-frequency modulation." Simultaneously, by combining a hierarchical Transformer structure with global modeling capabilities, cross-scale and cross-regional consistency modeling capabilities are introduced on top of the local frequency-sensitive representation provided by wavelet decomposition. This avoids unconstrained leakage of structural high-frequency information during generation, which could affect the physical appearance consistency of PET images. Furthermore, this invention addresses the common coarse-grained modulation approach in existing diffusion models, which typically employs full-channel bias or scale shift for time-step conditionalization. Instead, it proposes an Adaptive Wavelet Modulation (AdaWM) mechanism bound to local spectral units. This mechanism directly applies diffusion time-step information to the local wavelet spectral organization level, enabling the model to prioritize stabilizing low-frequency metabolic structures and suppressing redundant high-frequency interference in the early stages of diffusion. In later stages, necessary detail corrections are introduced in a controlled manner, thus achieving coordinated regulation of the time step and spectral response throughout the entire diffusion evolution process. Through this structured spectral decoupling and time-frequency coupled modeling, this invention achieves spectral tissue transfer from MRI to PET at the mechanistic level, effectively reducing the risk of high-frequency artifacts and structural distortion, and improving the reliability of the generated results in terms of metabolic distribution consistency and physical plausibility. To achieve the above objectives, this invention employs the following technical solution: In a first aspect, the present invention provides a method for generating PET images from MRI images based on multi-scale spectral alignment, comprising the following steps: Acquire MRI images and their corresponding PET images, and preprocess the acquired images; Multi-scale wavelet transforms were performed on MRI and PET images to obtain low-frequency subbands at each scale and high-frequency subbands in multiple directions. A wavelet diffusion bridge generation model is constructed, comprising an encoder, a decoder, and a diffusion bridge module. The encoder, based on a Transformer structure, is used to model the low-frequency subbands of MRI images at multiple scales and to pass the high-frequency subbands at each scale to the decoder via skip connections. The decoder is used to acquire the low-frequency features output by the encoder and the high-frequency subbands at the same scale passed by the skip connections, and directly reconstructs the predicted PET wavelet subbands to generate PET images through inverse wavelet transform. The diffusion bridge module is used to transition PET wavelet features to MRI wavelet features in the multi-scale wavelet feature space through a forward diffusion process, and to recursively predict PET wavelet features based on the MRI wavelet features as conditions through a reverse generation process. The model is trained using a comprehensive loss function. The trained wavelet diffusion bridge model is then used to sample the input MRI image and generate the corresponding PET image.
[0010] As an alternative implementation method, multi-scale wavelet transform is performed on both MRI and PET images, specifically as follows: Two-dimensional discrete wavelet transform is performed on MRI and PET images respectively to obtain low-frequency sub-bands and high-frequency sub-bands in three directions. The low-frequency sub-bands and high-frequency sub-bands in the three directions of each image are then spliced together in the channel dimension to form a multi-channel frequency domain feature representation.
[0011] As an alternative implementation, the wavelet diffusion bridge generation model adopts a U-shaped structure. In this model, the encoding path performs wavelet transform on the low-frequency sub-bands layer by layer, inputs the newly generated low-frequency sub-bands into the Transformer module of the current layer for feature modeling, and buffers the corresponding high-frequency sub-bands. The decoding path performs inverse wavelet transform on the upsampled low-frequency backbone features and the buffered high-frequency sub-bands to recover the new low-frequency backbone features, and then inputs them into the Transformer module of the corresponding layer for feature modeling.
[0012] As an alternative implementation, the Transformer module integrates a time-domain and frequency-domain adaptive modulation mechanism. This mechanism encodes the time step of the diffusion process into modulation parameters and performs affine transformations on the features at the current scale to achieve dynamic adjustment of low-frequency structural stability and high-frequency detail recovery.
[0013] As an alternative implementation, the model also includes a high-frequency refinement module, located in the bottleneck layer of the encoder and decoder, for selectively suppressing or enhancing high-frequency subbands from various scales to reduce high-frequency textures in the MRI source modality that do not match the PET modality.
[0014] As an alternative implementation, the diffusion bridge module employs a recursive prediction mechanism in the reverse generation process, iteratively updating the estimation of PET wavelet features multiple times within the same time step.
[0015] In a second aspect, the present invention provides a system for generating PET images from MRI images based on multi-scale spectral alignment, comprising: The data acquisition and preprocessing module is configured to acquire MRI images and their corresponding PET images, and to preprocess the acquired images. The wavelet transform module is configured to perform multi-scale wavelet transforms on MRI and PET images respectively to obtain low-frequency sub-bands at each scale and high-frequency sub-bands in multiple directions. The image reconstruction module is configured to: construct a wavelet diffusion bridge generation model, which includes an encoder, a decoder, and a diffusion bridge module; the encoder, based on a Transformer structure, is used to perform multi-scale feature modeling on the low-frequency subbands of the MRI image and to pass the high-frequency subbands of each scale to the decoder through skip connections; the decoder is used to obtain the low-frequency features output by the encoder and the high-frequency subbands of the same scale passed by the skip connections, and directly reconstruct the predicted PET wavelet subbands to generate a PET image through inverse wavelet transform; the diffusion bridge module is used to transition the PET wavelet features to MRI wavelet features in the multi-scale wavelet feature space through a forward diffusion process, and to recursively predict the PET wavelet features based on the MRI wavelet features as conditions through a reverse generation process; The image output module is configured to: train the model using a comprehensive loss function, generate a model using the trained wavelet diffusion bridge, sample the input MRI image, and generate the corresponding PET image.
[0016] Thirdly, the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.
[0017] Fourthly, the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in the first aspect.
[0018] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention proposes a method for generating PET images from MRI images based on multi-scale spectral alignment. It explicitly separates the low-frequency structure and high-frequency texture of MRI and PET images through wavelet multi-scale decomposition, effectively alleviating frequency domain mismatch and avoiding artifacts and structural distortion. Cross-scale frequency domain feature fusion and alignment are achieved based on wavelet pyramids and Transformer structures, improving the structural consistency and metabolic fidelity of the generated images. A time-step-driven frequency domain modulation mechanism is introduced, enabling the model to dynamically balance low-frequency stability and high-frequency detail recovery during diffusion, improving the stability and naturalness of the generation process. High-frequency refinement modules selectively suppress high-frequency noise unrelated to the PET modality, enhancing the modal specificity and diagnostic value of the generated images. A diffusion bridge model is constructed in the wavelet domain, reducing computational complexity and improving inference efficiency, making it suitable for high-resolution clinical image generation scenarios. The PET images generated by this invention exhibit excellent performance in terms of structural fidelity, metabolic distribution consistency, and visual realism, showing significant advantages over existing technologies. They can be widely applied in the auxiliary diagnosis, disease assessment, and clinical follow-up of various diseases, including neurodegenerative diseases and tumors, demonstrating promising clinical application prospects.
[0020] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0021] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0022] Figure 1 This is a flowchart of the method for generating PET images from MRI images based on multi-scale spectral alignment according to the present invention; Figure 2 This is an overall framework diagram of the system for generating PET images from MRI images based on multi-scale spectral alignment according to the present invention. Figure 3 This is a framework diagram of the Transformer encoding and decoding structure based on wavelet multi-scale pyramids of the present invention; Figure 4 This is a schematic diagram of the internal structure of the wavelet Transformer module of the present invention. Detailed Implementation
[0023] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0024] It should be noted that the following detailed description is exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0025] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments of the invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form as well. Furthermore, it should be understood that the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion, for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but includes other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0027] Terminology Explanation: Spectrum alignment refers to the coordination and constraint of energy distribution and structural response of subbands of different directions and scales in the multi-scale wavelet domain, rather than numerical matching for a single frequency component.
[0028] Example 1 like Figures 1 to 4 As shown, this embodiment provides a method for generating PET images from MRI images based on multi-scale spectral alignment, including the following steps: Acquire MRI images and their corresponding PET images, and preprocess the acquired images; Multi-scale wavelet transforms were performed on MRI and PET images to obtain low-frequency subbands at each scale and high-frequency subbands in multiple directions. A wavelet diffusion bridge generation model is constructed, comprising an encoder, a decoder, and a diffusion bridge module. The encoder, based on a Transformer structure, is used to model the low-frequency subbands of MRI images at multiple scales and to pass the high-frequency subbands at each scale to the decoder via skip connections. The decoder is used to acquire the low-frequency features output by the encoder and the high-frequency subbands at the same scale passed by the skip connections, and directly reconstructs the predicted PET wavelet subbands to generate PET images through inverse wavelet transform. The diffusion bridge module is used to transition PET wavelet features to MRI wavelet features in the multi-scale wavelet feature space through a forward diffusion process, and to recursively predict PET wavelet features based on the MRI wavelet features as conditions through a reverse generation process. The model is trained using a comprehensive loss function. The trained wavelet diffusion bridge model is then used to sample the input MRI image and generate the corresponding PET image.
[0029] The specific solution of the present invention is as follows: This invention provides a method for generating PET images from MRI images based on multi-scale spectral alignment. By introducing a frequency domain modeling mechanism, it addresses the spectral misalignment problem that traditional spatial domain generation methods easily encounter during cross-modal conversion, reducing the inherent differences between MRI and PET in low-frequency structural distribution and high-frequency texture representation, thereby improving the modal consistency and physical realism of the generated images. Specifically, a frequency domain modeling framework based on wavelet multi-scale decomposition is proposed. This involves explicitly decomposing low-frequency and high-frequency sub-bands of MRI and PET images respectively, constructing a frequency domain representation capable of simultaneously capturing global structure and multi-directional details. Furthermore, a multi-scale Transformer structure is constructed based on wavelet pyramids, jointly modeling low-frequency backbone features and multi-scale high-frequency sub-band features during encoding and decoding to achieve layer-by-layer alignment of cross-modal spectral differences. In addition, an adaptive wavelet modulation mechanism (AdaWM) is introduced into the Transformer backbone. By combining the diffusion time step with wavelet domain features, the frequency response at different diffusion stages is dynamically adjusted. This allows the model to focus on stabilizing low-frequency structures in the early stages of diffusion and gradually recover necessary high-frequency details in the later stages, thus obtaining a spectral evolution process consistent with PET imaging characteristics. Furthermore, in the bottleneck region of the wavelet pyramid, this embodiment proposes a high-frequency refinement module. By selectively suppressing and enhancing high-frequency subbands in each direction, it effectively removes structural high-frequency textures in MRI that do not belong to the PET modality, while retaining diagnostically significant edge and contour information, thus improving the modality specificity of generated signs. Finally, the predicted wavelet subbands are reconstructed into a spatial domain PET image through inverse wavelet transform. A joint optimization objective is constructed by combining frequency domain reconstruction loss, spatial domain content loss, adversarial loss, and high-frequency sparsity loss, thereby achieving comprehensive alignment of the PET modality in terms of structure, texture, and metabolic patterns, ensuring the authenticity and stability of cross-modal generated images.
[0030] The overall inventive concept of this embodiment is to address the problems of spectral misalignment, structural distortion, and texture artifacts in traditional spatial domain generation for cross-modal tasks from MRI to PET by introducing wavelet multi-scale frequency domain modeling into the diffusion model. The core idea is to rely on wavelet pyramids to complete multi-scale frequency decomposition of images, and combine this with self-attention structures, adaptive wavelet modulation mechanisms, and inverse wavelet reconstruction to achieve alignment between low-frequency structures in MRI and smooth metabolic patterns in PET, as well as selective suppression and recovery of high-frequency texture and noise components. Figure 1 and Figure 2As shown, firstly, wavelet multi-scale decomposition is performed on MRI and PET images. The original images are decomposed into a low-frequency (LL) backbone subband and three directional high-frequency subbands (LH / HL / HH) at each scale. The low-frequency subband preserves the global intensity structure, while the high-frequency subband contains edge textures in different directions, providing explicit frequency channels for subsequent cross-modal spectral alignment. Next, a wavelet Transformer encoding network is constructed. The encoder uses the low-frequency subband as the backbone input. Before entering the Transformer block, each layer performs a wavelet transform on the low-frequency features of the current scale, decomposing them into the low-frequency subband of the next scale and the corresponding high-frequency subband. The low-frequency backbone features are input into the Transformer block for modeling, while the high-frequency subbands at each scale are cached for skip connections in the subsequent decoding stage. Through hierarchical self-attention modeling, the network can capture long-range correlations between different scales, achieving stepwise alignment of MRI and PET in the frequency domain. Finally, a wavelet Transformer decoding network is constructed. In the decoding stage, inverse wavelet transform is performed on the upsampled low-frequency backbone features and the corresponding scale-preserved high-frequency subbands to recover new low-frequency backbone features. These features are then fed into the Transformer for further modeling, thereby recovering the metabolic structure distribution of the PET modality layer by layer. Simultaneously, a high-frequency refinement module is introduced in the bottleneck region to selectively suppress and enhance high-frequency subbands in each direction, effectively removing structural textures in MRI that are irrelevant to PET and improving modal consistency. Furthermore, an adaptive wavelet modulation mechanism (AdaWM) is embedded in the Transformer backbone. This mechanism maps the diffusion time step encoding to learnable modulation parameters. By affinely modulating the wavelet features within the Transformer, low-frequency features remain stable in the early stages of diffusion, while necessary high-frequency details are gradually enhanced in the later stages, ensuring consistency in the evolution of the generation process in both the time and frequency domains. Finally, PET images are generated through inverse wavelet reconstruction, and the entire model is jointly optimized using reconstruction loss, content loss, adversarial loss, and high-frequency sparsity regularization loss. This ensures that the generated images satisfy the characteristics of PET imaging in terms of structure, metabolic patterns, and texture, guaranteeing the realism and reliability of cross-modal image generation. In summary, this embodiment constructs a complete frequency domain diffusion generation framework for MRI to PET by synergistically optimizing multiple modules such as wavelet multi-scale expansion, self-attention feature modeling, time step modulation and frequency domain reconstruction, and realizes PET image generation with structural fidelity, metabolic consistency and controllable details.
[0031] Step 1: Acquire raw medical images, which include MRI and PET imaging data.
[0032] Step 2: Based on the frequency domain modeling of wavelet pyramids, wavelet transform is performed on MRI and PET images, and a wavelet multi-scale pyramid representation is constructed to provide explicit frequency decomposition of low-frequency backbone and high-frequency details for subsequent wavelet Transformer coding and diffusion bridge generation.
[0033] Step 2.1: Perform wavelet transform on MRI and PET images.
[0034] Step 2.1.1: Perform wavelet transform on the MRI images to obtain frequency domain subbands.
[0035] In one specific implementation, for MRI images, the preprocessed two-dimensional grayscale image The signal is fed into the wavelet transform module, where a two-dimensional discrete wavelet transform (DWT) is performed to obtain a low-frequency subband and three directional high-frequency subbands, namely: ; in This represents a low-frequency approximate subband of an MRI image, preserving the overall intensity distribution of the target structure; These represent high-frequency detail subbands in the horizontal, vertical, and diagonal directions, respectively, to depict tissue edges, textures, and local structural changes.
[0036] In one specific implementation, the two-dimensional wavelet transform can be expressed as: ; ; in For input images, Each is composed of a one-dimensional low-pass filter With high-pass filter It is obtained by combining in the row and column directions. This indicates a double downsampling.
[0037] Step 2.1.2: Perform wavelet transform on the PET image to obtain the frequency domain subband. In one specific implementation, for PET images, the preprocessed two-dimensional grayscale image The signal is fed into the wavelet transform module, where a two-dimensional discrete wavelet transform (DWT) is performed to obtain a low-frequency subband and three directional high-frequency subbands, namely: ; in This represents a low-frequency approximate subband of an MRI image, preserving the overall intensity distribution of the target structure; These represent high-frequency detail subbands in the horizontal, vertical, and diagonal directions, respectively, to depict tissue edges, textures, and local structural changes.
[0038] Step 2.1.3: Splice and standardize the sub-bands to form frequency domain input features.
[0039] In one specific implementation, the low-frequency and high-frequency sub-bands obtained from MRI and PET decomposition are stitched together along the channel dimension to construct multi-channel features. For example, for MRI images, this can be denoted as: ; For PET images, corresponding results can be obtained. Subsequently, the multi-channel sub-band features are normalized and dimension-mapped using a linear projection convolution or affine transformation layer to adapt them to the input dimension requirements of the subsequent wavelet Transformer encoder and diffusion bridge model.
[0040] Step 2.2: After obtaining the MRI and PET wavelet subbands, to obtain richer multi-scale structural information, a recursive wavelet transform is further performed on the low-frequency subbands obtained in Step 2.1. Specifically, for the... Low-frequency backbone characteristics of the layer We can continue with wavelet transform to obtain: ; in The low-frequency backbone features, representing lower resolution and a higher level of abstraction, are input to the corresponding layer's Trasnformer module. It is then preserved at this scale and used in inverse wavelet reconstruction via skip connections during the decoding stage.
[0041] In one specific implementation, the multi-scale wavelet pyramid is used not only for frequency domain downsampling at the encoding end but also to provide inverse transform constraints for reconstruction at the decoding end. The basic form of the inverse wavelet transform (IWT) can be expressed as: ; Through wavelet transform and pyramid construction described in step 2, this embodiment provides decomposable and alignable low-frequency backbone and high-frequency detail representations for MRI and PET in the frequency domain, laying a unified frequency space foundation for subsequent wavelet Transformer multi-scale encoding, decoding and diffusion bridge generation.
[0042] Step 3: Construct a wavelet pyramid-based Transformer generator network. Based on the wavelet pyramid frequency domain representation obtained in Step 2, construct a U-shaped Transformer generator network based on the wavelet pyramid. Simultaneously model low-frequency backbone features and high-frequency features at multiple scales. The Transformer module integrates a time-domain-frequency domain adaptive modulation mechanism and introduces a high-frequency refinement module in the bottleneck layer to improve the structural fidelity and spectral consistency of cross-modal generation from MRI to PET.
[0043] Step 3.1: In one specific implementation, this step uses a U-shaped encoder-decoder structure to construct the overall framework of the generator network. The downsampling path compresses the spatial resolution layer by layer by performing wavelet transforms on the low-frequency backbone features, and extracts low-frequency features at the current scale using the Transformer module at each layer. Correspondingly, the upsampling path restores the spatial resolution layer by layer through inverse wavelet reconstruction, and reconstructs the high-frequency subbands buffered at each scale from the encoder and the upsampled low-frequency backbone features from the decoder, thereby achieving collaborative modeling of multi-scale structure and details.
[0044] In this U-shaped structure, the wavelet pyramid is responsible for multi-scale decomposition and reconstruction in the frequency domain, the Transformer module is responsible for modeling structural dependencies and cross-modal feature relationships at each scale, and the high-frequency subbands are passed between the encoder and decoder through skip connections, so that the model retains the necessary high-frequency detail information while maintaining the stability of the global structure.
[0045] Step 3.2: In one specific implementation, each layer of the downsampling coding path consists of a "low-frequency wavelet transform + Transformer block". Specifically, for the... Low-frequency backbone characteristics of the layer First, perform wavelet transform to obtain Among them, the low-frequency subband Input this layer of Transformer module for feature modeling, high-frequency subband The skip connection is passed to the corresponding decoding layer.
[0046] In the decoding path, for the first... The low-frequency backbone features from the layer decoding stage are first upsampled to restore spatial resolution, and then fed together with the corresponding high-frequency subbands cached at the encoding end into the inverse wavelet transform module to obtain new low-frequency backbone features, which are then input into the Transformer module of that layer. Through this symmetrical design of "encoding-side wavelet transform + decoding-side inverse wavelet transform", the model can explicitly utilize frequency domain information at each scale to complete structure restoration and cross-modal spectrum alignment.
[0047] Step 3.3: Internal Structure of the Transformer Module and Time-Frequency Adaptive Modulation. In one specific implementation, each layer of the wavelet Transformer module includes a window-based multi-head self-attention unit, a feedforward network, and a normalization and residual connection structure combined with wavelet features, used for contextual modeling of low-frequency backbone features at the current scale. To adapt to the temporal evolution during the diffusion generation process, this embodiment introduces an adaptive wavelet modulation mechanism (AdaWM) into the Transformer module. Specifically, the diffusion time step... The input to the temporal coding network yields a temporal embedding vector, which is then mapped through a fully connected layer to parameters for an affine transformation of wavelet domain features. Before normalization layer or feature update, the backbone features are processed. implement: ; This allows for dynamic adjustment of the frequency domain response at different time steps. In the early stages of diffusion, high frequencies are suppressed to enhance the low-frequency stability structure, while in the later stages, high-frequency capabilities are gradually released to recover edges and details, achieving integrated time-frequency modulation. This modulation mechanism can be used in both the encoding and decoding stages of the Transformer module to ensure the consistency of the spectral evolution throughout the generation process. Step 3.4: High-Frequency Refinement Module (HF refinement) at the Bottleneck Layer. In one specific implementation, a high-frequency refinement module is set at the bottleneck layer of the U-shaped structure to selectively optimize the high-frequency subbands after multi-scale aggregation. This module performs lightweight filtering and gating operations on the wavelet high-frequency subbands, suppressing noise-dominated or PET mode-independent high-frequency texture responses, and retaining effective high-frequency information related to metabolic boundaries and structural contours.
[0048] The high-frequency subbands, after high-frequency refinement, are redistributed to the decoding path to participate in the inverse wavelet reconstruction process. This ensures the overall smoothness of the PET modality while preserving necessary diagnostic details and reducing the residue of MRI-specific structural textures in the PET domain, thereby improving the modal consistency and visual naturalness of the cross-modal generation results.
[0049] Step 4: Wavelet frequency domain-based diffusion bridge generation process. A diffusion bridge generation process is constructed within the wavelet pyramid Transformer feature space obtained in Steps 2 and 3. During the forward diffusion process, the wavelet sub-band of the target modality PET is used as the starting point, and the wavelet sub-band of the source modality MRI after noise perturbation is used as the ending point. Through forward diffusion and backward generation processes, cross-modal mapping from MRI to PET is achieved in a finite number of steps.
[0050] Step 4.1: In one specific implementation, the PET wavelet subband representation is denoted as... This corresponds to the target modal wavelet features obtained in steps 2–3 (e.g. The wavelet subband representation obtained after splicing); the MRI wavelet subband representation is denoted as This corresponds to the wavelet feature of the source mode.
[0051] This embodiment employs the diffusion bridge theory, within a finite time step... The definition above is from arrive The forward stochastic process. In this invention, the diffusion bridge generation framework differs from the traditional conditional diffusion model. Standard conditional diffusion methods typically introduce conditional information only during the denoising process, lacking explicit constraints on the generation path of the target mode, and are prone to problems such as conditional information decay or structural shift during multi-step stochastic denoising. In contrast, the diffusion bridge models the generation path at the stochastic process level, simultaneously constraining the start and end distributions of the generation process under the given source mode conditions, so that the generation process of the target mode maintains statistical consistency with the structure of the source mode.
[0052] Specifically, this invention constructs a diffusion bridge model in a multi-scale wavelet feature space. The forward diffusion process describes the gradual degradation of PET wavelet features to MRI wavelet features, while the reverse generation process recursively predicts PET wavelet features under the conditional constraints of MRI wavelet features. By introducing a recursive prediction mechanism, the target features are updated multiple times in the same time step, effectively suppressing the accumulation of errors during the diffusion process and improving the stability and consistency of cross-modal generation results in suppressing low-frequency metabolic structures and high-frequency textures. Combining the diffusion bridge framework with the wavelet pyramid frequency domain representation allows the generation process to explicitly model the cross-modal differences between MRI and PET at different scales and frequency bands. Compared to conditional diffusion directly in the pixel domain or a single latent space, this is more conducive to maintaining a balance between structural constraints and modal specificity in medical images. The specific process is as follows: Given a pair of wavelet features , No. intermediate state at each time step It is generated by the following formula: ; in Standard Gaussian noise, and Let be the weighting coefficients with respect to time, satisfying =1, used to mix the target mode and the source mode. The time-varying noise variance is used to smooth data distribution and improve robustness to measurement noise.
[0053] In implementation, the forward diffusion process is determined by a preset noise scheduling function and time step sampling rules. During the training phase, this process is used to progressively add noise to the real PET wavelet features to obtain the results at each time step. , which serves as the conditional input for the discriminator and the generator.
[0054] Step 4.2: Backdiffusion bridge generation process based on recursive estimation. In one specific implementation, the backdiffusion bridge generation process starts from the endpoint sample. Start, generate step by step This embodiment employs a generative network with a recursive mechanism. With discriminative networks Forming a recovery subnetwork, the direction transition probability is... Perform parameterization.
[0055] In the reverse direction Step 1, in order to obtain the target wavelet feature estimate corresponding to the current time step. This embodiment uses a recursive estimation method: initialization Then in a fixed Under MRI conditions, multiple forward calls to the generative network: ; in It is the number of recursions, when and When the change between the two values is less than a set threshold, convergence is considered achieved, and a consistent solution is taken. This serves as the target estimation mode for that time step. Compared to a one-step estimation that only performs sequential forward reasoning, this recursive mechanism allows the generator to correct errors multiple times within the same time step, improving accuracy. This improves the approximate accuracy, thereby enhancing the overall generation quality. In obtaining Subsequently, this embodiment constructs the posterior distribution based on the analytical form of the forward transition probability. The Gaussian distribution whose posterior mean and variance can be analytically expressed is: ; in The posterior mean function is derived from the forward noise scheduling and convex combination weights. This represents the corresponding posterior variance. Specifically, the posterior mean also depends on the current state. Source mode and recursive estimation This enables the joint utilization of MRI conditional information and PET target estimation information. During the training phase, real samples Samples are generated from the forward process. The discriminant network is obtained from posterior sampling. It is used to distinguish between real and generated intermediate states and to provide the generator with adversarial gradient signals to constrain the consistency between the backdiffusion process and the real diffusion bridge path.
[0056] Step 5: Loss Function Calculation and Model Optimization. This embodiment constrains the generator network through multiple loss functions to ensure that the generated PET images are consistent with the real PET modality in terms of spectral structure, detail texture, and overall authenticity. The overall loss consists of reconstruction loss, content loss, adversarial loss, and high-frequency sparse regularization loss.
[0057] Step 5.1: In one specific implementation, to ensure that the model can accurately reconstruct the multi-scale wavelet subbands of PET, this invention introduces a reconstruction loss based on the L1 norm to predict wavelet features. Compared with true wavelet features By constraining the differences between them, each frequency band generated has the spectral structure characteristics of a true PET.
[0058] Step 5.2: In one specific implementation, to ensure the quality of the PET image after inverse wavelet reconstruction... In terms of structure, brightness, and contrast, it is similar to a real PET image. To maintain consistency, this invention introduces a content loss based on the structural similarity index SSIM, through: ; Promote spatial domain structure consistency so that the generated images are more consistent with the physiological characteristics of PET in terms of macroscopic metabolic distribution.
[0059] Step 5.3: In one specific implementation, to improve the authenticity of the reverse-generated path of the diffusion bridge, this invention employs a discriminator to distinguish between true and false intermediate states, forming an adversarial learning structure of generator-discriminator. This is achieved through adversarial loss constraints. The back-diffusion process ensures that the generated image maintains the same distribution as the real PET wavelet features at each time step.
[0060] Step 5.4: In one specific embodiment, to suppress redundant high-frequency structural textures that may be introduced by MRI during PET generation, this invention optimizes the high-frequency subbands. Introducing sparse regularization loss This makes the high-frequency distribution of PET results more consistent with smooth metabolic characteristics and reduces the "artifacts" of MRI-specific structures in PET.
[0061] Step 5.5: In one specific implementation, the present invention combines the various losses in a weighted manner to obtain the overall generator loss: ; in, For the total loss, , These are the hyperparameters that balance the four objective functions.
[0062] By using this overall loss to train the generative network end-to-end, the model can achieve consistent optimization results in both the frequency and spatial domains, ultimately realizing high-fidelity MRI-to-PET mapping.
[0063] This invention proposes a method for generating PET images from MRI images based on multi-scale spectral alignment. By constructing a unified modeling framework between the frequency and time domains, high-fidelity generation is achieved in cross-modal imaging tasks. This invention fully utilizes the advantages of wavelet pyramids in frequency domain decomposition, explicitly separating and modeling the low-frequency structure and high-frequency texture of MRI and PET images. This allows the differences in frequency distribution between the two modalities to be directly aligned and compensated within the network. Furthermore, a Transformer module with wavelet features as input is introduced, enabling the model to simultaneously establish interactive expressions of global dependencies and local frequency features at various scales, thereby obtaining structure and texture reconstructions that better conform to the physiological metabolic characteristics of PET.
[0064] This invention integrates the diffusion bridge method into the wavelet Transformer framework, enabling the generation process to establish continuous intermediate states between MRI conditional information and PET target distribution. This allows the cross-modal transformation to gradually align structural and textural details in a multi-step progression. By introducing a recursive prediction mechanism, the generator can continuously correct its estimation of PET wavelet subbands at each time step, making the final inverse diffusion process more stable and consistent. Furthermore, this invention incorporates a time-step-driven adaptive wavelet modulation mechanism within the Transformer module, enabling the model to automatically adjust the intensity of low-frequency structure recovery and high-frequency detail compensation according to the diffusion stage, achieving joint time-domain and frequency-domain control and improving the smoothness and consistency of the generation path.
[0065] This invention incorporates a high-frequency refinement module at the bottleneck layer of the wavelet pyramid, specifically designed to suppress redundant high-frequency details or anatomical textures that may be introduced from the MRI source modality. This reduces artifacts or erroneous structures in PET images, improving modal consistency across modal generation. To further constrain the generated results, this invention designs a multi-objective optimization function including reconstruction loss, content structure loss, adversarial loss, and high-frequency sparsity regularization loss. This allows the model to be comprehensively optimized in terms of distribution consistency in both the frequency and spatial domains, ultimately generating more realistic, stable, and medically interpretable PET images.
[0066] Example 2 This embodiment provides a system for generating PET images from MRI images based on multi-scale spectral alignment, including: The data acquisition and preprocessing module is configured to acquire MRI images and their corresponding PET images, and to preprocess the acquired images. The wavelet transform module is configured to perform multi-scale wavelet transforms on MRI and PET images respectively to obtain low-frequency sub-bands at each scale and high-frequency sub-bands in multiple directions. The image reconstruction module is configured to: construct a wavelet diffusion bridge generation model, which includes an encoder, a decoder, and a diffusion bridge module; the encoder, based on a Transformer structure, is used to perform multi-scale feature modeling on the low-frequency subbands of the MRI image and to pass the high-frequency subbands of each scale to the decoder through skip connections; the decoder is used to obtain the low-frequency features output by the encoder and the high-frequency subbands of the same scale passed by the skip connections, and directly reconstruct the predicted PET wavelet subbands to generate a PET image through inverse wavelet transform; the diffusion bridge module is used to transition the PET wavelet features to MRI wavelet features in the multi-scale wavelet feature space through a forward diffusion process, and to recursively predict the PET wavelet features based on the MRI wavelet features as conditions through a reverse generation process; The image output module is configured to: train the model using a comprehensive loss function, generate a model using the trained wavelet diffusion bridge, sample the input MRI image, and generate the corresponding PET image.
[0067] It should be noted that the above modules correspond to the steps in Embodiment 1, and the examples and application scenarios implemented by the above modules and their corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above modules can be executed in a computer system as part of the system.
[0068] In further embodiments, the following is also provided: An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method described in Embodiment 1. For brevity, further details are omitted here.
[0069] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0070] A computer-readable storage medium for storing computer instructions that, when executed by a processor, perform the method of Embodiment 1.
[0071] The method in Example 1 can be directly executed by a hardware processor, or it can be executed by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.
[0072] A computer program product includes a computer program that, when executed by a processor, implements the method in Embodiment 1.
[0073] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which execute in a device on a target real or virtual processor to perform the processes / methods described above. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided among program modules as needed. The machine-executable instructions for the program modules can execute within a local or distributed device. In a distributed device, the program modules can reside in both local and remote storage media.
[0074] The computer program code used to implement the methods of the present invention may be written in one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the computer or other programmable data processing device, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a computer, partially on a computer, as a stand-alone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.
[0075] In the context of this invention, computer program code or related data may be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals may include electrical, optical, radio, sound, or other forms of propagation signals, such as carrier waves, infrared signals, etc.
[0076] Those skilled in the art will recognize that the units and algorithm steps described in conjunction with the embodiments herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0077] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for generating PET images from MRI images based on multi-scale spectral alignment, characterized in that, Includes the following steps: Acquire MRI images and their corresponding PET images, and preprocess the acquired images; Multi-scale wavelet transform was performed on MRI and PET images to obtain low-frequency sub-bands at each scale and high-frequency sub-bands in multiple directions. A wavelet diffusion bridge generation model is constructed, comprising an encoder, a decoder, and a diffusion bridge module. The encoder, based on a Transformer structure, is used to model the low-frequency subbands of MRI images at multiple scales and to pass the high-frequency subbands at each scale to the decoder via skip connections. The decoder is used to acquire the low-frequency features output by the encoder and the high-frequency subbands at the same scale passed by the skip connections, and directly reconstructs the predicted PET wavelet subbands to generate PET images through inverse wavelet transform. The diffusion bridge module is used to transition PET wavelet features to MRI wavelet features in the multi-scale wavelet feature space through a forward diffusion process, and to recursively predict PET wavelet features based on the MRI wavelet features as conditions through a reverse generation process. The model is trained using a comprehensive loss function. The trained wavelet diffusion bridge model is then used to sample the input MRI image and generate the corresponding PET image.
2. The method for generating PET images from MRI images based on multi-scale spectral alignment as described in claim 1, characterized in that, Multi-scale wavelet transforms were performed on both MRI and PET images, specifically as follows: Two-dimensional discrete wavelet transform is performed on MRI and PET images respectively to obtain low-frequency sub-bands and high-frequency sub-bands in three directions. The low-frequency sub-bands and high-frequency sub-bands in the three directions of each image are then spliced together in the channel dimension to form a multi-channel frequency domain feature representation.
3. The method for generating PET images from MRI images based on multi-scale spectral alignment as described in claim 1, characterized in that, The wavelet diffusion bridge generation model adopts a U-shaped structure. In the encoding path, wavelet transforms are performed on the low-frequency sub-bands layer by layer. The newly generated low-frequency sub-bands are input into the Transformer module of the current layer for feature modeling, and the corresponding high-frequency sub-bands are buffered. In the decoding path, inverse wavelet transforms are performed on the upsampled low-frequency backbone features and the buffered high-frequency sub-bands to recover the new low-frequency backbone features, and then input into the Transformer module of the corresponding layer for feature modeling.
4. The method for generating PET images from MRI images based on multi-scale spectral alignment as described in claim 3, characterized in that, The Transformer module integrates a time-domain and frequency-domain adaptive modulation mechanism. This mechanism encodes the time step of the diffusion process into modulation parameters and performs affine transformations on the features at the current scale to achieve dynamic adjustment of low-frequency structural stability and high-frequency detail recovery.
5. The method for generating PET images from MRI images based on multi-scale spectral alignment as described in claim 1, characterized in that, The model also includes a high-frequency refinement module, located in the bottleneck layer of the encoder and decoder, used to selectively suppress or enhance high-frequency subbands to reduce high-frequency textures in the MRI source modality that do not match the PET modality.
6. The method for generating PET images from MRI images based on multi-scale spectral alignment as described in claim 1, characterized in that, In the diffusion bridge module, the reverse generation process adopts a recursive prediction mechanism, which iteratively updates the estimation of PET wavelet features multiple times within the same time step.
7. A system for generating PET images from MRI images based on multi-scale spectral alignment, characterized in that, include: The data acquisition and preprocessing module is configured to acquire MRI images and their corresponding PET images, and to preprocess the acquired images. The wavelet transform module is configured to perform multi-scale wavelet transforms on MRI and PET images respectively to obtain low-frequency sub-bands at each scale and high-frequency sub-bands in multiple directions. The image reconstruction module is configured to: construct a wavelet diffusion bridge generation model, which includes an encoder, a decoder, and a diffusion bridge module; the encoder, based on a Transformer structure, is used to perform multi-scale feature modeling on the low-frequency subbands of the MRI image and to pass the high-frequency subbands of each scale to the decoder through skip connections; the decoder is used to obtain the low-frequency features output by the encoder and the high-frequency subbands of the same scale passed by the skip connections, and directly reconstruct the predicted PET wavelet subbands to generate a PET image through inverse wavelet transform; the diffusion bridge module is used to transition the PET wavelet features to MRI wavelet features in the multi-scale wavelet feature space through a forward diffusion process, and to recursively predict the PET wavelet features based on the MRI wavelet features as conditions through a reverse generation process; The image output module is configured to: train the model using a comprehensive loss function, generate a model using the trained wavelet diffusion bridge, sample the input MRI image, and generate the corresponding PET image.
8. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, perform the method described in any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the method described in any one of claims 1-6.