FFA image generation system, method, medium and device based on dual-domain constraint mamba diffusion model
By employing a method based on a dual-domain constrained Mamba diffusion model, combined with learnable wavelet frequency domain and dual-pyramid spatial domain feature extraction, the problems of insufficient vascular detail restoration, stability, and semantic expression in cross-modal generation from CFP to FFA are solved, enabling the generation of high-quality FFA images and supporting low-cost screening for retinal diseases.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI FIRST PEOPLES HOSPITAL
- Filing Date
- 2025-08-12
- Publication Date
- 2026-05-08
AI Technical Summary
Existing cross-modal generation methods from CFP to FFA suffer from insufficient ability to restore vascular details, poor stability and diversity of generated results, and insufficient semantic expression capabilities, making it difficult to meet the needs of clinical diagnosis.
A high-quality FFA image is generated by employing a dual-domain constrained Mamba diffusion model, combined with a learnable wavelet frequency domain extractor and a dual-pyramid spatial domain feature extractor, and denoising is performed through a dual-domain constrained Mamba module.
It improves the ability to model detailed vascular structures, enhances the stability of the generation process and the accuracy of semantic representation, and can generate high-quality FFA images to support low-cost screening of retinal diseases.
Smart Images

Figure CN121095098B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical image processing technology, and in particular to an FFA image generation system, method, medium and apparatus based on a dual-domain constrained Mamba diffusion model. Background Technology
[0002] In ophthalmological clinical diagnosis and treatment, retinal imaging technology plays a crucial and irreplaceable role, especially in the early screening, disease assessment, and monitoring of treatment effectiveness for retinal-related diseases. Currently, color fundus photography (CFP) and fundus fluorescein angiography (FFA) are the two most common and widely used retinal imaging techniques. They are significantly complementary in their imaging principles and clinical functions, together forming an important foundation for current retinal imaging diagnosis.
[0003] CFP (Catalytic Angiography) is an imaging technique that acquires color images of the fundus under visible light illumination. It requires no contrast agent injection and boasts advantages such as being non-invasive, rapid, economical, and easy to operate. It can clearly present the structural features of the fundus, such as the optic disc, retinal vascular distribution, and macular morphology. CFP is mainly used for preliminary screening of fundus diseases and observation of structural lesions, and has broad application prospects in clinical scenarios such as hypertensive retinopathy, myopia, and optic nerve diseases. However, CFP only provides static anatomical information and cannot reflect vascular permeability, microcirculatory function, and hemodynamic changes. Its diagnostic and assessment capabilities are limited in complex fundus diseases involving dynamic vascular changes or functional abnormalities (such as diabetic retinopathy and age-related macular degeneration). In contrast, FFA (Fluorescence Angiography) is a functional imaging technique based on the principle of angiography. It uses intravenous injection of sodium fluorescein as a contrast agent and acquires fundus fluorescence signals under excitation light of a specific wavelength, thereby recording the distribution and flow of fluorescein in the retinal vascular system in real time. FFA can directly reflect changes in vascular permeability and accurately identify pathological features such as microaneurysms, neovascularization, capillary non-perfusion areas, and vascular leakage. It has irreplaceable clinical value in the diagnosis, staging, prognostic assessment, and treatment efficacy evaluation of retinal vascular diseases (especially diabetic retinopathy, venous occlusive disease, and choroidal neovascularization).
[0004] However, as an invasive imaging technique, FFA relies on the injection of exogenous contrast agents, which may cause adverse reactions including nausea, vomiting, and allergic reactions. In severe cases, it can even lead to anaphylactic shock and respiratory failure, posing certain clinical risks. Therefore, how to obtain vascular function information equivalent to FFA in a non-invasive and low-risk manner without using sodium fluorescein has become one of the current hot research topics.
[0005] In recent years, with the continuous development of deep learning technology, cross-modal medical image generation methods have shown broad application prospects in fields such as assisted diagnosis and image enhancement, especially in achieving non-invasive acquisition of FFA images, providing a completely new technical path. These methods construct cross-modal generation models, using existing image data such as CFP as input, to generate images that are highly consistent with the target modality (e.g., FFA) in terms of structural and functional features, thereby achieving retinal vascular imaging without the need for contrast agent injection. This non-invasive imaging method not only effectively reduces examination risks and costs but also provides important support for large-scale screening and telemedicine, becoming an important direction in current ophthalmic artificial intelligence research. In existing research, scholars have proposed various deep learning-based methods for generating CFP-to-FFA images, which can be mainly divided into the following three categories: The first category of methods designs a dedicated image feature encoder-decoder network to extract and reconstruct features from CFP images, attempting to establish an explicit mapping relationship between CFP and FFA modalities, thereby generating output results that are structurally close to real FFA images; The second category of methods is based on the Generative Adversarial Network (GAN) framework, which learns the complex nonlinear mapping relationship from CFP to FFA through game optimization between the generator and the discriminator, in order to improve the realism and distribution consistency of the generated images; The third category of methods introduces a hybrid structure combining CNN and Transformer, which enhances the model's ability to model vascular details and potential lesion areas in CFP images by combining local feature modeling and global dependency modeling, while also integrating auxiliary information (such as vascular structure maps, lesion masks, or other guiding information), thereby improving the structural integrity and semantic expressiveness of the generated images.
[0006] Although existing methods have achieved initial success in generating FFA images from CFP images, they still face a series of technical bottlenecks and challenges, including the following: (1) Insufficient ability to restore vascular details: Most generative models have limited performance in microvascular modeling and are difficult to effectively reproduce the rich and detailed capillary network structure in real FFA images, resulting in the lack of key pathological details required in clinical practice in the generated images; (2) Poor stability and diversity of results: Generative models represented by GAN are easily affected by unstable convergence and pattern collapse during training, resulting in fluctuations in the diversity and reliability of the generated results, which limits their clinical usability; (3) Insufficient semantic expression ability: Some models are not accurate enough in expressing lesion areas (such as leakage points, microaneurysms, neovascularization, etc.) in the generated images, and have low semantic discrimination, which makes it difficult to meet the doctors' needs for interpreting fine pathological changes and affects their practical value in auxiliary diagnosis.
[0007] In summary, overcoming the aforementioned technical bottlenecks and improving the model's capabilities in areas such as detailed modeling of vascular structures, stability of the generation process, and accuracy of semantic expression have become key issues and technological development directions in current cross-modal generation research from CFP to FFA. There is an urgent need to propose a more efficient, robust, and medically interpretable image generation method to promote further development in this field. Summary of the Invention
[0008] In view of the shortcomings of the prior art described above, the purpose of this application is to provide an FFA image generation system, method, medium and device based on a dual-domain constrained Mamba diffusion model, which is used to solve the technical problems of how to improve the model's modeling of vascular structure details, generation process stability and semantic expression accuracy.
[0009] To achieve the above and other related objectives, a first aspect of this application provides an FFA image generation system based on a dual-domain constrained Mamba diffusion model, comprising: a VAE encoder, a dual-pyramid spatial domain feature extractor, and a learnable wavelet frequency domain extractor located in the forward diffusion stage; a CFP image is input to the pre-trained VAE encoder to be encoded into a latent representation, and Gaussian noise is progressively injected; the noisy latent representation is input to the dual-pyramid spatial domain feature extractor and the learnable wavelet frequency domain extractor respectively to extract spatial domain features for characterizing the spatial information of the lesion region and frequency domain features for characterizing the frequency domain contour of the vascular structure; a dual-domain conditionally constrained Mamba module and a VAE decoder located in the denoising stage; the input end of the dual-domain conditionally constrained Mamba module includes an input sequence composed of initial noise, spatial domain features, and frequency domain features; the input sequence is input to the dual-conditionally guided Mamba module for denoising guided by joint modeling in the spatial and frequency domains; the denoised latent representation is decoded and reconstructed by the VAE decoder to generate an FFA image.
[0010] In some embodiments of the first aspect of this application, the learnable wavelet frequency domain extractor combines learnable discrete wavelet transform with a subband energy balance mechanism and a Mamba sequence modeling module.
[0011] In some embodiments of the first aspect of this application, the learnable wavelet frequency domain extractor uses a set of learnable Haar wavelet filters as the initialization scheme for the learnable discrete wavelet transform, corresponding to low-pass and high-pass filtering operations respectively, and parameterizes them as backpropagable variables so that they can be adaptively adjusted during training to match the frequency domain distribution of the input image.
[0012] In some embodiments of the first aspect of this application, the Haar wavelet filter obtains low-frequency and high-frequency subbands from the input CFP image through convolution operations in four directions; and performs multi-level learnable wavelet decomposition on the low-frequency components to ultimately form multiple secondary subbands.
[0013] In some embodiments of the first aspect of this application, the learnable wavelet frequency domain extractor calculates the first-order spatial gradient difference between the reconstructed high-frequency subband and the original high-frequency subband in the horizontal and vertical directions based on the gradient preservation mechanism to construct the gradient preservation loss; and, through a learnable inverse wavelet transform module, it reconstructs all frequency band information layer by layer to restore it to the same resolution as the original image, so as to perform a reversible transformation from the frequency domain to the spatial domain.
[0014] In some embodiments of the first aspect of this application, the dual-pyramid spatial domain feature extractor includes a convolutional pyramid branch and a pooling pyramid branch; in the convolutional pyramid branch, multi-scale convolutional kernels are used to capture semantic features under different receptive fields, followed by channel compression, and the feature map is divided into patch sequences, which are then fed into the Mamba sequence modeling module for sequence modeling after introducing positional encoding; in the pooling pyramid branch, multi-scale adaptive average pooling is used to aggregate contextual information at multiple scales, and channel compression is combined with the Mamba sequence modeling module to integrate global semantic information; the local modeling of the convolutional pyramid branch and the global context-aware modeling of the pooling pyramid branch are fused to obtain the conditional priors for dual-pyramid spatial domain extraction.
[0015] In some embodiments of the first aspect of this application, the Mamba sequence modeling module employs a row-column scanning mechanism for image scanning.
[0016] In some embodiments of the first aspect of this application, the input to the dual-domain conditionally constrained Mamba module includes: modularizing the input noisy image into X. inpttokenThe patch sequence is obtained by fusing and embedding the patch with two types of conditional features respectively; wherein, the first type of conditional feature is the frequency domain feature extracted by the learnable wavelet frequency domain extractor; and the second type of conditional feature is the spatial domain feature extracted by the dual pyramid spatial domain feature extractor.
[0017] In some embodiments of the first aspect of this application, the patch sequence is fed into a biconditional attention module to model contextual relationships in the spatial and frequency domains and independently construct queries, keys, and values; the contextual features are input into a Mamba sequence modeling module, and through bidirectional state scanning and continuous activation operations, the operation results are trained by residual connection and normalization, and then the image structure is restored by Unpatchify operation to obtain the potential representation used to finally generate the FFA image.
[0018] In some embodiments of the first aspect of this application, the dual-pyramid spatial domain extractor is trained using an InfoNCE loss function; the learnable wavelet frequency domain extractor is trained based on a total loss function consisting of frequency domain contrast loss, subband energy conservation loss, and subband gradient total loss; during the training of the dual-domain conditional constraint Mamba module, the pre-training parameters of the dual-pyramid spatial domain extractor and the learnable wavelet frequency domain extractor are loaded and their structures are frozen, and their loss functions consist of spatial domain InfoNCE loss, frequency domain InfoNCE loss, and the total loss of the learnable wavelet frequency domain extractor.
[0019] To achieve the above and other related objectives, a second aspect of this application provides an FFA image generation method based on a dual-domain constrained Mamba diffusion model, comprising: in a forward diffusion stage, encoding a CFP image into a latent representation and progressively injecting Gaussian noise; extracting features from the injected noise latent representation to extract spatial domain features for characterizing the spatial information of the lesion region and frequency domain features for characterizing the frequency domain contour of the vascular structure; in a denoising stage, inputting an input sequence composed of initial noise, spatial domain features, and frequency domain features; denoising the input sequence under the guidance of joint modeling in the spatial and frequency domains; and decoding and reconstructing the denoised latent representation using the VAE decoder to generate an FFA image.
[0020] To achieve the above and other related objectives, a third aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the FFA image generation method based on the dual-domain constrained Mamba diffusion model.
[0021] To achieve the above and other related objectives, a fourth aspect of this application provides a computer program product comprising computer program code that, when executed on a computer, causes the computer to implement the FFA image generation method based on the dual-domain constrained Mamba diffusion model.
[0022] To achieve the above and other related objectives, a fifth aspect of this application provides a computer apparatus, including a memory, a processor, and a computer program stored in the memory; the processor executes the computer program to implement the FFA image generation method based on the dual-domain constrained Mamba diffusion model.
[0023] As described above, the FFA image generation system, method, medium, and apparatus based on the dual-domain constrained Mamba diffusion model of this application have the following beneficial effects: This invention proposes a spatial-frequency domain dual-constraint Mamba diffusion model (SFDC-MambaDiff), which integrates dual spatial and frequency domain constraints for cross-modal generation of FFA images from CFP images. Through a learnable wavelet feature extractor, key vascular frequency components are effectively extracted from fundus CFP images, enhancing the structured guidance for vascular edge representation in the generated FFA images. In the spatial domain, a dual-pyramid feature extractor effectively fuses local details and global semantics of fundus lesion regions, improving the consistency of lesion representation in the generated FFA images. Furthermore, the designed dual-domain conditional Mamba module, as the denoising core, achieves effective coordination of spatial and frequency information to improve image clarity and accuracy. Experimental validation on two public datasets, Hajeb and MPOS, and one private clinical dataset achieved PSNRs of 30.11 and 29.40, respectively. This demonstrates that the present invention performs excellently in terms of vascular detail preservation, structure generation, lesion restoration, and semantic consistency in cross-modal FFA image generation. It can make full use of fundus CFP image information and effectively generate high-quality FFA fundus images across modalities, providing new possibilities for low-cost screening of retinal diseases. Attached Figure Description
[0024] Figure 1 The diagram shown is a schematic representation of the overall architecture of an FFA image generation system based on a dual-domain constrained Mamba diffusion model, according to an embodiment of this application.
[0025] Figure 2 The diagram shown is a schematic representation of a learnable wavelet frequency domain extractor in one embodiment of this application.
[0026] Figure 3 This is a visual comparison diagram of different wavelet decomposition depths in one embodiment of this application.
[0027] Figure 4 The diagram shows the impact of different wavelet decomposition depths (Level 1, Level 2, Level 3) on model performance in one embodiment of this application.
[0028] Figure 5 The diagram shown is a structural schematic of a dual-pyramid spatial domain feature extractor in one embodiment of this application.
[0029] Figure 6 The diagram shown is a structural schematic of a dual-domain conditional constraint Mamba module in one embodiment of this application.
[0030] Figure 7 The diagram shows the performance test results of various modules on the Hajeb dataset in one embodiment of this application.
[0031] Figure 8 The diagram shows the results of one-way / multi-way variance statistical analysis of LWFE, DSFE, and DCCM in one embodiment of this application.
[0032] Figure 9 The diagram shown is a visualization analysis of the generated fundus FFA image in one embodiment of this application.
[0033] Figure 10 The diagram shown is a flowchart illustrating an FFA image generation method based on a dual-domain constrained Mamba diffusion model according to an embodiment of this application.
[0034] Figure 11 The diagram shown is a structural schematic of a computer device according to an embodiment of this application. Detailed Implementation
[0035] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.
[0036] Before providing a further detailed description of the present invention, the nouns and terms used in the embodiments of the present invention are explained, and the nouns and terms used in the embodiments of the present invention are subject to the following interpretations:
[0037] <1> The Mamba Diffusion Model is a generative artificial intelligence model based on the principle of diffusion, which generates high-quality images or data by progressively removing noise. It excels in multimodal generation tasks, capable of combining various inputs such as text and images to generate highly consistent and creative outputs.
[0038] <2> Fundus fluorescein angiography (FFA) is a medical examination method used to diagnose retinal vascular diseases. It involves intravenously injecting sodium fluorescein into the patient, then using light of a specific wavelength to excite the fluorescein and observe the filling and leakage of retinal blood vessels, thus aiding in the diagnosis of retinal vascular diseases, diabetic retinopathy, and other conditions.
[0039] <3> Color fundus photography (CFP) is an examination technique used to record color images of structures such as the retina and choroid at the back of the eye. It uses a special camera and optical system to capture images of the fundus, clearly showing details of retinal vessels, the optic disc, the macula, and other areas. It is widely used for screening and diagnosing eye diseases.
[0040] <4> The Dual-domain Conditional Constraint Mamba (DCCM) module integrates the frequency and spatial priors of CFP images. It achieves collaborative modeling of dual-domain features through a state-space modeling mechanism, thereby improving the model's ability to represent vascular structures and lesion regions and its cross-modal alignment accuracy.
[0041] <5> Learnable Wavelet Frequency-domain Extractor (LWFE): A learnable wavelet transform module used to extract frequency information from signals or images. It combines the multi-scale analysis capabilities of wavelet transform with the trainability of deep learning, enabling it to adaptively extract features at different frequency levels, making it suitable for complex signal processing and image analysis tasks.
[0042] <6> Dual-Pyramid Spatial-domain Feature Extractor (DSFE): A spatial domain feature extraction method based on a dual-pyramid structure. It extracts global and local features of an image by constructing two pyramid structures of different scales, and integrates these multi-level features through a fusion mechanism to improve the performance of image recognition and classification tasks.
[0043] <7> Spatial Domain: The spatial domain refers to the domain in image processing where operations are performed using pixels as the basic unit. In the spatial domain, each pixel value of an image directly participates in processing, such as filtering and edge detection, and is mainly used to process local features and spatial relationships of an image.
[0044] <8> Frequency Domain: The frequency domain is the domain in which signals or images are transformed from the spatial domain to the frequency domain for analysis and processing. Through methods such as Fourier transform, frequency domain analysis can reveal the frequency components of signals or images. It is commonly used for tasks such as filtering and spectrum analysis, and is suitable for handling global features and periodic changes.
[0045] <9> VAE (Variational Autoencoder) encoding is the encoding process of a variational autoencoder. It encodes input data into distribution parameters of the latent space, achieving dimensionality reduction and generation of data by introducing a probability distribution (such as a Gaussian distribution) while preserving the main features of the data. It is widely used in data generation, dimensionality reduction, and unsupervised learning tasks.
[0046] <10> Scale function: In wavelet transform, it is used to smooth signals and is a complementary function to the wavelet function.
[0047] <11> Wavelet function: A multi-resolution analysis function used in signal processing that can capture local features of a signal.
[0048] <12> PNSR (Peak Signal-to-Noise Ratio): Peak signal-to-noise ratio is a metric for measuring image quality, representing the ratio of the maximum possible power of an image to the power of noise.
[0049] <13> SSIM (Structural Similarity Index): A structural similarity index is a metric that measures the similarity between two images, taking into account brightness, contrast, and structural information.
[0050] <14> VS (Visual Similarity): Vascular similarity refers to the degree of similarity between vascular structures in two images.
[0051] <15> KID (Kernel Inception Distance): Kernel inception distance is a metric for evaluating the performance of generative models. It assesses the similarity between generated samples and real samples based on the distance in the feature space.
[0052] <16> FID (Fréchet Inception Distance): Fréchet distance is a metric that measures the difference in distribution between generated and real images, and is evaluated based on distance in feature space.
[0053] <17> LPIPS (Learned Perceptual Image Patch Similarity): Learned Perceptual Image Patch Similarity is an image similarity measurement method based on human perception.
[0054] <18> Haar Wavelet Filter: A simple wavelet transform filter used for signal decomposition and reconstruction in image processing.
[0055] <19> Unpatchify operation: refers to the process of reassembling image patches into a complete image, which is often used in image processing and computer vision tasks.
[0056] <20> InfoNCE Loss: A loss function based on contrastive learning, used to train neural networks to learn feature representations that distinguish different categories.
[0057] <21> Hajeb and MPOS datasets: Two distinct datasets of fundus FFA images. The Hajeb dataset focuses on diabetes, while the MPOS dataset is a dataset of multiple fundus diseases.
[0058] To address the challenges in the aforementioned background technology, this application provides a Mamba diffusion model (SFDC-MambaDiff) based on dual spatial-frequency domain constraints for generating accurate FFAs from CFP images across modalities. Specifically, this invention first designs a Learnable Wavelet Frequency-domain Extractor (LWFE) to extract frequency domain features from CFP images, serving as frequency domain constraints for the diffusion model to generate FFAs, thereby improving the high-fidelity representation of key structures in the generated images and enhancing the clarity of vascular contours. Secondly, a Dual-Pyramid Spatial-domain Feature Extractor (DSFE) is constructed, which effectively captures global and local spatial information from CFP images through convolution and pooling operations, serving as spatial domain constraints to improve the semantic consistency of lesion regions in FFAs. Finally, a dual-domain Conditional Constraint Mamba module (DCCM) is designed as the denoising core of the diffusion model, which can efficiently fuse prior information from the spatial and frequency domains, promoting the generation of clear and accurate FFA images. Experiments on public and private datasets demonstrate that this invention can fully utilize fundus CFP image information to generate high-quality FFA images across modalities.
[0059] To facilitate understanding of the embodiments of this application, firstly, in conjunction with Figure 1 Detailed explanation. Figure 1 This paper presents a schematic diagram of the overall architecture of an FFA image generation system based on a dual-domain constrained Mamba diffusion model according to an embodiment of the present invention. The system uses a Mamba diffusion model (SFDC-MambaDiff) based on dual constraints in the spatial and frequency domains to generate accurate FFA images (Fundus Fluorescein Angiography) from CFP cross-modal (Color Fundus Photography).
[0060] In the embodiments of this application, the FFA image generation system based on the dual-domain constrained Mamba diffusion model consists of two stages: forward diffusion and backward denoising.
[0061] The forward diffusion stage (also known as Stage 1: Conditional Prior Extraction) involves the following steps: First, the input CFP image is encoded into a latent representation by a pre-trained VAE encoder, and Gaussian noise is progressively injected to make the latent representation approximate a standard normal distribution. Second, a spatial domain feature extraction module and a frequency domain feature extraction module based on different Mamba scanning methods are constructed to extract the spatial information of the lesion region and the frequency domain contour of the vascular structure, respectively.
[0062] The denoising stage (also known as stage 2: Mamba Diffusion Generations) involves dividing the initial noise into several non-overlapping image patches, flattening them, and mapping them to fixed-dimensional embedding vectors. These vectors, along with the spatial domain feature embeddings and frequency domain feature embeddings output by the spatial domain feature extraction module and the frequency domain feature extraction module, form the input sequence. Each image patch in the input sequence is embedded with a learnable positional code. The input sequence is then fed into a dual-condition-guided Mamba module for efficient denoising under the guidance of joint modeling in the spatial and frequency domains, gradually approximating the target FFA image distribution through the latent representation. The denoised latent representation is reconstructed by the VAE decoder to generate a high-quality FFA image.
[0063] The forward diffusion stage includes a learnable wavelet frequency domain extractor 101 and a dual-pyramid spatial domain feature extractor 102. The specific structure and principle of these two modules will be explained in detail below with reference to specific embodiments.
[0064] The structure of the learnable wavelet frequency-domain extractor 101 (LWFE) is as follows: Figure 2As shown, this invention combines learnable discrete wavelet transform with a subband energy balance mechanism and a Mamba sequence modeling module to achieve more expressive frequency domain feature modeling. It is worth noting that accurate modeling of vascular structures remains a challenging problem in FFA image generation. Existing methods often suffer from blurred vessel orientation, unclear contours, and missing structural details when processing microvessels and edge regions, especially in the presence of complex backgrounds or lesion interference, making it difficult to restore the integrity and continuity of the vascular network. This problem mainly stems from the model's insufficient ability to express high-frequency information, limiting its perception and generation of key edge and contour features in the image. To address this, this invention introduces a learnable wavelet frequency domain extractor to perform multi-scale frequency domain decomposition on the input image, thereby extracting more discriminative vascular edge and structural contour information. The learnable wavelet frequency domain extractor, as a frequency domain information extractor for vascular structures, is embedded in the image generation backbone, providing structural frequency domain constraints to assist in the generation of potential representations of vascular structures.
[0065] It is worth noting that the learnable discrete wavelet transform and subband energy balancing mechanism are methods that combine learnable discrete wavelet transform and subband energy balancing. Learnable discrete wavelet transform is an innovative method that combines deep learning with traditional wavelet transform, breaking through the limitation of fixed wavelet bases in traditional wavelet transform. It automatically learns the wavelet basis functions best suited for the input signal through neural network training. The core of this transform lies in its trainable filter bank, whose parameters can be optimized through backpropagation to adapt to specific task requirements. The subband energy balancing mechanism is an important component of learnable discrete wavelet transform. During wavelet transform, the signal is decomposed into multiple subbands, each representing the signal's characteristics within a different frequency range. The subband energy balancing mechanism focuses on how to rationally distribute energy among these subbands to achieve better signal processing results. By adjusting the energy distribution of each subband, the performance of the signal processing task can be optimized.
[0066] Preferably, this application employs a set of learnable Haar wavelet filters as the initialization scheme for the learnable discrete wavelet transform, corresponding to low-pass and high-pass filtering operations respectively, and parameterizing them as backpropagable variables so that they can adaptively adjust during training to better match the frequency domain distribution of the input image. It should be noted that the definition of Haar wavelet filters is simple: the low-pass filter is [1,1], and the high-pass filter is [1,-1]. This simple structure makes it easy to implement. Due to its simplicity, the computational complexity of the Haar wavelet transform is low, making it suitable for rapid implementation and processing of large-scale data.
[0067] Specifically, the scaling function and wavelet function of the Haar wavelet filter can be expressed in the following form:
[0068]
[0069] Where φ(t) represents the low-pass filter (corresponding to the scaling function), and ψ(t) represents the high-pass filter (corresponding to the wavelet function). It is an interval indicator function, which takes the value 1 when t∈[a,b), and 0 otherwise. The learnable weights of the low-pass filter are initialized to α1 = α2 = 0.5. The learnable weights of the high-pass filter are represented by β1 = -0.5 and β2 = 0.5.
[0070] For the input CFP image x∈R B×C×H×W The Haar wavelet filter obtains low-frequency and high-frequency subbands, namely LL1, LH1, HL1, and HH1, through convolution operations in four directions; and extracts global structure, vertical edges, horizontal edges, and fine-grained textures, respectively.
[0071] It should be understood that in wavelet decomposition, LL1 refers to the low-frequency subband obtained after wavelet decomposition, which retains the low-frequency information of the image, namely the global structure and general outline of the image. The global structure refers to the relatively smooth and slowly changing parts of the image (such as the overall brightness distribution and the outlines of large objects). The low-frequency components in LL1 can well reflect this global structural information because it filters out the high-frequency details in the image and retains the main structure of the image. LH1 refers to the low-frequency-high-frequency subband obtained after wavelet decomposition, representing the high-frequency information of the image in the horizontal direction (i.e., the rapidly changing parts of the image in the horizontal direction). Since vertical edges produce high-frequency changes in the horizontal direction, the LH1 subband can capture vertical edges well (regions in the image where brightness or color changes abruptly in the vertical direction). HL1 refers to the high-frequency-low-frequency subband obtained after wavelet decomposition, representing the high-frequency information of the image in the vertical direction (i.e., the rapidly changing parts of the image in the vertical direction). Since horizontal edges produce high-frequency changes in the vertical direction, the HL1 subband can capture these horizontal edges well. HH1 refers to the high-frequency subband obtained after wavelet decomposition, representing the part of the image that contains high-frequency information in both the horizontal and vertical directions, i.e., the region in the image that is rich in detail and changes dramatically. Since image textures can produce high-frequency variations in both the horizontal and vertical directions, the HH1 subband can capture these fine-grained textures very well.
[0072] Preferably, to capture richer multi-scale frequency domain features, the low-frequency component LL1 is subjected to multi-level learnable wavelet decomposition, ultimately forming multiple secondary sub-bands. In this embodiment, there are seven secondary sub-bands: LL2, LH2, HL2, HH2, LH1, HL1, and HH1. It should be noted that multi-level decomposition provides more detailed frequency information, thereby enabling deeper analysis of image content. Among the seven sub-bands obtained from the decomposition, LL2 contains more detailed global structural information, while LH2, HL2, and HH2 sub-bands contain more detailed high-frequency information in the horizontal, vertical, and diagonal directions. These sub-bands can capture edge and texture details in the image. This application, through this multi-level wavelet decomposition, can extract rich multi-scale features from the image, which is of great significance for various applications such as image analysis, compression, denoising, and feature extraction. These features not only help improve the efficiency and quality of image processing but also provide strong support for image recognition and classification tasks.
[0073] For ease of understanding, Figure 3 A visual comparison chart of different wavelet decomposition depths is provided, demonstrating the impact of different decomposition depths on image processing. The left side of the chart shows the original image, and the right side shows images processed at different wavelet decomposition depths, divided into three levels: Level 1, Level 2, and Level 3, each corresponding to a different decomposition depth. Level 1 is the lowest level of decomposition, preserving the main outline and major structural features of the image, but simplifying the details. Level 2 is a medium-level decomposition, with richer details than Level 1, while still retaining the main structural features. Level 3 is the highest level of decomposition, with the richest details and texture information. At each level, the image is decomposed into four parts: LL_Level, LH_Level, HL_Level, and HH_Level. The meaning of each part has been explained above and will not be repeated here.
[0074] Figure 4This paper demonstrates the impact of different wavelet decomposition depths (Level 1, Level 2, and Level 3) on model performance and uses multiple performance metrics for evaluation. PNSR (Peak Signal-to-Noise Ratio) measures image reconstruction quality; a higher value indicates better image quality. SSIM (Structural Similarity Index) measures the structural similarity of images; a higher value indicates better preservation of structural information. VS (Visual Similarity Index) measures the visual quality of images; a higher value indicates better visual quality. LPIPS (Perceptual Loss Perception) measures the perceptual difference between images; a lower value indicates better perceptual quality. KID (Kernel Inception Distance) measures the difference in distribution between the generated and real images; a lower value indicates better generated image quality. FID (Fréchet Inception Distance) is a metric for evaluating generated image quality; it measures the similarity between the generated and real images based on the statistical properties of image features. A lower FID value indicates that the generated image is statistically closer to the real image, i.e., the generated image quality is higher.
[0075] from Figure 4 As can be seen, Level 2 scores the highest on the PNSR, SSIM, and VS metrics, indicating that Level 2 achieves the best image reconstruction quality, preservation of image structural information, and visual quality. Level 2 scores the lowest on the LPIPS, KID, and FID metrics, indicating that the generated image from Level 2 is closest to the real image in similarity, has the smallest difference in distribution between the generated and real images, and is statistically closest to the real image. This also explains why the embodiments in this application employ multi-level learnable wavelet decomposition and extract seven sub-bands from the second level.
[0076] Preferably, considering that different subbands may have excessively high or low energy proportions during training, leading to an imbalance in feature representation, a subband energy balancing mechanism is further introduced to ensure that the frequency domain subbands after wavelet decomposition have a stable and controllable energy distribution. Specifically, subband b i The frequency domain energy is E i Then the total energy can be:
[0077]
[0078] Among them, b i Let α represent the i-th sub-band. i This indicates that the energy-learnable weights of each subband are obtained through Softmax normalization, ||b i ||2 represents the energy of the i-th subband, calculated using the L2 norm (i.e., the Euclidean number). In this formula, the weighting coefficient α of each subband is adjusted...i To achieve subband energy balance, for example: the energy of a certain subband ||b i If ||2 is relatively large, then by reducing its weighting coefficient α i This can reduce the contribution of that subband to the total energy; conversely, if the energy of a certain subband ||b i If ||2 is relatively small, its weighting coefficient α can be increased. i This increases the contribution of the subband to the total energy. In this way, energy balance among the subbands can be achieved, thereby improving the performance and effectiveness of signal processing.
[0079] After completing two levels of learnable wavelet transform and energy conservation, this application connects all subbands along the channel dimension to construct a unified multi-scale frequency domain representation tensor. This multi-scale frequency domain representation tensor is then converted into a two-dimensional sequence suitable for the Mamba sequence modeling module. It is understood that in deep learning, different network modules require inputs in specific formats. For the Mamba sequence modeling module, it is necessary to convert the tensor into a two-dimensional sequence input, for example, by flattening or rearranging the tensor to transform it from a multi-dimensional array into a one-dimensional sequence.
[0080] The structure of the Mamba sequence modeling module is as follows: Figure 2 As shown, the Mamba sequence modeling module includes components such as LayerNorm, Linear layers, Depthwise Separable Convolutional Layers (DWconv), and activation functions (σ). The specific execution process is as follows:
[0081] First, the two-dimensional sequence, transformed from the multi-scale frequency domain representation tensor, is input into the LayerNorm layer. The LayerNorm layer normalizes the input data, helping the model more effectively handle variations in the input data's distribution. This normalization operation is particularly important for sequence modeling because it reduces internal covariate bias, enabling the model to learn cross-regional contextual information more stably. In this way, the LayerNorm layer provides the model with a more stable and consistent training environment, helping to improve the model's ability to capture features from different regions.
[0082] Secondly, the Linear layer in the Mamba sequence modeling module transforms the data feature space, increasing or decreasing the dimensionality of features, thereby helping the model learn the global correlations between different frequency bands. This linear transformation enables the model to better understand and represent the features of different frequency bands, laying the foundation for deep interaction of multi-frequency information.
[0083] Furthermore, the DWconv layer effectively reduces computation and the number of parameters by decomposing standard convolution operations into depthwise convolution and pointwise convolution. This efficient convolution operation not only maintains the model's expressive power but also enables the model to extract features from different frequency bands more accurately. In sequence modeling, DWconv helps the model establish connections between different frequency bands, achieving more effective feature interaction.
[0084] Finally, the Mamba sequence modeling module employs a spiral scan method, which, compared to traditional scanning methods, more effectively connects the relationships between frequency bands. The spiral scan method optimizes the interaction between frequency bands through specific data flows and computational sequences, enabling the model to more comprehensively capture and utilize information from different frequency bands. The spiral scan method in the Mamba sequence modeling module not only improves the model's ability to process multi-frequency information but also provides structurally consistent support for tasks such as inverse wavelet reconstruction and image generation.
[0085] Preferably, to enhance the ability to preserve edge and structural details (such as vascular texture) in frequency domain modeling, embodiments of this application calculate the first-order spatial gradient difference between the reconstructed high-frequency subband and the original high-frequency subband in the horizontal and vertical directions based on a gradient preservation mechanism to construct a gradient preservation loss, thereby effectively improving the model's perception and recovery accuracy of texture and edge changes. Finally, all frequency band information is reconstructed layer by layer through a learnable inverse wavelet transform module, restoring it to the same resolution as the original image, thus achieving a reversible transformation from the frequency domain to the spatial domain. Therefore, the frequency domain prior Y frequency Frequency can be expressed as:
[0086]
[0087] in, This indicates a learnable wavelet transform, where H and L represent learnable high-pass and low-pass filters, respectively, and 2 represents a second-order wavelet transform. Mamba spiral Indicates a spiral Mamba scan. This represents the inverse learnable wavelet transform.
[0088] Specifically, the first-order spatial gradients in the horizontal and vertical directions of the reconstructed and original high-frequency subbands can be calculated using convolutional operations (e.g., pooling or Batch Normalization). Then, the absolute or squared values of the gradient differences between the reconstructed and original high-frequency subbands in the horizontal and vertical directions are calculated to quantify the degree of detail preservation during reconstruction. Next, based on the calculated gradient differences, a gradient-preserving loss function is constructed, which measures the mean squared error (MSE) or other suitable metric of the gradient differences between the reconstructed and original high-frequency subbands in the horizontal and vertical directions. This loss function is used to train the model; minimizing this loss function optimizes the model, thereby better preserving edge and texture information in the image. Finally, after training, the model's performance is evaluated, particularly its accuracy in perceiving and recovering edge and texture details. This can be evaluated by comparing the high-frequency subbands of the reconstructed and original images. Therefore, through this gradient-preserving mechanism, the model's ability to preserve edge and structural details in frequency domain modeling is significantly enhanced, thereby improving visual quality in applications such as image reconstruction, compression, and enhancement.
[0089] It is worth noting that the learnable wavelet frequency domain extractor 101 not only supports end-to-end training, but also significantly improves the model's performance in frequency domain decoupling, multi-scale feature extraction, and structural fidelity, laying a solid foundation for accurately generating FFA vascular networks and structural contours.
[0090] The structure of the Dual-pyramid Spatial-domain Feature Extractor 102 (DSFE) is as follows: Figure 5 As shown, it mainly consists of convolutional pyramid branches and pooling pyramid branches, aiming to establish an efficient connection between local modeling and global context awareness through multi-scale structural fusion, thereby achieving deep modeling of key lesion regions in CFP images.
[0091] It is worth noting that CFP images typically contain multiple complex and heterogeneous lesion regions, such as hard exudates, retinal hemorrhages, and neovascularization. These lesions often exhibit high or low fluorescence responses in FFA images. However, existing FFA image generation methods frequently suffer from incomplete representation or semantic ambiguity when processing these regions, leading to weakened or omitted pathological information in the generated results. Therefore, this application proposes a dual-pyramid spatial domain feature extractor to enhance the perception and semantic representation capabilities of complex lesion regions in FFA images.
[0092] In the convolutional pyramid branch, multi-scale convolutional kernels of 3×3, 5×5, and 7×7 are used to capture semantic features under different receptive fields, such as neovascularization, localized hemorrhage, and vascular occlusion. Then, channel compression is performed using 1×1 convolution (1×1Conv), and the feature map is divided into patch sequences. After introducing positional encoding, these sequences are fed into the Mamba sequence modeling module for sequence modeling. The structure of the Mamba sequence modeling module includes components such as layer normalization (LayerNorm), linear layers, depthwise separable convolutional layers (DWconv), and activation functions (σ). The function of each component has been explained above and will not be repeated here.
[0093] Preferably, considering that lesion areas are typically sparse and localized, meaning these areas are unevenly distributed in the image and occupy only a small portion of the space, this application embodiment designs a row-column scanning mechanism (e.g., ...) to ensure computational efficiency while fully covering these important local areas. Figure 3 (As shown by the row and column arrows in the Mamba sequence modeling module on the left). The core idea of this row-column scanning mechanism is to systematically scan the image to ensure that all important information in the image is captured, especially sparsely distributed lesion areas. For example, the mechanism first scans the image row by row, analyzing the image content from top to bottom, ensuring that information from each row is considered. Then, it scans the image column by column, analyzing from left to right, further ensuring that information from each column is captured. Through this row-column scanning approach, the model can achieve comprehensive semantic coverage of the image while maintaining high computational efficiency. Simultaneously, computational efficiency is improved because it avoids intensive computation on the entire image, instead reducing computation by selectively focusing on rows and columns. Furthermore, this mechanism helps the model better understand local features of the image, especially sparsely distributed lesion areas, thereby improving the accuracy of lesion detection.
[0094] Finally, the spatial domain constraint Y of the convolution pyramid branch Conv_pyramid It can be represented as:
[0095] Y Cov_pyramid =Mamba R-C (Conv 1×1 (Conv cat_3×5×7 (x)));Formula (5)
[0096] Among them, Mamba R-C Indicates a row-column Mamba scan, Conv 1×1 Represents a 1×1 convolution, Conv cat_3×5×7 This represents pyramidal convolutions of 3×3, 5×5, and 7×7.
[0097] In the pooling pyramid branch, multi-scale aggregation of contextual information is achieved through multi-scale adaptive average pooling (3×3, 5×5, and 7×7), enhancing the ability to characterize large-scale lesions such as macular regions and diffuse exudates. Subsequently, channel compression and the Mamba sequence modeling module are combined to integrate global semantic information.
[0098] Preferably, in medical image analysis, lesion regions often exhibit a sparse distribution. To effectively detect and analyze these sparsely distributed lesion regions while maintaining computational efficiency and ensuring comprehensive coverage of image semantics, this application provides a multi-scale modeling method based on a row-column scanning mechanism. The core of this method lies in using a row-column scanning mechanism to segment the image into blocks, thereby capturing the features of lesion regions at different scales. Specifically, this mechanism first divides the image into multiple small blocks, and then scans these blocks sequentially by row and column. In this way, the model can perform detailed analysis of each region in the image while maintaining high computational efficiency, ensuring that no part that may contain lesion information is missed. During multi-scale feature modeling, the row-column scanning mechanism can effectively model features at different scales. This is because lesion regions may exhibit different features at different scales, and multi-scale analysis can capture these features more comprehensively. For example, at smaller scales, the model may focus more on capturing the detailed features of the lesion region; while at larger scales, it may focus more on identifying the overall shape and location of the lesion region. This multi-scale feature modeling method based on row and column scanning mechanism not only ensures comprehensive coverage of image semantics while maintaining computational efficiency, but also more accurately identifies and locates sparsely distributed lesion areas.
[0099] Ultimately, the spatial domain constraint Y of the pooling pyramid branch Avg_pyramid Represented as:
[0100] Y Avg_pyramid =Mamba z (Conv 1×1 (Avg cat_3×5×7 (x)));Formula (6)
[0101] Among them, Mamba z Indicates a zigzag scan, Conv 1×1 Represents a 1×1 convolution, Avg cat_3×5×7 This represents the average pooling of pyramids with dimensions of 3×3, 5×5, and 7×7.
[0102] Finally, the local modeling of the convolutional pyramid branch and the global context-aware modeling of the pooling pyramid branch are fused, and the conditional prior Y extracted from the fused dual-pyramid spatial domain is used. Spatial Represented as:
[0103] Y Spatial =Concat(Y Conv_pyramid ,Y Avg_pyramid )Formula (7)
[0104] Here, Concat represents the concatenation operation. It should be understood that feature map concatenation refers to connecting two or more feature maps along a certain dimension (usually the channel dimension), thereby increasing the expressive power of the network, as it allows the network to learn the information contained in different feature maps simultaneously.
[0105] It is worth noting that the spatial domain lesion features extracted by the dual-pyramid spatial domain feature extractor 102 serve as a semantic prior for diffusion-based generation, effectively guiding the learning of potential representations and ensuring that the generated FFA image exhibits higher structural accuracy and semantic consistency in the lesion region.
[0106] The denoising stage includes the dual-domain conditional constraint Mamba module 103. The specific structure and principle of this module will be explained in detail below with reference to specific embodiments.
[0107] The structure of the Dual-domain Conditional Constraint Mamba module 103 (DCCM) is as follows: Figure 6 As shown, in order to integrate the frequency domain and spatial domain priors of CFP images, a state space modeling mechanism is used to achieve collaborative modeling of dual-domain features, thereby improving the model's ability to represent vascular structures and lesion areas and its cross-modal alignment accuracy.
[0108] It should be noted that in FFA image generation tasks based on diffusion models, accurate reconstruction of fine-grained lesion regions and complex vascular structures is crucial for the clinical usability of the images. However, due to significant differences in structural features and visual representation between the input CFP image and the target FFA image, the generation model needs stronger structural modeling capabilities and cross-modal semantic alignment capabilities to achieve accurate reconstruction of details and pathological features. Traditional diffusion models often use U-Net as the denoising network. While its convolutional structure performs well in low-level feature extraction, it has inherent limitations in capturing long-range dependencies and global contextual information, making it difficult to meet the high requirements for structural fidelity and detail reproduction in FFA image generation. To address this, this application proposes a Mamba module that integrates spatial and frequency domain constraints as the core denoising backbone of the diffusion process, replacing the traditional U-Net architecture.
[0109] Specifically, the input to the dual-domain conditional constraint Mamba module 103 has the following three types: The input noisy image is modularized into X...inputtoken The features are then fused and embedded with two types of conditional features. The first type of conditional feature (Condition1) is the frequency domain feature Y extracted by the learnable wavelet frequency domain extractor 101. frequency It is used to characterize the edges and structural contours of blood vessels; the second type of conditional feature (Condition2) is the spatial domain feature Y extracted by the double pyramid spatial domain feature extractor 102. Spatial It focuses on the contextual semantics of the lesion region. The fused patch sequence is fed into a biconditional attention module (e.g., ...). Figure 4 The two Attention modules shown model contextual relationships in the spatial and frequency domains respectively. By independently constructing Query, Key, and Value and introducing a conditional guidance mechanism, they enhance salient region features and improve semantic consistency. By independently constructing these three components, target features can be defined and extracted more accurately. The conditional guidance mechanism referred to here guides the allocation of attention based on specific conditions or prior knowledge. For example, in image processing, if certain regions are known to be salient (i.e., regions more important to the final task), the conditional guidance mechanism can focus more attention on these regions, thereby enhancing salient region features.
[0110] Subsequently, the contextual features are incorporated into the Mamba sequence modeling module. Bidirectional forward and backward scanning, along with continuous activation operations, enhance the expressive power of the contextual information sequence, thereby more accurately characterizing complex lesion boundaries and vascular structures. Bidirectional forward and backward scanning involves scanning and analyzing the contextual feature sequence from both the forward and backward directions. Compared to unidirectional scanning, bidirectional scanning can more comprehensively capture the dependencies between contextual information. For example, when processing medical images, the features of lesions or vascular structures may have spatial correlations. Bidirectional scanning can simultaneously consider the forward and backward information of these features, thus providing a more complete understanding of their position and role in the entire sequence. Continuous activation operations further enhance the expressive power of the contextual information sequence. This operation ensures the continuity and stability of contextual features during transmission and processing, avoiding information loss or distortion caused by discretization. This continuous activation operation better preserves the detailed information of contextual features, enabling the model to more sensitively capture subtle changes in complex lesion boundaries and vascular structures. Therefore, the Mamba state space modeling module significantly enhances the expressive power of contextual information sequences through bidirectional state scanning and continuous activation operations, enabling the model to more accurately characterize the shape, location, and interrelationships of complex lesion boundaries and vascular structures.
[0111] Finally, the dual-domain conditionally constrained Mamba module 103, combined with residual connections and normalized stable training, recovers the image structure through the Unpatchify operation, yielding the latent representation Y. output This is used to generate high-quality FFA images. Therefore, Y output The output process can be represented as:
[0112]
[0113] Where Attention represents the conditional attention mechanism, Mamba for-back This indicates a bidirectional Mamba scan, and Unpatchify represents the reverse process of patching to recover the complete image.
[0114] It should be understood that the core function of the Unpatchify operation is to recombine processed image patches to restore the original layout and structure of the image. The Unpatchify operation rearranges the processed image patches and seamlessly stitches them together based on pre-stored positional information, while employing smoothing mechanisms (such as interpolation or fusion algorithms) to resolve inconsistencies in boundary pixels, thereby restoring the complete structure and visual effect of the image. This process not only preserves the useful information extracted during the processing stage but also ensures the structural integrity and accuracy of the image, which is crucial for applications such as medical image analysis.
[0115] It should be understood that the module division in the embodiments of this application is illustrative and only represents a logical functional division. In actual implementation, there may be other division methods. Furthermore, the functional modules in the various embodiments of this application can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0116] The preceding text describes a modular training framework for achieving high-quality FFA image generation under cross-modal conditions, comprising a spatial domain feature extractor, a frequency domain feature extractor, and a backbone generation network. Multiple loss functions were designed to address the functional characteristics of different modules, and each module was independently optimized during training to fully exploit multimodal information and improve the quality and expressive power of the generated images.
[0117] The dual-pyramid spatial domain extractor 102 employs an instance-level discriminative unsupervised contrastive learning strategy, trained using InfoNCE loss without requiring labels or data augmentation. It maximizes the feature space similarity between positive pairs while distinguishing negative samples, achieving effective self-supervised representation learning. (InfoNCE loss function) It can be represented as:
[0118]
[0119] Where N represents the batch size, z i Let i represent the i-th sample. Let z represent the i-th positive sample, sim(a, b) represent the similarity measure, and τ be the temperature coefficient controlling the smoothness of the distribution. j Let j represent the j-th sample.
[0120] The learnable wavelet frequency domain extractor 101 introduces three loss functions, namely frequency domain contrast loss. Subband energy conservation loss Total loss of sub-band gradient This is to jointly constrain the structural consistency and detail representation capability of the frequency domain module output. The total loss can be expressed as:
[0121]
[0122] Subband energy conservation loss It can be represented as:
[0123]
[0124] Among them, E i Indicates subband b i The frequency domain energy.
[0125] The gradient difference loss can be expressed as:
[0126]
[0127] Among them, Y h Indicates the original high-frequency subband. This represents the subband after reconstruction in dimension d∈{H,W}. Let ||d|| represent the first-order difference operator (i.e., discrete gradient) in the d-th spatial dimension, and ||·||1 represent the L1 norm. This loss is summed over all high-frequency subbands to obtain the total subband gradient loss. for:
[0128]
[0129] Where S is the total number of high-frequency sub-bands, λ grad The hyperparameter (set to 0.05 in the experiment) is used to balance the weight of gradient loss in the overall optimization objective.
[0130] During the training phase of the main generator network, pre-trained parameters of the dual-pyramid spatial domain extractor and the learnable wavelet frequency domain extractor are loaded, and their structure is frozen, optimizing only the backbone generator. This phase employs InfoNCE-based contrastive loss as the core supervision, guiding the generated images to possess cross-modal semantic consistency under both spatial and frequency domain conditions. Therefore, the structural summary of the entire network loss can be expressed as:
[0131]
[0132] The preceding text describes a Mamba diffusion model based on dual spatial and frequency domain constraints, including its forward diffusion and denoising stages, as well as the learnable wavelet frequency domain extractor, dual-pyramid spatial domain feature extractor, dual-domain constrained Mamba module, and loss function involved in these stages. To facilitate a better understanding of the technical solution provided by those skilled in the art, the following text will further illustrate the model's performance on a dataset.
[0133] Specifically, this invention is implemented on a PyTorch-based computing platform, running on a system equipped with an RTX 4090 GPU and 24GB of memory. A series of experiments were conducted on the Hajeb and MPOS datasets. During training, a conditional diffusion model was used as the base diffusion model, and the AdamW optimizer was employed. Horizontal flipping was used for data augmentation on both datasets, and a VAE with pre-trained weights was used to extract latent features from both datasets. The network architecture was trained using mini-batch standardized mean and standard deviation.
[0134] The specific experiment is as follows:
[0135] The first step, fundus datasets: The model from this application was subjected to a series of experiments on the MPOS and Hajeb datasets. It should be understood that the MPOS dataset is the first multi-disease paired CFP (color fundus photography) and FFA (fundus fluorescein angiography) dataset, covering four different fundus diseases. This dataset is designed to facilitate research on synthesizing FFA images from non-invasive CFPs, which is of great significance for enhancing ophthalmic diagnosis and patient care, especially by reducing harm to patients through non-invasive procedures. The MPOS dataset was also used to train a generative adversarial network (GAN) to generate high-fidelity FFA images, demonstrating superior performance to existing state-of-the-art methods across multiple evaluation metrics. The release of the MPOS dataset aims to support further research in this field.
[0136] For the MPOS dataset, with image resolution of 1920×991 pixels, the DSFE and LWFE modules were trained using the AdamW optimizer with a learning rate of 1e-4 and a batch size of 8, for 2000 and 1000 iterations respectively. The diffusion-based generative network was trained for 5000 iterations with the same batch size. For the Hajeb dataset, with image resolution of 720×576 pixels, the DSFE and LWFE modules were trained using the AdamW optimizer with learning rates of 2e4 and 1e-4 respectively. DSFE was trained with a batch size of 6 for 2000 iterations, while LWFE was trained with a batch size of 4 for 5000 iterations. The diffusion generative model was trained with a batch size of 6 for 1000 iterations. Horizontal flipping was used as a data augmentation strategy for both datasets. A pre-trained VAE model was used as the latent spatial encoder-decoder.
[0137] The second step involves model evaluation based on multiple metrics: six evaluation metrics were used to comprehensively assess the quality of the generated images: FID (↓), KID (↓), LPIPS (↓), PSNR (↑), SSIM (↑), and vessel similarity (VS↑). FID and KID assess overall fidelity and realism; LPIPS measures perceptual similarity; PSNR and SSIM assess image quality based on pixel-level error and structural similarity, respectively; and vessel similarity specifically quantifies the consistency of vascular structures. These metrics together constitute a comprehensive framework for evaluating the quality of generated FFA images.
[0138] For example, Figure 7 This table presents the performance test results of various modules on the Hajeb dataset, showing the performance data of the LDM, LWFE, DSFE, and DCCM modules on six metrics: FID, KID, LPIPS, SSIM, PSNR, and VS. The symbol "√" in the table indicates that the corresponding module is activated, and the symbol "-" indicates that the module is not used. Comparing the first and last rows, the first row uses only the LDM module, while the last row uses all four modules simultaneously. In terms of performance, the first row significantly outperforms the last row in FID, KID, and LPIPS, while the first row significantly underperforms the last row in SSIM, PSNR, and VS. Comparing the data from other rows also clearly demonstrates that the model's performance metrics generally improve with the addition of more modules. In particular, when all modules are activated (fifth row), the model achieves optimal performance across all metrics, demonstrating that the combination of modules such as LWFE, DSFE, and DCCM effectively improves the quality of image reconstruction or generation.
[0139] Among them, the results of one-way / multi-way ANOVA of LWFE, DSFE, and DCCM are as follows: Figure 8 As shown in the diagram, from top to bottom, the row and column factors have a significant impact on the three evaluation metrics: PSNR, SSIM, and VS (all P-values are less than 0.05). This indicates that the combination of the LWFE, DSFE, and DCCM modules has a significant impact on image quality.
[0140] The third step is to generate the result image: This invention performs visual analysis on the generated fundus FFA image, such as... Figure 9 As shown, from left to right, the first and third columns are the input CFP images, and the second and fourth columns are the generated FFA images. By comparing visualization examples of different FFA images, it can be clearly seen that the generated images, enhanced by frequency domain information, have clearer vessel contours, fuller structures, and higher overall visual quality. Furthermore, the lesion regions in the generated FFA images are effectively modeled using a double-pyramid structure, resulting in continuous lesion structures and rich details.
[0141] like Figure 10 The diagram illustrates a flowchart of an FFA image generation method based on a dual-domain constrained Mamba diffusion model according to an embodiment of this application, including the following steps:
[0142] Step S1001: In the forward diffusion stage, the CFP image is encoded into a latent representation and Gaussian noise is injected step by step; features are extracted from the latent representation of the injected noise to extract spatial domain features to characterize the spatial information of the lesion area and frequency domain features to characterize the frequency domain contour of the vascular structure.
[0143] Step S1002: In the denoising stage, the input sequence composed of initial noise, spatial domain features, and frequency domain features is denoised under the guidance of joint modeling in the spatial and frequency domains; the denoised latent representation is decoded and reconstructed by the VAE decoder to generate an FFA image.
[0144] It should be understood that the specific process of performing the corresponding steps in each method step has been described in detail in the above system embodiments, and will not be repeated here for the sake of brevity. It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" mean examples, illustrations or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or more advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner. In addition, in the embodiments of this application, "at least one" means one or more, and "more" means two or more. "And / or" describes the relationship between related objects, indicating that there can be three relationships, for example, A and / or B, which can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can be represented as: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single or multiple.
[0145] Figure 11 This is a schematic block diagram of a computer device provided in an embodiment of this application. Figure 11 As shown, the computer device includes at least one processor 1101, a memory 1102, at least one network interface 1103, and a user interface 1105. The various components in the device are coupled together via a bus system 1104. It is understood that the bus system 1104 is used to implement communication between these components. In addition to a data bus, the bus system 1104 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 11 The general will label all buses as bus systems.
[0146] The user interface 1105 may include a monitor, keyboard, mouse, trackball, clicker, button, touchpad, or touch screen.
[0147] It is understood that memory 1102 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM) or programmable read-only memory (PROM), used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memories described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable categories of memory.
[0148] In this embodiment of the invention, the memory 1102 is used to store various types of data to support the operation of the electronic terminal 1100. Examples of this data include: any executable program for operation on the electronic terminal 1100, such as the operating system 11021 and application programs 11022; the operating system 11021 contains various system programs, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks. The application program 11022 may contain various applications, such as a media player, browser, etc., for implementing various application services. The implementation of the FFA image generation method based on the dual-domain constrained Mamba diffusion model provided in this embodiment of the invention can be included in the application program 11022.
[0149] The methods disclosed in the above embodiments of the present invention can be applied to processor 1101, or implemented by processor 1101. Processor 1101 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 1101 or by instructions in the form of software. The processor 1101 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 1101 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. General-purpose processor 1101 may be a microprocessor or any conventional processor, etc. The steps of the accessory optimization method provided in the embodiments of the present invention can be directly reflected as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in a memory. The processor reads the information in the memory and combines it with its hardware to complete the steps of the aforementioned method.
[0150] In an exemplary embodiment, the electronic terminal 1100 may be used by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs) to execute the aforementioned method.
[0151] According to the method provided in the embodiments of this application, this application also provides a computer program product, which includes: computer program code, which, when run on a computer, causes the computer to execute the above-described FFA image generation method based on the dual-domain constrained Mamba diffusion model.
[0152] According to the method provided in the embodiments of this application, this application also provides a computer-readable storage medium storing program code, which, when run on a computer, causes the computer to perform the above-described method.
[0153] As used in this specification, the terms "component," "module," "system," etc., are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. As illustrated, applications running on computing devices and computing devices can both be components. One or more components may reside in a process and / or an execution thread, and components may be located on a single computer and / or distributed among two or more computers. Furthermore, these components can be executed from various computer-readable media on which various data structures are stored. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).
[0154] Those skilled in the art will recognize that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0155] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0156] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0157] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0158] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0159] In the above embodiments, the functions of each functional unit can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. A computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs, DVDs), or semiconductor media (e.g., solid-state disks, SSDs, etc.).
[0160] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0161] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0162] In summary, this application provides an FFA image generation system, method, medium, and apparatus based on a dual-domain constrained Mamba diffusion model. This invention proposes a spatial-frequency domain dual-constraint Mamba diffusion model (SFDC-MambaDiff), which integrates dual spatial and frequency domain constraints for cross-modal FFA image generation from CFP images. A learnable wavelet feature extractor effectively extracts key vascular frequency components from fundus CFP images, enhancing the structured guidance for vascular edge representation in the generated FFA image. In the spatial domain, a dual-pyramid feature extractor effectively fuses local details and global semantics of fundus lesion regions, improving the consistency of lesion representation in the generated FFA image. Furthermore, the designed dual-domain conditional Mamba module serves as the denoising core, effectively coordinating spatial and frequency information to improve image clarity and accuracy. Experimental validation on two public datasets, Hajeb and MPOS, and one private clinical dataset achieved PSNRs of 30.11 and 29.40, respectively. This demonstrates that the present invention performs excellently in preserving vascular details, generating structures, restoring lesions, and ensuring semantic consistency in cross-modal FFA image generation. It can fully utilize fundus CFP image information and effectively generate high-quality FFA fundus images across modalities, providing new possibilities for low-cost screening of retinal diseases. Therefore, this application effectively overcomes the various shortcomings of the prior art and has high industrial application value.
[0163] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. An FFA image generation system based on a dual-domain constrained Mamba diffusion model, characterized in that, include: The system comprises a VAE encoder, a dual-pyramid spatial domain feature extractor, and a learnable wavelet frequency domain extractor located in the forward diffusion stage. A CFP image is input into the pre-trained VAE encoder to be encoded into a latent representation, and Gaussian noise is progressively injected. The noisy latent representation is then input into the dual-pyramid spatial domain feature extractor and the learnable wavelet frequency domain extractor to extract spatial domain features representing the spatial information of the lesion region and frequency domain features representing the frequency domain contour of the vascular structure, respectively. The learnable wavelet frequency domain extractor uses a set of learnable Haar wavelet filters as the initialization scheme for the learnable discrete wavelet transform, corresponding to low-pass and high-pass filtering operations, and parameterizes them as backpropagation-capable variables, allowing them to adaptively adjust during training to match the frequency domain distribution of the input image. The Haar wavelet filters perform convolution operations in four directions on the input CFP image to obtain low-frequency and high-frequency subbands. Furthermore, multi-level learnable wavelet decomposition is performed on the low-frequency components to ultimately form multiple secondary subbands. The dual-domain conditional constraint Mamba module and VAE decoder are located in the denoising stage; the input of the dual-domain conditional constraint Mamba module includes an input sequence composed of initial noise, spatial domain features, and frequency domain features; the input sequence is input to the dual-conditional guided Mamba module for denoising guided by joint modeling in the spatial and frequency domains; the denoised latent representation is decoded and reconstructed by the VAE decoder to generate an FFA image.
2. The FFA image generation system based on the dual-domain constrained Mamba diffusion model according to claim 1, characterized in that, The learnable wavelet frequency domain extractor combines learnable discrete wavelet transform with subband energy balance mechanism and Mamba sequence modeling module.
3. The FFA image generation system based on the dual-domain constrained Mamba diffusion model according to claim 2, characterized in that, The learnable wavelet frequency domain extractor calculates the first-order spatial gradient difference between the reconstructed high-frequency subband and the original high-frequency subband in the horizontal and vertical directions based on the gradient preservation mechanism to construct the gradient preservation loss; and, through the learnable inverse wavelet transform module, it reconstructs all frequency band information layer by layer to restore it to the same resolution as the original image, so as to perform a reversible transformation from the frequency domain to the spatial domain.
4. The FFA image generation system based on the dual-domain constrained Mamba diffusion model according to claim 1, characterized in that, The dual-pyramid spatial domain feature extractor includes a convolutional pyramid branch and a pooling pyramid branch; In the convolutional pyramid branch, multi-scale convolutional kernels are used to capture semantic features under different receptive fields, followed by channel compression, and the feature map is divided into patch sequences. After introducing positional encoding, it is sent to the Mamba sequence modeling module for sequence modeling. In the pooling pyramid branch, multi-scale adaptive average pooling is used to aggregate contextual information at multiple scales, and channel compression and Mamba sequence modeling modules are combined to integrate global semantic information. The local modeling of the convolutional pyramid branch and the global context-aware modeling of the pooling pyramid branch are fused to obtain the conditional priors for the extraction of the dual-pyramid spatial domain.
5. The FFA image generation system based on the dual-domain constrained Mamba diffusion model according to claim 4, characterized in that, The Mamba sequence modeling module uses a row-column scanning mechanism for image scanning.
6. The FFA image generation system based on the dual-domain constrained Mamba diffusion model according to claim 1, characterized in that, The input to the dual-domain conditional constraint Mamba module includes: modularizing the input noisy image into... The patch sequence is obtained by fusing and embedding the patch with two types of conditional features respectively; wherein, the first type of conditional feature is the frequency domain feature extracted by the learnable wavelet frequency domain extractor; and the second type of conditional feature is the spatial domain feature extracted by the dual pyramid spatial domain feature extractor.
7. The FFA image generation system based on the dual-domain constrained Mamba diffusion model according to claim 6, characterized in that, The patch sequence is fed into the biconditional attention module, which models the contextual relationships in the spatial and frequency domains and independently constructs queries, keys, and values. The contextual features are input into the Mamba sequence modeling module, which performs bidirectional state scanning and continuous activation operations. The operation results are trained by residual connection and normalization, and then the image structure is restored by the Unpatchify operation to obtain the potential representation used to finally generate the FFA image.
8. The FFA image generation system based on the dual-domain constrained Mamba diffusion model according to claim 1, characterized in that, The dual-pyramid spatial domain feature extractor is trained using the InfoNCE loss function; the learnable wavelet frequency domain extractor is trained based on the frequency domain contrast loss, subband energy conservation loss, and subband gradient total loss as the total loss function; during the training of the dual-domain conditional constraint Mamba module, the pre-trained parameters of the dual-pyramid spatial domain feature extractor and the learnable wavelet frequency domain extractor are loaded and their structures are frozen, and their loss functions consist of the spatial domain InfoNCE loss, the frequency domain InfoNCE loss, and the total loss of the learnable wavelet frequency domain extractor.
9. A method for generating FFA images based on a dual-domain constrained Mamba diffusion model, characterized in that, An FFA image generation system based on a dual-domain constrained Mamba diffusion model as described in any one of claims 1 to 8, the system comprising: a VAE encoder, a dual-pyramid spatial domain feature extractor, and a learnable wavelet frequency domain extractor located in the forward diffusion stage; and a dual-domain conditionally constrained Mamba module and a VAE decoder located in the denoising stage; the method comprising: In the forward diffusion stage, the CFP image is encoded into a latent representation and Gaussian noise is progressively injected. Feature extraction is performed on the latent representation with injected noise to extract spatial domain features representing the spatial information of the lesion region and frequency domain features representing the frequency domain contour of the vascular structure. Specifically, the latent representation with injected noise is input into the dual-pyramid spatial domain feature extractor and the learnable wavelet frequency domain extractor to extract spatial domain features representing the spatial information of the lesion region and frequency domain features representing the frequency domain contour of the vascular structure. The learnable wavelet frequency domain extractor uses a set of learnable Haar wavelet filters as the initialization scheme for the learnable discrete wavelet transform, corresponding to low-pass and high-pass filtering operations, and parameterizes them as backpropagation variables so that they can adaptively adjust during training to match the frequency domain distribution of the input image. The Haar wavelet filters perform convolution operations in four directions on the input CFP image to obtain low-frequency and high-frequency subbands. Furthermore, multi-level learnable wavelet decomposition is performed on the low-frequency components to ultimately form multiple secondary subbands. In the denoising stage, the input sequence is composed of initial noise, spatial domain features, and frequency domain features; the input sequence is denoised under the guidance of joint modeling in the spatial and frequency domains; the denoised latent representation is decoded and reconstructed by the VAE decoder to generate an FFA image.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the FFA image generation method based on the dual-domain constrained Mamba diffusion model as described in claim 9.
11. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the FFA image generation method based on the dual-domain constrained Mamba diffusion model as described in claim 9.