OCTA fundus image generation system, method, medium, and apparatus based on mamba and diffusion model

By using an image generation system based on Mamba and diffusion models, and leveraging cross-order attention and adaptive feature conditional cues, the problems of lost vascular details and insufficient contrast in OCTA fundus image generation are solved, resulting in high-quality OCTA images suitable for the diagnosis of ophthalmic diseases.

CN120198318BActive Publication Date: 2026-04-24SHANGHAI FIRST PEOPLES HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI FIRST PEOPLES HOSPITAL
Filing Date
2025-03-07
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing OCTA fundus image generation technology suffers from problems such as loss of vascular details, insufficient contrast, and numerous artifacts, making the generated images unsuitable for the diagnosis of ophthalmic diseases.

Method used

An image generation system based on Mamba and a diffusion model is adopted. By combining Mamba blocks and a diffusion model, cross-order attention conditional prompts and adaptive feature conditional prompts are used to generate OCT fundus images. Cross-order attention conditional prompts are used to extract OCT fundus image features, and adaptive feature conditional prompts are embedded in the diffusion model for supervised prompting, thus generating high-quality OCT fundus images.

Benefits of technology

It improves the contrast of OCTA images, reduces artifacts, enhances the spatial continuity of blood vessels, and significantly improves the quality of generated OCTA images, meeting the needs of ophthalmic disease diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198318B_ABST
    Figure CN120198318B_ABST
Patent Text Reader

Abstract

The application provides an OCTA fundus image generation system, method, medium and device based on Mamba and diffusion model, the application effectively predicts the diffusion noise of the accurate OCTA image from the potential noise of the OCT by designing a unique Mamba block, reduces the artifacts of the generated image, and improves the contrast of the OCTA image generation. Secondly, the proposed bidirectional Bow scanning scheme effectively alleviates the mechanism challenge of the original Mamba scanning, enhances the spatial continuity of the generated OCTA image, and improves the blood vessel extension of the generated image. In addition, the designed conditional prompt improves the understanding of the OCT image in combination with the Mamba block and the diffusion model, and further improves the authenticity of the generated fundus OCTA image. The experimental verification on two public data sets of OCTA500-3M and OCTA500-6M and a private data set respectively reaches 87.5% and 88.1% of SSIM, indicating that the application can fully utilize the fundus OCT image information and effectively generate high-quality fundus OCTA images, providing a new possibility for low-cost screening of retinal diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an OCTA fundus image generation system, method, medium, and apparatus based on Mamba and diffusion models. Background Technology

[0002] The incidence of retinal diseases is rising sharply with increasing life stress. To aid in the diagnosis of complex retinal diseases, numerous retinal imaging technologies have emerged. For example, CFP (Color Fundus Photography) and OCT (Optical Coherence Tomography) are two common retinal imaging techniques that have made significant contributions to ophthalmological research. However, for some complex retinal diseases, such as diabetic retinopathy and other diseases involving intervascular connections in the retina, existing retinal imaging techniques prove insufficient. Therefore, new imaging technologies or methods for generating retinal images are crucial.

[0003] Currently, based on OCT imaging, a new retinal microvascular imaging technology has emerged—OCTA (Optical Coherence Tomography Angiography). Its imaging principle is mainly based on the combination of optical coherence tomography and blood flow imaging technologies. It captures the movement of red blood cells in blood vessels and uses special computational methods for signal analysis and vascular imaging. OCTA can display the microvascular structure of the retina and choroid at high resolution, providing strong support for the diagnosis and treatment of ophthalmic diseases. However, it is worth noting that although OCTA fundus vascular imaging technology has brought a major breakthrough in the field of ophthalmic diagnosis, its image acquisition is highly dependent on specific, highly specialized medical equipment. The purchase, maintenance, and operating costs of these devices are extremely high, constituting a significant economic burden for many medical institutions. This, to some extent, limits the speed and scope of the popularization of OCTA fundus vascular imaging technology, becoming a major challenge on its path to widespread application. Therefore, how to explore solutions to reduce costs and improve equipment accessibility while maintaining technological advancement is a key issue that urgently needs to be addressed in the development of OCTA fundus vascular imaging technology. To reduce costs, researchers have developed deep learning-based methods to generate OCTA fundus images from existing OCT fundus images. These methods can be mainly divided into three categories: the first category uses a unique encoder-decoder structure to convert OCT B-Scans into OCTAB-Scans; the second category uses an unsupervised 3D domain adaptive method to map OCT B-Scans to the feature space to directly generate OCTA fundus images; and the third category uses a generative model to directly generate OCTA images from OCT images. Although these methods have made some progress in OCT to OCTA conversion, they still have limitations: (1) Loss of vascular details: existing models have difficulty capturing microvascular structures, resulting in a lack of key information in the generated OCTA fundus images. (2) Insufficient contrast: the contrast between the vascular region and the background in the generated images is low, which is not conducive to doctors' interpretation. (3) More artifacts: background artifacts are easily introduced during the generation process, and poor vascular extensibility affects medical applications. Summary of the Invention

[0004] In view of the shortcomings of the prior art described above, the purpose of this application is to provide an OCTA fundus image generation system, method, medium and device based on Mamba and diffusion models, to solve the technical problems of loss of vascular details, insufficient contrast and many artifacts encountered in the process of converting OCT to OCTA.

[0005] To achieve the above and other related objectives, a first aspect of this application provides an OCTA fundus image generation system based on Mamba and a diffusion model, comprising: a conditional prompting module for generating cross-order attention conditional prompts and adaptive feature conditional prompts; the cross-order attention conditional prompts are configured to extract image features from OCT fundus images, and the adaptive feature conditional prompts are configured to extract image features from OCTA fundus images; an image generation module deploying a Mamba block and a diffusion model; the Mamba block integrating the cross-order attention conditional prompts; the OCT fundus image being converted into a latent space representation by a variational autoencoder and then input into the Mamba block; the Mamba block generating noise prediction results based on the cross-order attention conditional prompts and the Mamba scanning method; and supervising the diffusion model by combining the noise prediction results output by the Mamba block with the adaptive feature conditional prompts to generate OCTA fundus images.

[0006] In some embodiments of the first aspect of this application, the generation method of the cross-order attention condition cues includes: dividing a number of B-Scans that make up an OCT fundus image into a number of image patches; performing max pooling and average pooling on the divided image patches respectively; generating at least two attention matrices by calculating the max pooling results and average pooling results using different computational operations; and resampling the features of the image patches respectively using the attention matrices to obtain the spatial information and channel information of the OCT fundus image.

[0007] In some embodiments of the first aspect of this application, the process of generating spatial information of the OCT fundus image includes: generating a first attention matrix by multiplying the max pooling result and the average pooling result of the image patch element by element, and resampling the normalized image patch with the attention matrix to obtain the spatial information of the fundus image.

[0008] In some embodiments of the first aspect of this application, the process of generating channel information of the OCT fundus image includes: adding the max pooling result and average pooling result of the image patch element by element, mapping them through a fully connected layer to generate a second attention matrix, and resampling the second attention matrix with the original features to obtain the channel information of the fundus image.

[0009] In some embodiments of the first aspect of this application, the adaptive conditional cue generation process includes: encoding vascular details of OCTA fundus images into dynamic conditional signals using learnable cue blocks and embedding them into the image generation process of a diffusion model.

[0010] In some embodiments of the first aspect of this application, in the image generation module, the Mamba block generates a noise prediction result based on cross-order attention conditional cues and the Mamba scanning method, which includes: inputting the latent space representation into the Mamba block after normalization and performing bi-branch processing:

[0011] In the first branch, cross-order attention conditional cues are used, combined with spatial and channel information from OCT fundus images, and processed through a multi-head self-attention mechanism to obtain the first branch processing result; in the second branch, the latent spatial representation is scanned sequentially in different directions to obtain the second branch processing result; the noise prediction result is obtained by fusing the first branch processing result and the second branch processing result.

[0012] In some embodiments of the first aspect of this application, the scanning method in the second branch includes: a bidirectional scanning mechanism based on alternating forward and backward scanning, and scanning by automatically switching scanning paths in different Mamba layers using preset multiple dynamic scanning methods and modular arithmetic.

[0013] To achieve the above and other related objectives, a second aspect of this application provides a method for generating OCTA fundus images based on Mamba and a diffusion model, comprising: generating cross-order attention conditional cue and adaptive feature conditional cue; the cross-order attention conditional cue is configured to extract image features from an OCT fundus image, and the adaptive feature conditional cue is configured to extract image features from an OCTA fundus image; the OCT fundus image is converted into a latent space representation by a variational autoencoder and then input into a Mamba block; the Mamba block integrates the cross-order attention conditional cue; the Mamba block generates a noise prediction result based on the cross-order attention conditional cue and the Mamba scanning method; and the diffusion model is supervised and prompted by combining the noise prediction result output by the Mamba block with the adaptive feature conditional cue to generate the OCTA fundus image.

[0014] To achieve the above and other related objectives, a third aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the OCTA fundus image generation method based on the Mamba and diffusion model.

[0015] To achieve the above and other related objectives, a fourth aspect of this application provides a computer program product comprising computer program code that, when executed on a computer, causes the computer to implement the OCTA fundus image generation method based on the Mamba and diffusion model.

[0016] To achieve the above and other related objectives, a fifth aspect of this application provides a computer apparatus, including a memory, a processor, and a computer program stored in the memory; the processor executes the computer program to implement the OCTA fundus image generation method based on the Mamba and diffusion model.

[0017] As described above, the OCTA fundus image generation system, method, medium, and apparatus based on Mamba and diffusion models of this application have the following beneficial effects: By designing unique Mamba blocks, the diffusion noise of accurate OCTA images is effectively predicted from potential OCT noise, reducing artifacts in the generated images and improving the contrast of the generated OCTA images. Secondly, the proposed bidirectional Bow scanning scheme effectively alleviates the mechanistic challenges of the original Mamba scan, enhances the spatial continuity of the generated OCTA images, and improves the vascular extension of the generated images. In addition, the designed conditional cueing, combined with Mamba blocks and diffusion models, improves the understanding of OCT images, further enhancing the realism of the generated fundus OCTA images. Experimental verification on two public datasets, OCTA500-3M and OCTA500-6M, and one private dataset achieved SSIMs of 87.5% and 88.1%, respectively, indicating that this invention can fully utilize fundus OCT image information to effectively generate high-quality fundus OCTA images, providing new possibilities for low-cost screening of retinal diseases. Attached Figure Description

[0018] Figure 1 The diagram shown is a schematic representation of the framework of an OCTA fundus image generation system based on the Mamba and diffusion model in one embodiment of this application.

[0019] Figure 2 The diagram shown is a structural schematic of a condition prompt module in one embodiment of this application.

[0020] Figure 3 The diagram shown is a structural schematic of a Mamba block in one embodiment of this application.

[0021] Figure 4 The diagram shown is a schematic representation of four Bow scanning schemes in one embodiment of this application.

[0022] Figure 5 The diagram shows the performance of the present invention on two datasets in one embodiment of this application.

[0023] Figure 6 The diagram shown is a flowchart illustrating an OCTA fundus image generation method based on Mamba and a diffusion model in one embodiment of this application.

[0024] Figure 7The diagram shown is a structural schematic of a computer device according to an embodiment of this application. Detailed Implementation

[0025] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.

[0026] Before providing a further detailed description of the present invention, the nouns and terms used in the embodiments of the present invention are explained, and the nouns and terms used in the embodiments of the present invention are subject to the following interpretations:

[0027] <1> The Mamba model is an advanced state-space model (SSM) that employs a linear time series modeling architecture combined with a selective state space. Its core is the selective state space layer, which allows the model to selectively propagate or suppress information based on the input at each step. It can process sequences linearly according to their length and has applications in language processing, genomics, audio analysis, and image generation.

[0028] <2> Diffusion Model: Based on the diffusion process in physics, it simulates the forward diffusion process of data from an ordered state to a disordered state by adding noise, and the reverse diffusion process of data from a disordered and noisy state back to an ordered original data state. By learning the inverse process, it can model the data distribution and generate new data samples.

[0029] <3> OCT (Optical Coherence Tomography) images: Images generated by optical coherence tomography are visualization results obtained by performing tomographic imaging of biological tissues using near-infrared light and optical interference principles.

[0030] <4> OCTA (Optical Coherence Tomography Angiography) images: Images presented by optical coherence tomography angiography are three-dimensional images that can image the retina and choroidal vessels.

[0031] To address the technical problems mentioned above, this invention proposes an OCTA fundus image generation system based on Mamba and a diffusion model. Specifically, a unique Mamba block is designed, which, by combining multiple Mamba blocks, effectively predicts the diffusion noise of the OCTA fundus image from the latent noise of OCT, thereby improving the contrast of the generated OCTA fundus image and reducing artifacts. Secondly, to ensure the continuity of blood vessels in the generated OCTA fundus image, a bidirectional Bow scanning scheme is further proposed. Simultaneously, this invention also designs a special conditional cue, consisting of a cross-order attention cue and an adaptive conditional cue, designed to guide the diffusion model in generating clear and vivid blood vessels in the OCTA fundus image. Experiments on public and private datasets demonstrate that this invention can fully utilize OCT fundus image information to generate high-quality OCTA fundus images.

[0032] The following section will provide a detailed explanation of the conditional prompt generation system based on the Mamba model and diffusion model provided by the present invention, with reference to specific embodiments.

[0033] Figure 1 A schematic diagram of the framework structure of an OCT fundus image generation system based on the Mamba and diffusion model is shown in an embodiment of the present invention. This system is used to generate OCT fundus images from OCT fundus images.

[0034] The conditional prompting module is used to generate cross-order attention conditional prompts and adaptive feature conditional prompts; the cross-order attention conditional prompts are configured to extract image features from OCT fundus images, and the adaptive feature conditional prompts are configured to extract image features from OCT fundus images.

[0035] Cross-order attention conditional prompts extract image features from OCT fundus images by combining spatial attention and channel attention mechanisms. Adaptive feature conditional prompts construct an adaptive learning sequence to dynamically learn and extract image features from OCTA fundus images. It is worth noting that the OCTA fundus images generated in this invention need to maintain a high degree of consistency with OCT fundus images in morphology, structure, and features to meet the accuracy requirements for fundus disease diagnosis. However, existing unconditional diffusion models can only generate samples based on random noise, lacking effective control over the generation process and making it difficult to accurately match specific features of fundus images. Therefore, their generation results cannot meet the stringent requirements for fundus disease diagnosis. Furthermore, traditional conditional diffusion models often fail to effectively learn the features of specific regions in datasets with uneven distribution or complex structures, leading to biased generated samples and reducing the overall quality of the generated samples, thus affecting their application value in fundus disease diagnosis. In view of this, this invention designs unique conditional prompts to assist the generation process and ensure the consistency of generated OCTA fundus images.

[0036] Figure 2 This diagram illustrates the structure of the conditional prompt module. The input OCT fundus image consists of multiple B-scans (B-scans), each a depth scan of the fundus tissue along a specific direction, reflecting two-dimensional tomographic information of the fundus structure. It should be understood that B-scans (B-scans) are a type of two-dimensional cross-sectional image in optical coherence tomography (OCT), providing two-dimensional cross-sectional information of the sample's internal structure, capable of displaying the structural features of the sample at different depths and lateral positions. In ophthalmology, B-scans can clearly display the hierarchical structure and pathological changes of ocular tissues such as the retina, aiding doctors in disease diagnosis and research.

[0037] To achieve effective processing of OCT fundus images, the generation process of cross-order attention conditional cues is as follows:

[0038] First, the B-scans are divided into multiple 7×7 pixel image patches. This division method aims to decompose complex OCT fundus images into smaller, more manageable units, allowing for targeted analysis or processing of each patch subsequently, thereby improving the efficiency and accuracy of image processing.

[0039] Subsequently, the segmented image patches are subjected to max pooling (Max) and average pooling (Avg) respectively. At least two attention matrices are generated based on the max pooling and average pooling results of the image patches. The image patches are then resampled using the attention matrices to obtain the spatial and channel information of the fundus image.

[0040] Specifically, the image patch is input into max pooling and average pooling respectively, and the local feature representation (avg) of the patch is extracted respectively. out and global feature representation max out As a subsequent parameter sharing, the patch region x input max pooling and average pooling can be expressed as:

[0041] max out =Max(x); Formula (1)

[0042] avg out =Avg(x); Formula (2)

[0043] In this context, Max(·) and Avg(·) represent max pooling and average pooling, respectively.

[0044] In some examples, a first attention matrix is ​​generated by element-wise multiplying the max pooling and average pooling results of the image patches. This attention matrix is ​​then used to resample the normalized image patches to obtain the spatial attention information of the fundus image. Specifically, the average pooling (avg) is used to... out and max out Element-wise multiplication is performed to form a first attention matrix. Then, feature resampling is performed between this first attention matrix and the normalized patch image. The weights of important feature regions in the overlapping areas are amplified to highlight salient information at specific locations, i.e., the spatial information of the fundus image. att For example, the characteristic information between the sclera and other membranes. Spatial information. att It can be described as:

[0045]

[0046] Where σ1 represents the sigmoid function and σ2 represents LayerNorm normalization.

[0047] In some examples, a second attention matrix is ​​generated by element-wise summing of the max pooling and average pooling results of the image patches and mapping them through a fully connected layer. This second attention matrix is ​​then resampled with the original features to obtain the channel attention information of the fundus image. Specifically, the average pooling (avg) is... out and max out Element-wise summation is performed, followed by further mapping through a fully connected layer to form a second attention matrix. By resampling the features of the original patch channel information with the second attention matrix, the local features and global distribution are balanced to generate smoother and more robust weights, i.e., the channel information of the fundus image. attThis allows the features of each channel to retain important local information while better representing global characteristics. (Channel information) att It can be represented as:

[0048]

[0049] σ1 represents the sigmoid function, and σ3 represents the fully connected layer.

[0050] Finally, the spatial and channel information of the fundus image is weighted and fused, and then normalized after skip connection with the original image patch to extract features that are significant in both spatial and channel dimensions, forming a cross-order attention conditional cue. This cross-order attention conditional cue serves as a noise prediction cue in the Mamba block, assisting in the noise prediction of OCTA generation. Specifically, the cross-order attention conditional cue... att It can be represented as:

[0051]

[0052] Where σ² represents LayerNorm normalization, spatial att Represents spatial information of fundus images, channel att This represents the channel information of the fundus image.

[0053] To achieve effective processing of OCTA fundus images, the adaptive feature-based conditional cue generation process includes: This invention encodes vascular details from OCTA fundus images into dynamic conditional signals using learnable cue blocks and embeds them into the image generation process of a diffusion model. This allows for fine-grained control over the generated results. The learnable cue blocks mentioned here refer to those configured with a size of X∈R... C×H×W The cue block is defined by X, a three-dimensional tensor; C, the number of channels, representing the feature dimension at each spatial location; and H×W, the height and width, corresponding to the spatial resolution of the image. The cue block is a learnable parameter matrix whose core task is to adaptively learn vascular details from OCTA fundus images. Each parameter (i.e., an element in the tensor) can be trained independently, allowing the model to dynamically adjust its focus on different vascular features.

[0054] Adaptive conditional prompts for fundus images are input into the diffusion model as conditional signals, providing supervised prompts when the model generates fundus images from predicted noise. Through adaptive conditional learning, the model captures the original attribute features of the fundus image itself, and then uses these attribute features as supervision to improve the realism of the generated fundus images. Specifically, the adaptive conditional prompts for fundus images function in two phases: first, during the training phase, gradients are backpropagated to update the parameters of the prompt blocks by comparing the generated fundus images with real images; second, during the generation phase, the learned prompt block parameters guide the diffusion model to generate images that more closely resemble the real vascular structure. The goal of the diffusion model is progressive denoising, and the prompt blocks supervise this process by interacting with the intermediate features of the diffusion model at each denoising step, such as through attention mechanisms or feature concatenation. This interaction causes the model to prioritize preserving the vascular details and attribute features learned by the prompt blocks when generating fundus images.

[0055] The image generation module deploys a Mamba block and a diffusion model; the Mamba block integrates the cross-order attention conditional cue; the OCT fundus image is converted into a latent space representation by a variational autoencoder and then input into the Mamba block; the Mamba block generates noise prediction results based on the cross-order attention conditional cue and the Mamba scanning method; the diffusion model is supervised and prompted by combining the noise prediction results output by the Mamba block with the adaptive feature conditional cue to generate the OCT fundus image.

[0056] It should be understood that the image generation module in this embodiment combines the Mamba model and the diffusion model, used for segmentation and image generation tasks respectively. The Mamba model is a deep learning architecture based on a state-space model (SSM), adept at processing long sequence data and efficiently capturing long-range dependencies, excelling at handling both global and local information in image processing. The diffusion model is a generative model that generates high-quality, diverse images by progressively removing noise. The image generation module leverages Mamba's ability to capture long-range dependencies and the diffusion model's image generation capabilities to improve the performance of both image generation and segmentation.

[0057] OCT fundus images are encoded using a Variational Autoencoder (VAE) to convert them into a latent space representation. It should be understood that in a VAE, the encoder is responsible for compressing the input data (e.g., OCT fundus images) into a latent space representation. This latent space representation is not a fixed value, but a distribution, typically a Gaussian distribution. The encoder outputs the mean and standard deviation of the latent variables, rather than directly providing the specific values ​​of the latent variables. This means that the latent representation generated by the VAE is a probability distribution, not a deterministic value.

[0058] like Figure 1 The lower half of the diagram shows the following: In the input phase, the noised latent variable serves as the model's input, representing the latent spatial representation after noise interference. The timestep (Timestep t) marks the current temporal position in the diffusion / denoising process. In the feature preprocessing phase, the patchify operation segments the high-dimensional noised latent variable into tractable local feature patches, achieving structured processing of the spatial dimension. The embed operation embeds features at the timestep, converting the scalar timestep into a continuous vector representation with semantic information. The concatenation of N Mamba blocks involves sequential feature processing through multiple Mamba blocks; the processing details will be explained below in conjunction with the appendix. Figure 3 A detailed explanation follows. LinearReshape performs a linear transformation on the high-dimensional features output by the Mamba block, restoring the processed feature tensor to the spatial dimension of the original noisy latent variables. Finally, it outputs predicted noise, which is an estimate of the original input noise, used for subsequent denoising iterations.

[0059] After transforming the potential noise into a sequence, training it is crucial. Traditional diffusion models train the noise to gradually transform into Gaussian noise when making positive noise predictions. However, the learning difficulty is uneven across different time steps, resulting in low efficiency and error accumulation, which affects the model's convergence. To stabilize noise prediction and improve its convergence, this invention designs a Mamba Block with conditional cues for noise prediction.

[0060] The execution process of the Mamba block includes: inputting the latent spatial representation into the Mamba block after normalization and performing two-branch processing; in the first branch, using cross-order attention conditional cues and combining spatial and channel information from OCT fundus images, processing is performed through a multi-head self-attention mechanism to obtain the first branch processing result; in the second branch, the latent spatial representation is scanned sequentially in different directions to obtain the second branch processing result; and the noise prediction result is obtained by fusing the first branch processing result and the second branch processing result.

[0061] The specific process of dual-branch processing is as follows: Figure 3 As shown:

[0062] In the first branch, the input sequence (Input Tokens) is first normalized by a LayerNorm layer. Then, the normalized input sequence, along with the cross-order attention conditional cue, is input to the Scale & Shift feature adjustment module. These cuees are adjusted through scaling and shifting operations, helping the model learn more flexible feature representations. Next, the adjusted features are input to the Multi-Head Self-Attention module, which processes the data based on the multi-head self-attention mechanism to capture the relationships between different parts of the sequence. This allows the model to focus on key parts of the sequence, thus better understanding the structure of the input data. The first branch primarily utilizes conditional cues to enhance the model's understanding of image features, thereby improving the accuracy and efficiency of noise prediction.

[0063] In the second branch, the input sequence (Input Tokens) is first normalized by a LayerNorm layer. The normalized input sequence is then fed into a Scale & Shift module for scaling and shifting, which helps the model learn more flexible feature representations. Next, the adjusted features are scanned. The second branch is primarily responsible for scanning the spatial sequence of the latent representation to capture the spatial structure information of the image.

[0064] Preferably, the scan in the Mamba block is based on a bidirectional scanning mechanism and a Bow scanning method, specifically including: dynamically allocating multiple Bow scanning schemes, performing path switching across network layers through modular arithmetic; combining forward and backward bidirectional scans to generate multi-directional feature sequences; and reusing scan path memory to reduce memory usage.

[0065] It should be noted that while the traditional Mamba model can efficiently process one-dimensional (1D) sequence data, its row-column storage and row / column scanning mechanism have significant limitations. Firstly, it ignores spatial continuity. In 2D data scenarios such as images and videos, traditional row-column scanning (e.g., row-by-row or column-by-column traversal) destroys the local spatial relationships between pixels or features, making it difficult for the model to capture the interactions between adjacent elements. Secondly, it suffers from low memory efficiency. Existing improved solutions generate multiple sets of sequences through various scanning paths (e.g., serpentine, Z-shaped), requiring independent calls to the State-Space Model (SSM) for each set of sequences. This results in a linear increase in the number of random sequences and memory usage, failing to meet the demands of high-risk image processing.

[0066] In view of this, the scanning of the Mamba block in this embodiment of the invention adopts a bidirectional scanning mechanism and a Bow scanning scheme. The bidirectional scanning mechanism refers to enhancing the modeling ability of local spatial structure by alternating forward and backward scanning paths. The Bow scanning scheme refers to designing four dynamic scanning modes, which automatically switch scanning paths in different network layers through modular arithmetic to achieve balanced extraction of global and local features.

[0067] Figure 4 Four Bow scanning schemes were demonstrated using R. j Let Ω represent (where j∈[0,3]), the scanning scheme Ω of the i-th layer of Mamba. i =R {i%4} %, is the modulo operator. Specifically, when i = 0, the index is 0%4 = 0; when i = 1, the index is 1%4 = 1; when i = 2, the index is 2%4 = 2; when i = 3, the index is 3%4 = 3; and when i = 4, the index is 4%4 = 0.

[0068] Option 0: Snake-like scan from left to right, with odd-numbered rows moving forward and even-numbered rows moving backward. Red and blue arrows represent alternating forward and backward scan paths.

[0069] Option 1: Snake-like scan from right to left, odd rows in reverse and even rows in forward. Red and blue arrows represent alternating forward and backward scan paths.

[0070] Option 2: Vertical serpentine scan from top to bottom, odd columns in the forward direction and even columns in the reverse direction; the forward direction is from top to bottom. Red and blue arrows represent alternating forward and backward scan paths.

[0071] Option 3: Vertical serpentine scan from bottom to top, odd columns reversed, even columns forward; forward direction is from top to bottom. Red and blue arrows represent alternating forward and backward scan paths.

[0072] This scanning scheme enables the generation of multi-directional features in a single scan path, avoiding independent processing of multiple sequences and significantly reducing memory consumption. The bidirectional scanning mechanism reduces redundant computation and greatly improves inference speed. The Bow scanning scheme preserves the spatial continuity of the image, resulting in a significant improvement in average accuracy in tasks such as image segmentation and object detection.

[0073] Furthermore, the specific process of scanning in the second branch is as follows: Figure 3 As shown in the right half: the input features are normalized using LayerNorm, and then divided into two processing paths. The first processing path uses a one-dimensional convolutional layer (Conv1D) to extract local features from the sequence; the sequence after one-dimensional convolution is rearranged using a rearrangement operation Ω (Reaarange); then, bidirectional scanning (Forward and Backward) is performed, and after scanning, a reverse rearrangement is performed again. The first processing step, reverse rearrange, restores the original sequence, iterating through all input sequences to complete the scanning of each layer of the Mamba block. The second processing path uses a one-dimensional convolutional layer (Conv1D). The sequence after one-dimensional convolution undergoes a non-linear transformation using an activation function. This activation function introduces non-linearity, enabling the model to learn more complex feature representations. The reverse-rearranged sequence is multiplied by the activated sequence, and the result is then residual-connected with the normalized features to obtain the final scanning result of the Mamba block. Specifically, for each layer of the Mamba block, the input sequence Z... i The output Z of the bidirectional scan block i+1 The process can be represented as:

[0074] Z Ωi =recorder(Z i ,Ω i ); formula (6)

[0075]

[0076] Among them, Ω i and These represent the rearrangement and derearrangement of the input sequence at layer i, respectively. `forward` and `backward` represent the forward and backward directions from which the scan begins, respectively. The input sequence undergoes rearrangement, bidirectional scanning, and derearrangement to ensure that the input and output of each Mamba layer follow the original image sampling order.

[0077] Finally, after the Mamba module outputs the noise prediction results, these results are combined with adaptive feature conditional prompts to supervise the diffusion model in order to generate OCTA fundus images. It should be understood that the diffusion model is a generative model that generates samples by simulating a gradual process of noise addition (forward diffusion) and denoising (backward diffusion) of data. Its typical applications, besides image generation, include audio synthesis, molecular structure design, and video prediction.

[0078] It should be noted that the embodiments of the present invention use a Conditional Diffusion Model as the basic architecture. Unlike the traditional diffusion model, its reverse process can incorporate external conditional information, that is, conditional input is added on the basis of the traditional diffusion model.

[0079] The noise predicted by the Mamba block output is the completely noisy image generated by forward propagation. Adaptive feature conditional cueing is used to extract feature information from the existing OCTA image, forming an adaptive feature matrix. This extracted OCTA adaptive feature matrix is ​​used as a condition in the subsequent image generation process. Specifically, in the backpropagation of the conditional diffusion model: first, the conditional diffusion model is initialized; then, the OCTA adaptive feature matrix is ​​loaded into the backpropagation process; next, during backpropagation, the model uses the OCTA adaptive feature matrix as supervision information to guide the gradual denoising. Finally, through the backpropagation of the conditional diffusion model, an OCTA image that meets the target is generated.

[0080] It should be understood that the specific processes by which each module performs the corresponding steps described above have been detailed in the above method embodiments, and will not be repeated here for the sake of brevity. Furthermore, the module division in the embodiments of this application is illustrative and merely a logical functional division; in actual implementation, there may be other division methods. Additionally, the functional modules in the various embodiments of this application can be integrated into a single processor, exist as separate physical entities, or have two or more modules integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0081] The preceding text has provided a detailed explanation of the structure and principle of an OCTA fundus image generation system based on Mamba and a diffusion model, as provided in the embodiments of the present invention. The following text further illustrates the technical effects achievable by the present invention through experimental verification on two public datasets, OCTA500-3M and OCTA500-6M, and one private dataset.

[0082] This invention is implemented on a PyTorch-based computing platform, running on a system equipped with an RTX 4090 GPU and 24GB of memory. A series of experiments were conducted on the OCTA500_3M and OCTA500_6M datasets. During training, a conditional diffusion model was used as the base diffusion model, and the AdamW optimizer was employed. Horizontal flipping was used for data augmentation on both datasets, and a VAE with pre-trained weights was used to extract latent features from both datasets. The mean squared error between the predicted noise and the Gaussian noise of the label sampling was used as the primary loss to optimize the model, and the network architecture was trained using mini-batch standardized mean and standard deviation.

[0083] The specific experiment is as follows:

[0084] (1) Fundus datasets: A series of experiments were conducted on the OCTA-500_3M and OCTA-500_6M datasets. The experimental results are as follows: Figure 5 As shown, the subset was split into three parts based on the 3M dataset experimental parameters: a training set containing 140 data points, a validation set containing 10, and a test set containing 50. The image size was set to 304 pixels × 304 pixels, the batch size to 10, and the training run for 50,000 epochs, using the AdamW optimizer with a learning rate of 1e-4. On the 6M dataset, the subset was split into three parts: a training set containing 180 data points, a validation set containing 20, and a test set containing 100. The image size was set to 400 pixels × 400 pixels, the batch size to 8, and the training run for 80,000 epochs, using the AdamW optimizer with a learning rate of 2e-4.

[0085] (2) Evaluation Metrics: The generated OCTA images are evaluated from two aspects: the OCTA B-Scan and the OCTA mapping. To evaluate the B-Scan of each OCTA image, the mean squared error (MSE), structural similarity index (SSIM), and peak signal-to-noise ratio (PSNR) are directly used to measure image quality. Furthermore, to measure the vascular status of the generated fundus OCTA mapping, this invention designs two evaluation metrics: Vessel Proportion Error (VPE) and Vessel Similarity (VPE). VPE is mainly used to compare the absolute difference in vascular proportion between the generated OCTA mapping and the ground-based mapping. The calculation formula is as follows:

[0086]

[0087] Where M is the size of the test set, R i,G R represents the proportion of blood vessel area generated in the i-th image. i,T The proportion of the actual blood vessel area in the i-th image is expressed by the formula:

[0088]

[0089] Among them, R G R represents the average proportion of blood vessel area across all generated images. T It is the average of the proportion of blood vessel area in all real images; N G N is the number of images generated; T It represents the number of real images.

[0090] R i The ratio of the area of ​​blood vessels in a single B-Scan image to the total area of ​​the image pixels is expressed as:

[0091]

[0092] The better the generated OCTA mapping, the smaller the VPE value.

[0093] VS is used to measure the similarity between blood vessels in the generated OCTA mapping and the actual ground mapping. Its formula can be expressed as:

[0094]

[0095] Where n represents the total number of samples; G i and T i Let i represent the generated image and the ground truth map, respectively. It is the average value of blood vessel feature values ​​from all generated images. It is the average of the vascular feature values ​​of all real images. The closer the VS value is to 1, the more similar the generated OCTA mapping is to the real OCTA ground map.

[0096] (3) Generated Result Images: This invention performs visualization analysis on the generated fundus OCTA B-Scan and Mapping images. By comparing visualization examples of different OCTA B-Scan images and OCTA mapping images, the significant advantages in detail processing, image realism, and clarity of vascular structures can be clearly seen.

[0097] For actual testing, by retaining the weights from the training process, inputting an OCT B-Scan image or an OCT Mapping image, loading the retained pre-trained weights into the pre-trained network, and running the network will yield the prediction results for the corresponding OCT B-Scan or OCT Mapping image.

[0098] like Figure 6 The diagram illustrates a flowchart of an OCTA fundus image generation method based on the Mamba and diffusion model, according to an embodiment of the present invention. The method includes the following steps:

[0099] Step S61: Generate cross-order attention conditional cue and adaptive feature conditional cue; the cross-order attention conditional cue is configured to extract image features from OCT fundus images, and the adaptive feature conditional cue is configured to extract image features from OCT fundus images.

[0100] Step S62: The OCT fundus image is converted into a latent space representation by a variational autoencoder and then input into a Mamba block; the Mamba block integrates the cross-order attention conditional cue; the Mamba block generates a noise prediction result based on the cross-order attention conditional cue and the Mamba scanning method; the diffusion model is supervised and prompted by combining the noise prediction result output by the Mamba block with the adaptive feature conditional cue to generate the OCT fundus image.

[0101] It should be noted that the OCTA fundus image generation method based on Mamba and diffusion model provided in this embodiment of the invention is similar in implementation process and principle to the system described above, and will not be repeated here.

[0102] Furthermore, in the embodiments of this application, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. For example, "first XX" and "second XX" are merely to distinguish different XXs and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" do not necessarily imply that they are different. It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner. In the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any single or multiple items. For example, "at least one of a, b, or c" can be expressed as: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0103] Figure 7 This is a schematic block diagram of a computer device / equipment / system provided in an embodiment of this application. For example... Figure 7 As shown, the computer device includes at least one processor 701, a memory 702, at least one network interface 703, and a user interface 705. The various components in the device are coupled together via a bus system 704. It is understood that the bus system 704 is used to implement communication between these components. In addition to a data bus, the bus system 704 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 7 The general will label all buses as bus systems.

[0104] The user interface 705 may include a monitor, keyboard, mouse, trackball, clicker, button, touchpad, or touch screen.

[0105] It is understood that memory 702 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM) or programmable read-only memory (PROM), used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memories described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable categories of memory.

[0106] In this embodiment of the invention, the memory 702 is used to store various types of data to support the operation of the electronic terminal 700. Examples of this data include: any executable program for operation on the electronic terminal 700, such as the operating system 7021 and application program 7022; the operating system 7021 contains various system programs, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks. The application program 7022 may contain various applications, such as a media player, browser, etc., for implementing various application services. The OCTA fundus image generation method based on the Mamba and diffusion model provided in this embodiment of the invention can be included in the application program 7022.

[0107] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 701. Processor 701 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 701 or by instructions in software form. The processor 701 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 701 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. General-purpose processor 701 may be a microprocessor or any conventional processor, etc. The steps of the accessory optimization method provided in the embodiments of the present invention can be directly reflected as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in memory. The processor reads the information in the memory and combines it with its hardware to complete the steps of the aforementioned method.

[0108] In an exemplary embodiment, the electronic terminal 700 may be used by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs) to execute the aforementioned method.

[0109] According to the method provided in the embodiments of this application, this application also provides a computer program product, which includes: computer program code, which, when run on a computer, causes the computer to execute the OCTA fundus image generation method based on the Mamba and diffusion model of the above embodiments.

[0110] According to the method provided in the embodiments of this application, this application also provides a computer-readable storage medium storing program code, which, when run on a computer, causes the computer to perform the above-described method.

[0111] As used in this specification, the terms "component," "module," "system," etc., are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. As illustrated, applications running on computing devices and computing devices can both be components. One or more components may reside in a process and / or an execution thread, and components may be located on a single computer and / or distributed among two or more computers. Furthermore, these components can be executed from various computer-readable media on which various data structures are stored. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).

[0112] Those skilled in the art will recognize that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0113] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0114] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0115] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0116] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0117] In the above embodiments, the functions of each functional unit can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. A computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs, DVDs), or semiconductor media (e.g., solid-state disks, SSDs, etc.).

[0118] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0119] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0120] In summary, this application provides an OCTA fundus image generation system, method, medium, and apparatus based on Mamba and a diffusion model. By designing a unique Mamba block, this invention effectively predicts diffusion noise in accurate OCTA images from potential OCT noise, reducing artifacts in the generated images and improving the contrast of the generated OCTA images. Secondly, the proposed bidirectional Bow scanning scheme effectively alleviates the mechanistic challenges of the original Mamba scan, enhances the spatial continuity of the generated OCTA images, and improves the vascular extension of the generated images. Furthermore, the designed conditional cueing, combined with the Mamba block and diffusion model, improves the understanding of OCT images, further enhancing the realism of the generated fundus OCTA images. Experimental verification on two public datasets (OCTA500-3M and OCTA500-6M) and one private dataset achieved SSIMs of 87.5% and 88.1%, respectively, demonstrating that this invention can fully utilize fundus OCT image information to effectively generate high-quality fundus OCTA images, providing new possibilities for low-cost screening of retinal diseases. Therefore, this application effectively overcomes the various shortcomings of the prior art and has high industrial applicability.

[0121] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.

Claims

1. An OCTA fundus image generation system based on the Mamba and diffusion model, characterized in that, include: The conditional suggestion module is used to generate cross-order attention conditional suggestions and adaptive feature conditional suggestions; The cross-order attention conditional cue is configured to extract image features from OCT fundus images, and the adaptive feature conditional cue is configured to extract image features from OCT fundus images; The image generation module is equipped with a Mamba block and a diffusion model; the Mamba block integrates the cross-order attention conditional cue; the OCT fundus image is converted into a latent space representation by a variational autoencoder and then input into the Mamba block; the Mamba block generates noise prediction results based on the cross-order attention conditional cue and the Mamba scanning method; The diffusion model is supervised and prompted by combining the noise prediction results of the Mamba block output with adaptive feature conditional cueing to generate OCTA fundus images; The generation method of the cross-order attention condition cues includes: dividing several B-scans that make up the OCT fundus image into several image patches; performing max pooling and average pooling on the divided image patches respectively; generating at least two attention matrices by calculating the max pooling and average pooling results using different computational operations; and resampling the features of the image patches using the attention matrices to obtain the spatial information and channel information of the OCT fundus image.

2. The OCTA fundus image generation system based on the Mamba and diffusion model according to claim 1, characterized in that, The process of generating spatial information from the OCT fundus image includes: The first attention matrix is ​​generated by multiplying the max pooling result and the average pooling result of the image patch element by element. The attention matrix is ​​then used to resample the normalized image patch to obtain the spatial information of the fundus image.

3. The OCTA fundus image generation system based on the Mamba and diffusion model according to claim 1, characterized in that, The process of generating the channel information of the OCT fundus image includes: By element-wise summing of the max pooling and average pooling results of the image patches and mapping them through a fully connected layer to generate a second attention matrix, the channel information of the fundus image is obtained by resampling the second attention matrix with the original features.

4. The OCTA fundus image generation system based on the Mamba and diffusion model according to claim 1, characterized in that, The adaptive feature conditional cue generation process includes: encoding vascular details of OCTA fundus images into dynamic conditional signals through learnable cue blocks and embedding them into the image generation process of the diffusion model.

5. The OCTA fundus image generation system based on the Mamba and diffusion model according to claim 1, characterized in that, In the image generation module, the Mamba block generates noise prediction results based on cross-order attention conditional cues and the Mamba scanning method, which includes: inputting the latent space representation into the Mamba block after normalization and performing bi-branch processing: In the first branch, cross-order attention conditional cues are used, combined with spatial and channel information from OCT fundus images, and processed through a multi-head self-attention mechanism to obtain the processing result of the first branch; In the second branch, the latent space representation is scanned sequentially in different directions to obtain the result of the second branch processing; The noise prediction result is obtained by fusing the results of the first branch processing and the second branch processing.

6. The OCTA fundus image generation system based on the Mamba and diffusion model according to claim 5, characterized in that, The scanning methods in the second branch include: a bidirectional scanning mechanism based on alternating forward and backward scanning, and scanning by automatically switching scanning paths in different Mamba layers using preset dynamic scanning methods and modular arithmetic.

7. A method for generating OCTA fundus images based on the Mamba and diffusion models, characterized in that, include: Generate cross-order attention conditional cues and adaptive feature conditional cues; The cross-order attention conditional cue is configured to extract image features from OCT fundus images, and the adaptive feature conditional cue is configured to extract image features from OCT fundus images; The OCT fundus image is converted into a latent spatial representation by a variational autoencoder and then input into a Mamba block; the Mamba block integrates the cross-order attention conditional cue; the Mamba block generates noise prediction results based on the cross-order attention conditional cue and the Mamba scanning method; The diffusion model is supervised and prompted by combining the noise prediction results of the Mamba block output with adaptive feature conditional cueing to generate OCTA fundus images; The generation method of the cross-order attention condition cues includes: dividing several B-scans that make up the OCT fundus image into several image patches; performing max pooling and average pooling on the divided image patches respectively; generating at least two attention matrices by calculating the max pooling and average pooling results using different computational operations; and resampling the features of the image patches using the attention matrices to obtain the spatial information and channel information of the OCT fundus image.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the OCTA fundus image generation method based on the Mamba and diffusion model as described in claim 7.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the OCTA fundus image generation method based on the Mamba and diffusion model as described in claim 7.

Citation Information

Patent Citations

  • Macular edema diagnosis system, method and device based on color fundus photography, medium and program product

    CN119560133A

  • Automated detection of shadow artifacts in optical coherence tomography angiography

    US20200273218A1