Three-dimensional angiography synthesis method and system based on state space diffusion model
Through the method based on the state space diffusion model, a three-dimensional latent spatial representation of angiography is generated using a variational autoencoder and a three-dimensional vascular tree state space diffusion model, which solves the problems of fracture and topological inconsistency in three-dimensional angiography synthesis and achieves high-quality three-dimensional angiography synthesis.
Patent Information
- Application Number
- CN202510448826.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art has problems of 3D vascular fracture and topological inconsistency in three-dimensional angiography synthesis. The traditional 2D slice generation method has insufficient continuity and consistency in the three-dimensional vascular structure, which cannot effectively model the dynamic cross-layer anatomical characteristics of the vascular tree.
Using a method based on the state space diffusion model, a variational autoencoder and a three-dimensional vascular tree state space diffusion model is used to gradually add noise to generate slices by the forward diffusion module, and combined with the embedded representation of masks and markers and the cross-slice attention mechanism, a minimum spanning tree topology of vascular vascularity is constructed to realize the generation of angiography three-dimensional potential spatial representation.
The geometric continuity and anatomical consistency of multimodal vascular data were achieved, and the problems of 3D vascular fracture and topological inconsistency in traditional methods were solved, and the quality and efficiency of three-dimensional angiography synthesis were improved.
Smart Images

Figure CN120495506A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a three-dimensional angiography synthesis method and system based on a state space diffusion model. Background Art
[0002] With the increasing penetration of vascular imaging technology in clinical diagnosis and medical research, three-dimensional angiography has become the gold standard for vascular disease detection and is widely used to assist in the diagnosis and treatment of diseases such as coronary heart disease, atherosclerosis, and aneurysms. However, traditional angiography relies on iodine-containing contrast agents (such as X-ray angiography (XRA), computed tomography angiography (CTA), and magnetic resonance angiography (MRA), which may pose a risk of nephrotoxicity and ionizing radiation exposure, leading to potential carcinogenicity. Despite the availability of non-invasive imaging methods, their vascular structure contrast and resolution remain significantly limited. While deep learning-based medical image generation technologies (such as GANs, VAEs, and diffusion models) have made progress in single-modality image conversion, they face a critical bottleneck in 3D angiography synthesis. Existing methods are primarily based on 2D slice generation (e.g., SynDiff and Fast-DDPM), which results in a break in the continuity of 3D vascular topology and a significant decrease in vascular morphological consistency during cross-modality conversion. Traditional sequence modeling methods (such as Zigzag-scan and Local-scan) cannot effectively model the dynamic cross-layer anatomical characteristics of the vascular tree due to parameter redundancy and insufficient spatial perception, resulting in a loss in the spatial connectivity of the synthesized blood vessels, which is far below clinical needs.
[0003] Several technical approaches are currently being explored to address the problem of synthesizing high-quality vascular angiography with geometric structure from non-contrast 3D medical data. These include: 1) Generative Adversarial Network (GAN)-based modality conversion methods: MedGAN, for example, achieves multimodal image conversion through end-to-end mapping, but the generated vascular structure lacks continuity in 3D space and only supports conversion from CT to CTA, failing to adapt to other modalities such as MRI. 2) Diffusion-based cross-modal synthesis techniques: SynDiff generates high-fidelity medical images through an unsupervised diffusion process, but its 2D generation method fails to capture the geometric topological relationships of blood vessels, leading to broken vascular branches after 3D stacking. 3) Sequential modeling methods using fixed state space models (SSMs): These methods model long sequence dependencies through structured state spaces, but rely on a fixed scanning mechanism, making them difficult to adapt to the spatial dynamic characteristics of 3D blood vessels and unable to effectively model the branching structure and cross-slice consistency of vascular trees.
[0004] Furthermore, Wang et al.'s work, titled "Angio-Diff: Learning a Self-Supervised Adversarial Diffusion Model for Angiographic Geometry Generation," achieved unsupervised 2D angiography synthesis. However, this method is not suitable for 3D angiography synthesis, and its application to 3D angiography synthesis can lead to vascular fragmentation and topological inconsistencies. Continuing the principles of unsupervised 2D angiography synthesis to achieve high-quality and generalizable geometric synthesis of angiography in 3D volumes is a key challenge that needs to be addressed, potentially resolving technical issues 1) through 3) of the aforementioned related technical routes. Summary of the Invention
[0005] Technical problem to be solved by the present invention: In response to the above-mentioned problems of the prior art, a three-dimensional angiography synthesis method and system based on a state-space diffusion model are provided. The present invention aims to solve the problems of 3D blood vessel breakage and topological inconsistency caused by traditional 2D slice generation methods, and improve the quality and efficiency of 3D angiography synthesis.
[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: A three-dimensional angiography synthesis method based on a state-space diffusion model comprises the following steps: S1, using variational autoencoder VAE, the three-dimensional images with and without contrast vessels are converted into angiographic slices along the cross-sectional slices. and non-angiographic sections The three-dimensional latent space representation ; S2, represents the three-dimensional latent space Generating 3D latent space representation of angiography using visual encoder and 3D vascular tree state space diffusion model VasTSD , including: using visual encoders to encode non-angiographic slices The mask M and the embedded representation of the label T corresponding to the potential position of the intra-layer and inter-layer blood vessels are extracted. The three-dimensional vascular tree state space diffusion model VasTSD is composed of a forward diffusion module and a three-dimensional denoising module based on the vascular tree state space. First, the angiography slice is subjected to the forward diffusion module. Forward diffusion gradually adds noise to generate fixed-size slices, and then constructs the vascular minimum spanning tree topology through a three-dimensional denoising module based on the vascular tree state space. Combining the embedding representation of mask M and label T and the cross-slice attention mechanism, the slice is scanned and gradually denoised to finally generate a three-dimensional latent space representation of the angiography. ; S3, represents the angiography 3D latent space The final angiography 3D image is obtained by decoding using the decoder corresponding to the variational autoencoder VAE.
[0007] Optionally, in step S2, a visual encoder is used to encode non-angiographic slices. Extracting the embedding representation of the mask M and the label T corresponding to the potential position of the intra-layer and inter-layer blood vessels includes: Features are extracted through two-dimensional convolution, and then global features are extracted using average pooling and local features are extracted using maximum pooling. The extracted global features and local features are mapped to a specific embedding space through a multi-layer perceptron (MLP) to generate an embedded representation of the mask M corresponding to the potential position of blood vessels within and between layers. The embedded representation of the mask M is subjected to layer normalization to obtain the embedded representation of the label T, thereby obtaining the embedded representation of the mask M and label T corresponding to the blood vessel position and structural information.
[0008] Optionally, in step S2, the angiographic slices are processed by the forward diffusion module. When forward diffusion gradually adds noise to generate slices of fixed size, the function expression of the forward diffusion module for diffusion at any time step t is: , in, and are slices of time step t and time step t-1 respectively, is the scaling factor, is the noise sampled from a Gaussian distribution.
[0009] Optionally, in step S2, the embedded representation and slices of the mask M and the marker T are gradually denoised by a 3D denoising module based on the state space of the vascular tree, and finally a 3D latent space representation of the angiography is generated. The method includes: firstly, the embedded representation and slices of the mask M and the marker T are input into the multi-level Mamba module, and the multi-level Mamba module sequentially performs three-dimensional denoising based on the vascular tree state space using a dynamic tree topology scanning algorithm and a cross-slice attention mechanism. The final output features are passed through a linear layer and rearranged to obtain a three-dimensional latent space representation of the angiography. .
[0010] Optionally, the multi-level Mamba module includes a multi-level Mamba module connected in cascade sequence, wherein the multi-level Mamba module is composed of a plurality of first Mamba modules for extracting intra-slice features and a second Mamba module for extracting cross-slice features connected in cascade sequence, and the first Mamba module includes: The linearization layer is used to linearize the embedding representation and slice of the mask M and the tag T respectively; A deep convolutional network is used to perform a depth-wise separable convolution operation on the embedded representation of the mask M and the tag T and the embedded features obtained after the slice linearization process to extract features; SiLU activation function layer, used to extract features from deep convolutional networks using SiLU activation function to obtain features ; The vascular tree generation module is used to construct the minimum spanning tree topology of the vascular system using the Kruskal algorithm for each feature from the SiLU activation function layer, including: calculating the similarity between the elements in the input features, constructing a complete graph based on the similarity of each input feature, where each node represents an element in the input feature and the edge represents the similarity between two nodes, and using the four-connected graph G with the minimum similarity in the complete graph to generate a tree topology structure as the vascular tree in the slice; and generating an updated state transition matrix A based on the input features. - And the feature map matrix B - , where the updated state transfer matrix A - Used to describe the propagation path of the input features, the updated feature map matrix B - Used to map the input features into the vascular tree state space of the four-connected graph G corresponding to the vascular tree in the slice; The tree scanning module is used to perform two rounds of state propagation on the vascular tree in the slice. The first round is the state propagation from the leaf node to the root node of the vascular tree: the features are first propagated from the leaf node to the root node of the vascular tree, and in the propagation process, the edges of the vascular tree are traversed and the updated feature map matrix B is used. - The state of each node is updated by using the state transition matrix A- to aggregate the states of its child nodes, and the state update of each node is calculated based on the states and weights of its child nodes. The second round is the propagation from the root node to the leaf nodes of the vascular tree: the state is propagated from the root node to the leaf nodes, and during the propagation process, the information of the root node is combined with the states of its child nodes so that the state of each node in the vascular tree can reflect the global information of the input features. The state of each node in the vascular tree after two rounds of state propagation is then used as the output feature in the slice. The normalization module layer is used to normalize the features in the slice to serve as the output features of the Mamba module; The second Mamba module includes: The linearization layer is used to linearize the embedding representation and slice of the mask M and the tag T respectively; The cross-slice attention mechanism module is used to calculate the cross-slice attention coefficient using the attention mechanism based on the embedded representation of the mask M and the tag T and the embedded features obtained after the slice linearization processing and the output features from other Mamba modules; A deep convolutional network is used to perform a depth-wise separable convolution operation on the embedded representation of the mask M and the tag T and the embedded features obtained after the slice linearization process to extract features; SiLU activation function layer, used to extract features from deep convolutional networks using SiLU activation function to obtain features ; The vascular tree generation module is used to construct the vascular minimum spanning tree topology using the Kruskal algorithm for the features from the cross-slice attention mechanism module and the SiLU activation function layer, including: calculating the similarity between the elements in the input features, constructing a complete graph based on the similarity of each input feature, where each node represents an element in the input feature, and the edge represents the similarity between two nodes, and generating a four-connected graph G with a tree topology structure with the edges with the minimum similarity in the complete graph as the vascular tree within or across slices; and generating an updated state transition matrix A based on the input features. - And the feature map matrix B - , where the updated state transfer matrix A - Used to describe the propagation path of the input features, the updated feature map matrix B - Used to map input features into the vascular tree state space of the four-connected graph G corresponding to the vascular tree within or across slices; The tree scanning module is used to perform two rounds of state propagation on the vascular tree within and across slices. The first round is the state propagation from the leaf node to the root node of the vascular tree: first, the features are propagated from the leaf node to the root node of the vascular tree, and in the propagation process, the edges of the vascular tree are traversed and the updated feature map matrix B is used. - The state of each node is updated using the state transition matrix A- to aggregate the states of its child nodes, and the state update of each node is calculated based on the states and weights of its child nodes. The second round is the propagation from the root node to the leaf nodes of the vascular tree: the state is propagated from the root node to the leaf nodes, and during the propagation process, the information of the root node is combined with the states of its child nodes so that the state of each node in the vascular tree can reflect the global information of the input features. The state of each node in the vascular tree after two rounds of state propagation is then used as the feature output within and across slices. The normalization module layer is used to concatenate the features output within a slice and across slices and perform normalization operations to serve as the output features of the Mamba module.
[0011] Optionally, the calculation function expression of the cross-slice attention coefficient is: , in, For slices and The cross-slice attention coefficient, is the Leaky ReLU activation function, is a hyperparameter, 、 and Slice 、 and The embedded representation of the mask M is, and For slices Separate and slice ,slice The similarity of , and the calculation function expression of the similarity is: , in, For slices The eigenvector of For slices The eigenvector of and for and The norm of .
[0012] Optionally, the functional expression of the loss function used in the training of the three-dimensional vascular tree state space diffusion model VasTSD is: , in, is the loss function, 、 and is the weight, is the InfoNCE loss of the visual encoder, is the MSE loss of the 3D denoising module, is the tree scan loss, and: , , , in, Non-angiographic slices extracted from 3D images without angiographic vessels by variational autoencoder VAE , for The same 3D image of a non-contrast vascular is represented by a token obtained by a multi-layer perceptron MLP. is the cosine similarity, is the temperature coefficient, is the number of negative samples, As negative samples, Other non-angiographic sections from the same batch ; is the blood vessel image obtained by diffusion at time step t, is the original blood vessel image, which is a clean image without noise; Indicates the expected value, represents noise sampled from a standard Gaussian distribution, Represents the nodes in the tree scanning algorithm Gradient of the hidden state.
[0013] In addition, the present invention also provides a three-dimensional angiography synthesis system based on a state-space diffusion model, comprising a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the three-dimensional angiography synthesis method based on the state-space diffusion model.
[0014] In addition, the present invention also provides a computer-readable storage medium, which stores a computer program or instruction. The computer program or instruction is programmed or configured to execute the three-dimensional angiography synthesis method based on the state-space diffusion model through a processor.
[0015] In addition, the present invention also provides a computer program product, comprising a computer program or instructions, which are programmed or configured to execute the three-dimensional angiography synthesis method based on the state-space diffusion model through a processor.
[0016] Compared with the prior art, the present invention can achieve the following beneficial effects: the present invention represents the three-dimensional latent space Generating 3D latent space representation of angiography using visual encoder and 3D vascular tree state space diffusion model VasTSD The three-dimensional vascular tree state space diffusion model VasTSD consists of a forward diffusion module and a three-dimensional denoising module based on the vascular tree state space. Forward diffusion gradually adds noise to generate fixed-size slices, and constructs the minimum spanning tree topology of the blood vessels through a three-dimensional denoising module based on the vascular tree state space. Combined with the embedding representation of the mask M and label T extracted by the visual encoder and the cross-slice attention mechanism, the slice is scanned and gradually denoised in the vascular tree state space to finally generate a three-dimensional latent space representation of the angiography. It can uniformly encode multimodal vascular data into a geometrically continuous state space, realize anatomically consistent 3D vascular continuous synthesis, solve the problems of 3D vascular breakage and topological inconsistency caused by traditional 2D slice generation methods, and improve the quality and efficiency of 3D angiography synthesis. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Schematic diagram of the basic principle of the method in an embodiment of the present invention.
[0018] Figure 2 Schematic diagram of the network structure of the three-dimensional denoising module in an embodiment of the present invention.
[0019] Figure 3 Schematic diagram of the structure of the Mamba module in an embodiment of the present invention.
[0020] Figure 4 Schematic diagram of the principle of dynamically scanning and updating the state space in an embodiment of the present invention.
[0021] Figure 5 Schematic diagram of the principle of the cross-slice attention mechanism in an embodiment of the present invention.
[0022] Figure 6 Figure 2 shows a synthetic comparison of 2D slice angles in T1-Flash to MRA mode between the method according to the present invention and the existing method, where (a) is the label of the first sample, (b) is the result of the first sample obtained by the cGAN method, (c) is the result of the first sample obtained by Syndiff, and (d) is the result of the first sample obtained by the method according to the present invention; (e) is the label of the second sample, (f) is the result of the second sample obtained by the cGAN method, (g) is the result of the second sample obtained by Syndiff, and (h) is the result of the second sample obtained by the method according to the present invention.
[0023] Figure 7 Figure 2 shows a synthetic comparison of 2D slice angles in T1-MPRAGE to MRA conversion between the method of the present invention and the existing method, where (a) is the label of the first sample, (b) is the result of the first sample obtained by the cGAN method, (c) is the result of the first sample obtained by Syndiff, and (d) is the result of the first sample obtained by the method of the present embodiment; (e) is the label of the second sample, (f) is the result of the second sample obtained by the cGAN method, (g) is the result of the second sample obtained by Syndiff, and (h) is the result of the second sample obtained by the method of the present embodiment.
[0024] Figure 8 Figure 2 is a synthetic comparison of 2D slice angles in the T2-to-MRA mode between the method of the embodiment of the present invention and the existing method, where (a) is the label of the first sample, (b) is the result of the first sample obtained by the cGAN method, (c) is the result of the first sample obtained by Syndiff, and (d) is the result of the first sample obtained by the method of the present embodiment; (e) is the label of the second sample, (f) is the result of the second sample obtained by the cGAN method, (g) is the result of the second sample obtained by Syndiff, and (h) is the result of the second sample obtained by the method of the present embodiment.
[0025] Figure 9This is a schematic diagram of the three-dimensional angiography synthesis perspective of the embodiment method of the present invention, where (a) is the spatial structure of the blood vessel, and (b) is a schematic diagram of four perspectives: perspective 1 to perspective 4.
[0026] Figure 10 Schematic diagram of the 3D angiography synthesis effect of the method of an embodiment of the present invention, wherein (a1) and (a2) are the 3D angiography synthesis effects of view 1 and view 2 obtained by the Syndiff method; (a3) and (a4) are two 2D vascular images corresponding to the four view angles in the Syndiff method; (a5) and (a6) are the 3D angiography synthesis effects of view 3 and view 4 obtained by the Syndiff method; (b1) and (b2) are the 3D angiography synthesis effects of view 1 and view 2 obtained by the DiffMa method ; (b3) and (b4) are two two-dimensional vascular images corresponding to the four viewing angles in the DiffMa method; (b5) and (b6) are the three-dimensional angiography synthesis effects of viewing angle 3 and viewing angle 4 obtained by using the DiffMa method; (c1) and (c2) are the three-dimensional angiography synthesis effects of viewing angle 1 and viewing angle 2 obtained by using the method of this embodiment; (c3) and (c4) are two two-dimensional vascular images corresponding to the four viewing angles in the method of this embodiment; (c5) and (c6) are the three-dimensional angiography synthesis effects of viewing angle 3 and viewing angle 4 obtained by using the method of this embodiment.
[0027] Figure 11 Figure 3 is a comparison of the 3D angiography synthesis effects of the method according to the embodiment of the present invention, where (a1) to (e1) are comparison diagrams of cGAN, SynDiff, DiffMa, the method according to the embodiment, and the true effect in the axial plane; (a2) to (e2) are comparison diagrams of cGAN, SynDiff, DiffMa, the method according to the embodiment, and the true effect in the coronal plane; and (a3) to (e3) are comparison diagrams of cGAN, SynDiff, DiffMa, the method according to the embodiment, and the true effect in the sagittal plane. DETAILED DESCRIPTION
[0028] The present invention's state-space diffusion model-based 3D angiography synthesis method aims to synthesize a corresponding 3D angiographic image (angiographic 3D volume) from a non-contrast medical 3D volume of a blood vessel. This method allows a user to input a 3D image (3D volume) of a non-contrast blood vessel and automatically synthesize the corresponding 3D angiographic image (angiographic 3D volume). To facilitate a better understanding of the present invention, the following detailed description is provided in conjunction with the accompanying drawings illustrating embodiments of the present invention.
[0029] like Figure 1As shown, the three-dimensional angiography synthesis method based on the state-space diffusion model in this embodiment includes the following steps: S1, using variational autoencoder VAE, the three-dimensional images (3D volume) with and without contrast vessels are converted into angiographic slices along the cross-sectional slices. and non-angiographic sections The three-dimensional latent space representation ; S2, represents the three-dimensional latent space Generating 3D latent space representation of angiography using visual encoder and 3D vascular tree state space diffusion model VasTSD , including: using visual encoders to encode non-angiographic slices The mask M and the embedded representation of the label T corresponding to the potential position of the intra-layer and inter-layer blood vessels are extracted. The three-dimensional vascular tree state space diffusion model VasTSD is composed of a forward diffusion module and a three-dimensional denoising module based on the vascular tree state space. First, the angiography slice is subjected to the forward diffusion module. Forward diffusion gradually adds noise to generate fixed-size slices, and then constructs the vascular minimum spanning tree topology through a three-dimensional denoising module based on the vascular tree state space. Combining the embedding representation of mask M and label T and the cross-slice attention mechanism, the slice is scanned and gradually denoised to finally generate a three-dimensional latent space representation of the angiography. ; S3, represents the angiography 3D latent space Using the decoder corresponding to the variational autoencoder VAE ( Figure 1 The final angiographic 3D image (angiographic 3D volume) is decoded (omitted in the figure). It should be noted that the variational autoencoder (VAE) and its decoder are both well-known methods, so their implementation details are not detailed here. During the training phase, 3D images (3D volumes) with and without angiographic vessels are used as input for the visual encoder and the 3D vascular tree state-space diffusion model VasTSD. During the inference phase, only the 3D image without angiographic vessels is input to obtain the corresponding angiographic 3D image using the visual encoder and the 3D vascular tree state-space diffusion model VasTSD.
[0030] To automatically synthesize a corresponding angiographic 3D volume from a non-contrast-enhanced 3D image (angiographic 3D volume), this example uses a pretrained visual encoder and the 3D vascular tree state space diffusion model (VasTSD) to achieve high-quality 3D angiographic synthesis. The visual encoder extracts vascular location and structural information from non-contrast-enhanced slices and generates embedded representations of masks (M) and labels (T).
[0031] In step S1 of this embodiment, a variational autoencoder (VAE) is used to convert the three-dimensional images (3D volume) of the angiographic blood vessels and the non-angiographic blood vessels into the angiographic blood vessels along the Z-axis direction. and non-angiographic sections The three-dimensional latent space representation In step S3, the angiography three-dimensional latent space is represented When the decoder corresponding to the variational autoencoder VAE is used to decode and obtain the final three-dimensional angiography image, it is also a reverse operation in the Z-axis direction.
[0032] The visual encoder in this embodiment uses a pre-trained visual embedding mechanism that ensures the geometric consistency of blood vessels and performs continuous structure representation in angiography through directional semantic propagation. Figure 1 As shown, in step S2 of this embodiment, the non-angiography slice is imaged using a visual encoder. Extracting the embedding representation of the mask M and the label T corresponding to the potential position of the intra-layer and inter-layer blood vessels includes: Features are extracted through two-dimensional convolution, followed by global features extracted using average pooling and local features extracted using maximum pooling. The extracted global and local features are then mapped to a specific embedding space using a multi-layer perceptron (MLP) to generate embedding representations of the mask M corresponding to the potential locations of blood vessels within and between layers. The embeddings of the mask M are then normalized layer by layer to obtain embedding representations of the label T, thereby obtaining embedding representations of the mask M and label T corresponding to the location and structure of the blood vessels. The contrastive learning of the visual encoder introduces an exponential moving average mechanism to smooth the blood vessel sequence, and the InfoNCE loss is used to calculate the similarity and difference between samples.
[0033] The three-dimensional vascular tree state space diffusion model VasTSD consists of a forward diffusion module and a three-dimensional denoising module based on the vascular tree state space. The diffusion module gradually introduces noise through the forward diffusion process, and then the three-dimensional denoising process of the three-dimensional denoising module gradually removes the noise to generate high-quality angiography.
[0034] In step S2 of this embodiment, the angiography slice is processed by the forward diffusion module. When forward diffusion gradually adds noise to generate slices of fixed size, the function expression of the forward diffusion module for diffusion at any time step t is: , in, and are the patches of time step t and time step t-1 respectively, is the scaling factor, The noise is sampled from the Gaussian distribution and is gradually introduced through forward diffusion. , generating a noisy slice of a fixed size. For example, as an optional implementation, this embodiment generates a noisy slice of a size of 3×3.
[0035] During the denoising process, the three-dimensional denoising module based on the vascular tree state space uses a dynamic tree topology scanning algorithm and a cross-slice attention mechanism to ensure that the geometric structure and cross-layer interaction information of the blood vessels are effectively captured. Among them, the three-dimensional denoising module based on the vascular tree state space is used to synthesize angiography, and realize angiography synthesis with high structural fidelity from non-angiography data of multiple modalities and anatomical regions. The three-dimensional denoising module based on the vascular tree state space dynamically generates a tree based on the angiography state space method of dynamic programming to achieve continuous vascular geometry and semantic remote interaction. The three-dimensional denoising module based on the vascular tree state space uses the mean square error (MSE) loss to minimize the generated image x t The difference between it and the target noise-free image x0.
[0036] like Figure 2 As shown, in step S2 of this embodiment, the vascular minimum spanning tree topology is constructed by a 3D denoising module based on the vascular tree state space, and the vascular tree state space is scanned and gradually denoised on the slices by combining the embedded representation of the mask M and the label T and the cross-slice attention mechanism to finally generate a 3D latent space representation of the angiography. The method includes: firstly, the embedded representation and slices of the mask M and the marker T are input into the multi-level Mamba module, and the multi-level Mamba module sequentially performs three-dimensional denoising based on the vascular tree state space using a dynamic tree topology scanning algorithm and a cross-slice attention mechanism, and the final output features are passed through a linear layer (Linear) and rearranged (Rearrange) to obtain a three-dimensional latent space representation of the angiography. .
[0037] like Figure 1 As shown, the multi-level Mamba module in this embodiment includes multi-level Mamba modules connected in cascade sequence. The multi-level Mamba module is composed of a plurality of first Mamba modules for extracting intra-slice features and a second Mamba module for extracting cross-slice features, which are cascaded in sequence. The first Mamba module includes: The linearization layer is used to linearize the embedding representation and slice of the mask M and the tag T respectively; A deep convolutional network is used to perform a depth-wise separable convolution operation on the embedded representation of the mask M and the tag T and the embedded features obtained after the slice linearization process to extract features; SiLU activation function layer, used to extract features from deep convolutional networks using SiLU activation function to obtain features ; The vascular tree generation module is used to construct the minimum spanning tree topology of the vascular system using the Kruskal algorithm for the features from the SiLU activation function layer. The module calculates the similarity between the elements in the input features, constructs a complete graph based on the similarity of each input feature, where each node represents an element in the input feature and the edge represents the similarity between two nodes. The four-connected graph G with the edge-to-edge tree topology with the minimum similarity in the complete graph is used as the vascular tree in the slice, which can be expressed as: ,in, is a collection of nodes, is a set of edges; and generates an updated state transfer matrix A based on the input features - And the feature map matrix B - , where the updated state transfer matrix A - Used to describe the propagation path of the input features, the updated feature map matrix B - Used to map the input features into the vascular tree state space of the four-connected graph G corresponding to the vascular tree in the slice; The tree scanning module is used to perform two rounds of state propagation on the vascular tree in the slice. The first round is the state propagation from the leaf node to the root node of the vascular tree: the features are first propagated from the leaf node to the root node of the vascular tree, and in the propagation process, the edges of the vascular tree are traversed and the updated feature map matrix B is used. - The state of each node is updated by using the state transition matrix A- to aggregate the states of its child nodes, and the state update of each node is calculated based on the states and weights of its child nodes. The second round is the propagation from the root node to the leaf nodes of the vascular tree: the state is propagated from the root node to the leaf nodes, and during the propagation process, the information of the root node is combined with the states of its child nodes so that the state of each node in the vascular tree can reflect the global information of the input features. The state of each node in the vascular tree after two rounds of state propagation is then used as the output feature in the slice. The normalization module layer is used to normalize the features in the slice to serve as the output features of the Mamba module; like Figure 3 As shown, the second Mamba module includes: The linearization layer is used to linearize the embedding representation and slice of the mask M and the tag T respectively; The cross-slice attention mechanism module is used to calculate the cross-slice attention coefficient using the attention mechanism based on the embedded representation of the mask M and the tag T and the embedded features obtained after the slice linearization processing and the output features from other Mamba modules; A deep convolutional network is used to perform a depth-wise separable convolution operation on the embedded representation of the mask M and the tag T and the embedded features obtained after the slice linearization process to extract features; SiLU activation function layer, used to extract features from deep convolutional networks using SiLU activation function to obtain features ; The vascular tree generation module is used to construct the minimum spanning tree topology of the vascular tree using the Kruskal algorithm for the features from the cross-slice attention mechanism module and the SiLU activation function layer, including: calculating the similarity between the elements in the input features, constructing a complete graph based on the similarity of each input feature, where each node represents an element in the input feature, and the edge represents the similarity between two nodes. The edges with the minimum similarity in the complete graph are used to generate a four-connected graph G with a tree topology structure as the vascular tree within or across slices. It can also be expressed as: , in, is a collection of nodes, is a set of edges; when the Kruskal algorithm is used to construct the minimum spanning tree topology of blood vessels, the acquisition of the edge spanning tree topology structure with the minimum similarity in the complete graph includes: aggregating the edges in the complete graph into multiple sets, traversing the edge (u, v) formed by any node pair u, v in the complete graph, and if the node pair u, v is in different sets, then the edge (u, v) is added to the minimum spanning tree topology MST, and finally ends and exits when the minimum spanning tree topology MST contains |V|-1 edges, thereby obtaining the edge spanning tree topology structure with the minimum similarity in the complete graph; generating the updated state transition matrix A according to the input features - And the feature map matrix B - , where the updated state transfer matrix A - Used to describe the propagation path of the input features, the updated feature map matrix B - Used to map input features into the vascular tree state space of the four-connected graph G corresponding to the vascular tree within or across slices; The tree scanning module is used to perform two rounds of state propagation on the vascular tree within and across slices to ensure that information flows from top to bottom and bottom to top in the tree: the first round is the state propagation from the leaf node to the root node of the vascular tree: the features are first propagated from the leaf node to the root node of the vascular tree, and in the propagation process, the edges of the vascular tree are traversed and the updated feature map matrix B is used. -The state of each node is updated using the state transition matrix A- to aggregate the states of its child nodes, and the state update of each node is calculated based on the states and weights of its child nodes. The second round is the propagation from the root node to the leaf nodes of the vascular tree: the state is propagated from the root node to the leaf nodes, and during the propagation process, the information of the root node is combined with the states of its child nodes so that the state of each node in the vascular tree can reflect the global information of the input features. The state of each node in the vascular tree after two rounds of state propagation is then used as the feature output within and across slices. The normalization module layer is used to concatenate the features output within a slice and across slices and perform normalization operations to serve as the output features of the Mamba module.
[0038] As an optional implementation, Figure 1 As shown, in this embodiment, the 1st to 3rd level Mamba modules are the first Mamba modules, the 4th to 5th level Mamba modules are the second Mamba modules, and the output features of other Mamba modules inputted by the cross-slice attention mechanism module of the 4th level Mamba module are specifically the output features of the 2nd level Mamba module, and the output features of other Mamba modules inputted by the cross-slice attention mechanism module of the 5th level Mamba module are specifically the output features of the 1st level Mamba module. In addition, the output features of other Mamba modules can also be selected as needed and adjusted to the required size by upsampling or downsampling, which can also be used as the input of the cross-slice attention mechanism module in the second Mamba module.
[0039] like Figure 4 As shown, in this embodiment, the updated state transfer matrix A is generated according to the input features. - And the feature map matrix B - When including: Step 1: Pass the input features through the linear layer to generate the feature map matrix B; By characteristics For example, the corresponding feature map matrix obtained through the linear layer according to the input features can be expressed as: , in Characterized by The corresponding feature map matrix, is the weight matrix used to transform the input features from the dimension Mapped to the state space dimension N, is the input feature vector of the i-th node, is the bias term, , , The feature map matrix B controls the conversion of input features into node states to generate a preliminary representation of the node state.
[0040] Step 2: Obtain the state transition matrix A based on the mask M output by the visual encoder; The embedded representation (mask M) obtained by visual encoding, each node corresponds to the embedded representation of an input feature. State transition matrix Controls the state propagation of nodes, which determines how information is passed from parent nodes to child nodes in the tree structure. is the latitude of the hidden space.
[0041] Step 3: Randomly generate the initial output matrix C and feedback matrix D. The output matrix C and feedback matrix D are updated through the loss function during the training process. The output matrix C is the transformation matrix used to calculate the final output, while the feedback matrix D is responsible for linear transformation of the input.
[0042] Step 4: After the input features are activated by the linear layer and the activation function is used, the state transfer matrix A and the feature map matrix B are updated to obtain the updated state transfer matrix A. - And the feature map matrix B - , so that the updated feature map matrix B can be used - And the state transition matrix A- is used to update the state of each node. The update process is divided into two stages: bottom-up gradient aggregation and top-down gradient propagation: Among them, in the bottom-up aggregation process from child nodes to parent nodes.
[0043] Input feature map It can be expressed as: , in, L Indicates the length of the sequence.
[0044] For nodes , whose hidden state is , defined as: , in, For nodes The set of neighbor nodes of is the edge weight of the node. For nodes The feature map matrix, The input node Features, nodes For nodes The nodes in the neighbor node set of , and have: , in, is a hyperparameter that controls the exponential decay rate of similarity. The larger the value, the sharper the similarity distribution. is the cosine function, The input node Features, The input node Features, nodes For nodes In the tree state space model, the state of each node is propagated along the tree structure. Specifically, the state of a node is affected by the state of its parent node and child nodes. The gradient update formula for state propagation is: , in, For nodes The gradient, is a node The feature map matrix, Is the parent node To child node The state transition matrix of . The gradient of the node is passed through the gradient of the child node (that is, the node Gradient ) and the hidden state of the parent node (i.e. node The hidden state of ) is calculated as: .
[0045] in, is the loss function, is the loss function for the state transfer matrix gradient.
[0046] In the top-down propagation process from parent node to child node, the gradient is transferred from parent node to child node. and its child nodes The gradients of will be updated as follows: , , , , , , in, is the state transfer matrix after bottom-up update, Original node The state transition matrix to its child nodes, is the learning rate, which controls the parameter update step size, is the loss-to-state transfer matrix The gradient, is the updated input feature map matrix, is the loss-to-state transfer matrix The gradient, For loss Output matrix The gradient, For loss Feedback Matrix The gradient, is the output matrix, is the feedback matrix, " represents the update operation. In the process of aggregation and propagation, the function expression for calculating the scanning loss is: , in, Scanning loss.
[0047] Finally, the updated feature map matrix B is used - To update the state of each node to achieve the state of the aggregated child nodes, each input feature x will be combined with the feature map matrix B - By performing linear transformation, the state of the vascular tree state space can be updated.
[0048] In this embodiment, a cross-slice attention mechanism is introduced, combining pre-trained non-vascular image embeddings (e.g., non-vascular image features) with spatial constraints in the slice mask. The vascular region of each slice is represented by a feature vector, while the eigenvalues of the non-vascular region are lower. Specifically, the calculation function expression of the cross-slice attention coefficient in this embodiment is: , in, For slices and The cross-slice attention coefficient, is the Leaky ReLU activation function, is a hyperparameter, 、 and Slice 、 and The embedded representation of the mask M (indicating the distribution of vascular and non-vascular areas), and For slices Separate and slice ,slice The similarity of , and the calculation function expression of the similarity is: , in, For slices The eigenvector of For slices The eigenvector of and for and The norm of . Figure 5 As shown, the cross-slice attention module uses an attention mechanism to calculate the cross-slice attention coefficient based on the embedded representation of the mask M and the marker T, the embedded features obtained after slice linearization, and the output features from other Mamba modules. This allows the fusion of blood vessels within a single slice and across slices, making the 3D vascular structure more continuous and smooth. Through the cross-slice attention mechanism, the algorithm can capture the spatial continuity between slices, thereby better capturing the 3D vascular structure. This method overcomes the limitation of traditional serialization methods that process a single slice at a time and can more accurately reconstruct 3D vascular models.
[0049] In this embodiment, there is loss in visual encoding in the visual encoder, and there is also loss in the three-dimensional vascular tree state space diffusion model VasTSD. In this embodiment, the loss function used in the training of the three-dimensional vascular tree state space diffusion model VasTSD is expressed as follows: , in, is the loss function, 、 and is the weight, is the InfoNCE loss of the visual encoder, is the MSE loss of the 3D denoising module, is the tree scan loss, and: , , , in, Non-angiographic slices extracted from 3D images without angiographic vessels by variational autoencoder VAE , for The same 3D image of a non-contrast vascular is represented by a token obtained by a multi-layer perceptron MLP. is the cosine similarity, is the temperature coefficient, is the number of negative samples, As negative samples, Other non-angiographic sections from the same batch ; is the blood vessel image obtained by diffusion at time step t, is the original blood vessel image, which is a clean image without noise; Indicates the expected value, represents noise sampled from a standard Gaussian distribution, Represents the nodes in the tree scanning algorithm Gradient of the hidden state.
[0050] In order to verify the modality universality of the method in this embodiment, the existing cGAN and SynDiff methods are used as comparisons in this embodiment. The data of three different magnetic resonance imaging (MRI) modalities, T1-Flash, T1-MPRAGE, and T2, are input on the brain ITKTubeTK dataset. The results are as follows: Figure 6 、 Figure 7 and Figure 8 As shown, Figure 6 The 2D slice angle synthesis comparison between the method of this embodiment and the existing method in T1-Flash to MRA mode is shown. Figure 7 The 2D slice angle synthesis comparison between the method of this embodiment and the existing method in the T1-MPRAGE to MRA mode is shown. Figure 8 This is a comparison of the 2D slice angle synthesis of the method of this embodiment and the existing method in the T2-to-MRA mode. Figure 6 、 Figure 7 and Figure 8 It can be seen that the method of this embodiment achieved better results than the existing cGAN and SynDiff methods on data of different imaging modalities of three magnetic resonance imaging (MRI) images, T1-Flash, T1-MPRAGE, and T2, input on the brain ITKTubeTK dataset.
[0051] In order to verify the 3D angiography synthesis effect of the method of this embodiment, the following are selected in this embodiment: Figure 9 The four viewing angles shown in the figure, View 1 to View 4, are 45°, 135°, 225°, and 325°, respectively. A 3D angiography synthesis test was performed on the ITKTubeTK brain dataset, and the existing SynDiff and DiffMa methods were used for comparison. The results are shown in the figure below. Figure 10As shown, (a1) and (a2) are the three-dimensional angiography synthesis effects of view 1 and view 2 obtained by the Syndiff method; (a3) and (a4) are two two-dimensional vascular images corresponding to the four view angles in the Syndiff method; (a5) and (a6) are the three-dimensional angiography synthesis effects of view 3 and view 4 obtained by the Syndiff method; (b1) and (b2) are the three-dimensional angiography synthesis effects of view 1 and view 2 obtained by the DiffMa method; (b3) and (b4) are Two two-dimensional vascular images corresponding to the four viewing angles in the DiffMa method; (b5) and (b6) are the three-dimensional angiography synthesis effects of viewing angle 3 and viewing angle 4 obtained by using the DiffMa method; (c1) and (c2) are the three-dimensional angiography synthesis effects of viewing angle 1 and viewing angle 2 obtained by using the method of this embodiment; (c3) and (c4) are two two-dimensional vascular images corresponding to the four viewing angles in the method of this embodiment; (c5) and (c6) are the three-dimensional angiography synthesis effects of viewing angle 3 and viewing angle 4 obtained by using the method of this embodiment. Figure 10 The results show that the method of this embodiment achieves better results than the existing SynDiff and DiffMa methods at four viewing angles of 45°, 135°, 225°, and 325° in the ITKTubeTK brain dataset.
[0052] The method of this embodiment was tested on the ISICDM 2020 lung dataset with the existing cGAN, SynDiff, and DiffMa methods. Three angles in the human body coordinate system, namely the axial plane, the coronal plane, and the sagittal plane, were selected. The results are shown in the figure below. Figure 11 As shown, (a1) to (e1) are comparison diagrams of cGAN, SynDiff, DiffMa, the method of this embodiment and the true effect in the axial plane; (a2) to (e2) are comparison diagrams of cGAN, SynDiff, DiffMa, the method of this embodiment and the true effect in the coronal plane; (a3) to (e3) are comparison diagrams of cGAN, SynDiff, DiffMa, the method of this embodiment and the true effect in the sagittal plane. Figure 11 It is shown that the method of this embodiment achieves better results than the existing cGAN, SynDiff, and DiffMa methods at three angles of the axial plane, coronal plane, and sagittal plane of the ISICDM 2020 lung dataset.
[0053] In summary, the method of this embodiment has the following characteristics: (1) In response to the problem that the traditional method generates angiography based on 2D slices and lacks spatial correlation modeling across slices, resulting in the problem of breakage or discontinuity in the synthesized 3D blood vessels, which affects the accuracy of clinical diagnosis, the three-dimensional angiography synthesis method based on the state-space diffusion model of this embodiment solves the problem of 3D blood vessel structure discontinuity and improves the geometric integrity of blood vessels in the following ways: Dynamic tree state space model: The blood vessel tree topology (minimum spanning tree) is constructed through dynamic programming, and the spatial relationship modeling of the blood vessel feature map is performed using the Kruskal algorithm to retain the geometric continuity of the blood vessel branches. Cross-slice attention mechanism: The mask-guided cross-slice similarity calculation is introduced, and the feature interaction of adjacent slices is dynamically adjusted through cosine similarity and mask weight to ensure the anatomical consistency of the blood vessels in 3D space. (2) Traditional methods only convert a single modality (such as CT to CTA) and are limited by the differences in data distribution between different imaging devices, making it difficult to generalize to multimodal scenarios. This embodiment solves the problem of 3D angiography synthesis method based on the state space diffusion model by the following means: Pre-trained visual encoder: Extracts cross-modal common features of vascular position and structure through contrastive learning (InfoNCE loss), maps non-angiography data (such as MRI, CT) to a unified state space, and eliminates inter-modal differences. 3D denoising module of the conditional diffusion framework: In the reverse denoising process of the 3D denoising module, the embedded features (mask M and token T) are fused with the vascular potential representation to guide the anatomical constraints of multimodal angiography. (3) To address the problem that traditional state-space models (such as S4 / S6) use a fixed scanning mechanism, are unable to capture the long-range spatial relationship of 3D blood vessels, and have a large number of parameters and low computational efficiency, the three-dimensional angiography synthesis method based on the state-space diffusion model in this embodiment solves this problem through the following means: Tree scanning algorithm: Based on dynamic programming, a minimum spanning tree is constructed to achieve hierarchical aggregation of vascular features with linear complexity (O(L)), avoiding the high computational overhead of the fully connected graph. Lightweight architecture: The model parameter count (79.24M) is optimized through depthwise separable convolution and residual connection, which reduces the parameter count by 40% compared to Zigzag-scan-B (133.8M), thereby improving inference efficiency. (4) In order to solve the problem that traditional generative models (such as GAN and DDPM) are sensitive to noise, easily leading to blurred vascular details or artifacts, and lack explicit constraints on vascular physiological characteristics, the three-dimensional angiography synthesis method based on the state space diffusion model in this embodiment solves this problem by the following means: Vascular geometric state space: explicitly modeling vascular volume attributes in back diffusion, and strengthening the physical rationality of branch curvature and tube diameter through directional semantic propagation. Multi-objective joint optimization: combining contrast loss, denoising loss and tree scanning loss to improve the generation quality from three aspects: feature alignment, noise elimination and topological constraints.The three-dimensional angiography synthesis method based on the state-space diffusion model in this embodiment achieves high-fidelity synthesis of multimodal 3D angiography for the first time through dynamic tree state-space modeling, cross-modal visual embedding and lightweight tree scanning algorithm, solving the core defects of traditional methods in continuity, cross-modal adaptability and computational efficiency, and providing an innovative solution for non-invasive vascular imaging and low-dose angiography.
[0054] In addition, this embodiment also provides a three-dimensional angiography synthesis system based on a state-space diffusion model, including a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the three-dimensional angiography synthesis method based on the state-space diffusion model.
[0055] In addition, this embodiment also provides a computer-readable storage medium, which stores a computer program or instructions. The computer program or instructions are programmed or configured to execute the three-dimensional angiography synthesis method based on the state-space diffusion model through a processor.
[0056] In addition, this embodiment also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute the three-dimensional angiography synthesis method based on the state-space diffusion model through a processor.
[0057] Those skilled in the art should understand that the technical solution provided by the present invention may be in the form of a method, a system, or a computer program product. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the functions described in the process. Figure 1 a process or multiple processes and / or boxes Figure 1These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including the instruction device, which implements the function specified in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0058] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A three-dimensional angiography synthesis method based on a state-space diffusion model, characterized in that: The steps include: S1, using variational autoencoder VAE, the three-dimensional images with and without contrast vessels are converted into angiographic slices along the cross-sectional slices. and non-angiographic sections The three-dimensional latent space representation ; S2, represents the three-dimensional latent space Generating 3D latent space representation of angiography using visual encoder and 3D vascular tree state space diffusion model VasTSD , including: using visual encoders to encode non-angiographic slices The mask M and the embedded representation of the label T corresponding to the potential position of the intra-layer and inter-layer blood vessels are extracted. The three-dimensional vascular tree state space diffusion model VasTSD is composed of a forward diffusion module and a three-dimensional denoising module based on the vascular tree state space. First, the angiography slice is subjected to the forward diffusion module. Forward diffusion gradually adds noise to generate fixed-size slices, and then constructs the vascular minimum spanning tree topology through a three-dimensional denoising module based on the vascular tree state space. Combining the embedding representation of mask M and label T and the cross-slice attention mechanism, the slice is scanned and gradually denoised to finally generate a three-dimensional latent space representation of the angiography. ; S3, represents the angiography 3D latent space The final angiography 3D image is obtained by decoding using the decoder corresponding to the variational autoencoder VAE.
2. The three-dimensional angiography synthesis method based on the state-space diffusion model according to claim 1, characterized in that: In step S2, the non-angiography slices are encoded using a visual encoder. Extracting the embedding representation of the mask M and the label T corresponding to the potential position of the intra-layer and inter-layer blood vessels includes: Features are extracted through two-dimensional convolution, and then global features are extracted using average pooling and local features are extracted using maximum pooling. The extracted global features and local features are mapped to a specific embedding space through a multi-layer perceptron (MLP) to generate an embedded representation of the mask M corresponding to the potential position of blood vessels within and between layers. The embedded representation of the mask M is subjected to layer normalization to obtain the embedded representation of the label T, thereby obtaining the embedded representation of the mask M and label T corresponding to the blood vessel position and structural information.
3. The three-dimensional angiography synthesis method based on the state-space diffusion model according to claim 1, characterized in that: In step S2, the angiographic slices are processed by the forward diffusion module. When forward diffusion gradually adds noise to generate slices of fixed size, the function expression of the forward diffusion module for diffusion at any time step t is: , in, and are slices of time step t and time step t-1 respectively, is the scaling factor, is the noise sampled from a Gaussian distribution.
4. The three-dimensional angiography synthesis method based on the state-space diffusion model according to claim 3, characterized in that: In step S2, the vascular minimum spanning tree topology is constructed by a 3D denoising module based on the vascular tree state space. The vascular tree state space is scanned and denoised step by step on the slices by combining the embedding representation of the mask M and the label T and the cross-slice attention mechanism to finally generate a 3D latent space representation of the angiography. The method includes: firstly, the embedded representation and slices of the mask M and the marker T are input into the multi-level Mamba module, and the multi-level Mamba module sequentially performs three-dimensional denoising based on the vascular tree state space using a dynamic tree topology scanning algorithm and a cross-slice attention mechanism. The final output features are passed through a linear layer and rearranged to obtain a three-dimensional latent space representation of the angiography. .
5. The three-dimensional angiography synthesis method based on the state-space diffusion model according to claim 4, characterized in that: The multi-level Mamba module includes a multi-level Mamba module connected in cascade sequence, wherein the multi-level Mamba module is composed of a plurality of first Mamba modules for extracting intra-slice features and a second Mamba module for extracting cross-slice features connected in cascade sequence, and the first Mamba module includes: The linearization layer is used to linearize the embedding representation and slice of the mask M and the tag T respectively; A deep convolutional network is used to perform a depth-wise separable convolution operation on the embedded representation of the mask M and the tag T and the embedded features obtained after the slice linearization process to extract features; SiLU activation function layer, used to extract features from deep convolutional networks using SiLU activation function to obtain features ; The vascular tree generation module is used to construct the minimum spanning tree topology of the vascular system using the Kruskal algorithm for each feature from the SiLU activation function layer, including: calculating the similarity between the elements in the input features, constructing a complete graph based on the similarity of each input feature, where each node represents an element in the input feature and the edge represents the similarity between two nodes, and using the four-connected graph G with the minimum similarity in the complete graph to generate a tree topology structure as the vascular tree in the slice; and generating an updated state transition matrix A based on the input features. - And the feature map matrix B - , where the updated state transfer matrix A - Used to describe the propagation path of the input features, the updated feature map matrix B - Used to map the input features into the vascular tree state space of the four-connected graph G corresponding to the vascular tree in the slice; The tree scanning module is used to perform two rounds of state propagation on the vascular tree in the slice. The first round is the state propagation from the leaf node to the root node of the vascular tree: the features are first propagated from the leaf node to the root node of the vascular tree, and in the propagation process, the edges of the vascular tree are traversed and the updated feature map matrix B is used. - The state of each node is updated by using the state transition matrix A- to aggregate the states of its child nodes, and the state update of each node is calculated based on the states and weights of its child nodes. The second round is the propagation from the root node to the leaf nodes of the vascular tree: the state is propagated from the root node to the leaf nodes, and during the propagation process, the information of the root node is combined with the states of its child nodes so that the state of each node in the vascular tree can reflect the global information of the input features. The state of each node in the vascular tree after two rounds of state propagation is then used as the output feature in the slice. The normalization module layer is used to normalize the features in the slice to serve as the output features of the Mamba module; The second Mamba module includes: The linearization layer is used to linearize the embedding representation and slice of the mask M and the tag T respectively; The cross-slice attention mechanism module is used to calculate the cross-slice attention coefficient using the attention mechanism based on the embedded representation of the mask M and the tag T and the embedded features obtained after the slice linearization processing and the output features from other Mamba modules; A deep convolutional network is used to perform a depth-wise separable convolution operation on the embedded representation of the mask M and the tag T and the embedded features obtained after the slice linearization process to extract features; SiLU activation function layer, used to extract features from deep convolutional networks using SiLU activation function to obtain features ; The vascular tree generation module is used to construct the vascular minimum spanning tree topology using the Kruskal algorithm for the features from the cross-slice attention mechanism module and the SiLU activation function layer, including: calculating the similarity between the elements in the input features, constructing a complete graph based on the similarity of each input feature, where each node represents an element in the input feature, and the edge represents the similarity between two nodes, and generating a four-connected graph G with a tree topology structure with the edges with the minimum similarity in the complete graph as the vascular tree within or across slices; and generating an updated state transition matrix A based on the input features. - And the feature map matrix B - , where the updated state transfer matrix A - Used to describe the propagation path of the input features, the updated feature map matrix B - Used to map input features into the vascular tree state space of the four-connected graph G corresponding to the vascular tree within or across slices; The tree scanning module is used to perform two rounds of state propagation on the vascular tree within and across slices. The first round is the state propagation from the leaf node to the root node of the vascular tree: first, the features are propagated from the leaf node to the root node of the vascular tree, and in the propagation process, the edges of the vascular tree are traversed and the updated feature map matrix B is used. - The state of each node is updated using the state transition matrix A- to aggregate the states of its child nodes, and the state update of each node is calculated based on the states and weights of its child nodes. The second round is the propagation from the root node to the leaf nodes of the vascular tree: the state is propagated from the root node to the leaf nodes, and during the propagation process, the information of the root node is combined with the states of its child nodes so that the state of each node in the vascular tree can reflect the global information of the input features. The state of each node in the vascular tree after two rounds of state propagation is then used as the feature output within and across slices. The normalization module layer is used to concatenate the features output within a slice and across slices and perform normalization operations to serve as the output features of the Mamba module.
6. The three-dimensional angiography synthesis method based on the state-space diffusion model according to claim 5, characterized in that: The calculation function expression of the cross-slice attention coefficient is: , in, For slices and The cross-slice attention coefficient, is the Leaky ReLU activation function, is a hyperparameter, 、 and Slice 、 and The embedded representation of the mask M is, and For slices Separate and slice j ,slice k The similarity of , and the calculation function expression of the similarity is: , in, For slices The eigenvector of For slices The eigenvector of and for and The norm of .
7. The three-dimensional angiography synthesis method based on the state-space diffusion model according to claim 5, characterized in that: The functional expression of the loss function used in the training of the three-dimensional vascular tree state space diffusion model VasTSD is: , in, is the loss function, 、 and is the weight, is the InfoNCE loss of the visual encoder, is the MSE loss of the 3D denoising module, is the tree scan loss, and: , , , in, Non-angiographic slices extracted from 3D images without angiographic vessels by variational autoencoder VAE , for The same 3D image of a non-contrast vascular is represented by a token obtained by a multi-layer perceptron MLP. is the cosine similarity, is the temperature coefficient, is the number of negative samples, As negative samples, Other non-angiographic sections from the same batch ; is the blood vessel image obtained by diffusion at time step t, is the original blood vessel image, which is a clean image without noise; Indicates the expected value, represents noise sampled from a standard Gaussian distribution, Represents the nodes in the tree scanning algorithm Gradient of the hidden state.
8. A three-dimensional angiography synthesis system based on a state-space diffusion model, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the three-dimensional angiography synthesis method based on the state-space diffusion model according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute the three-dimensional angiography synthesis method based on the state-space diffusion model according to any one of claims 1 to 7 through a processor.
10. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute the three-dimensional angiography synthesis method based on the state-space diffusion model according to any one of claims 1 to 7 through a processor.