Single-view projection fast fluorescence tomography reconstruction system based on double-path condition guide diffusion model
The single-view projection rapid fluorescence tomography reconstruction system based on a dual-path condition-guided diffusion model solves the pathological problem caused by single-view data sparsity, achieves high-quality fluorescence tomography reconstruction, improves the spatial resolution and robustness of imaging, and is suitable for in vivo biomedical imaging.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIHANG UNIV
- Filing Date
- 2026-02-02
- Publication Date
- 2026-05-12
AI Technical Summary
Existing single-view fluorescence tomography technology suffers from severe pathological problems under extremely sparse projection data conditions, resulting in poor image quality and insufficient spatial resolution, making it difficult to achieve rapid dynamic three-dimensional imaging of living organisms.
A fast fluorescence tomography reconstruction system based on a dual-path condition-guided diffusion model is adopted. The system acquires raw data through a view acquisition module, constructs a conditional model and a diffusion model, and combines a dual-path condition-guided mechanism to extract features using lightweight U-Net and ResNet structures. It then performs iterative denoising and reconstruction, and optimizes the model to achieve high-quality reconstruction.
This technology enables rapid, dynamic, in vivo three-dimensional imaging of target molecules, improving imaging quality, enhancing spatial resolution and robustness, mitigating pathological issues, and obtaining highly generalizable and high-quality fluorescence tomography reconstruction images.
Smart Images

Figure CN122023596A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of optical molecular imaging technology, and in particular to a single-view projection fast fluorescence tomography reconstruction system based on a dual-path condition-guided diffusion model. Background Technology
[0002] Fluorescence molecular tomography (FMT) is an important non-invasive three-dimensional (3D) optical imaging technique. By measuring the fluorescence signal on the surface of a living organism and reconstructing the three-dimensional distribution and concentration of fluorescent probes within the body, it enables quantitative monitoring and visualization of molecular activity in living animal tissues. Compared to structural imaging techniques such as CT, MRI, and photoacoustic imaging, and functional imaging techniques such as PET, FMT offers advantages such as being radiation-free, low-cost, highly specific, and highly sensitive. In recent years, FMT has been widely applied in various biomedical fields, including cardiovascular and cerebrovascular diseases, tumor detection and evaluation, drug development, and intraoperative navigation.
[0003] Because of the need to capture rapid biomolecular dynamics, finite-projection imaging (FMT) must operate at extremely high speeds, typically requiring the acquisition of projected images from only a single angle to minimize image acquisition time. However, trading extremely sparse projected data for temporal resolution exacerbates the underdeterminacy and ill-conditioned nature of the inverse problem. Although emerging second near-infrared (NIR-II) based FMT frameworks (1000-1700 nm) have mitigated ill-conditionedness to some extent by reducing scattering, scattering effects still dominate and hinder fundamental improvements in image quality. To address the severe ill-conditioned nature of finite-projection FMT, researchers have developed two main reconstruction strategies. The first is the traditional iterative regularization method, which addresses the inverse problem by introducing regularization terms or sparse priors. While these methods are theoretically rigorous, they suffer from high computational costs, poor condition numbers, the need for manual parameter tuning, and the requirement of ≥3 projections. Furthermore, their performance is limited by the inherently severe ill-conditioned nature of the inverse problem, resulting in generally poor reconstruction fidelity and stability. The second approach is deep learning, which directly learns the mapping from measurement data to fluorescence distribution through neural networks, such as MAP-PGAN and SVRNet. These methods better conform to optical nonlinearity and mitigate ill-conditioned phenomena, significantly improving the positioning accuracy and shape recovery capability of finite / single-view FMTs. However, current techniques still suffer from drawbacks such as insufficient spatial resolution (≥2 mm) and limited generalization ability. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a single-view projection rapid fluorescence tomography reconstruction system based on a dual-path condition-guided diffusion model. This system solves the pathological problem exacerbated by the extremely sparse single-view data and overlapping depth information, enabling rapid dynamic in vivo three-dimensional imaging of target molecules and achieving high-quality reconstruction.
[0005] To achieve the above objectives, the present invention provides the following solution: a single-view projection fast fluorescence tomography reconstruction system based on a dual-path condition-guided diffusion model, comprising: The view acquisition module is used to acquire the raw dataset consisting of fluorescence three-dimensional distribution images and single-view fluorescence projection measurement images, and to acquire single-view fluorescence projection data in real experiments. The model building module is used to define a conditional model framework for extracting conditional features from boundary fluorescence projection measurements, a diffusion model framework for iterative denoising to generate a three-dimensional fluorescence distribution, and a dual-path conditional guidance mechanism for connecting the two model frameworks, so as to obtain an initial single-view projection reconstruction model. The model optimization module is used to train, update, and set and optimize hyperparameters for the initial single-view projection reconstruction model using the original dataset to obtain a single-view projection reconstruction model. The single-view projection reconstruction model is then validated and evaluated using a test set.
[0006] Optionally, the view acquisition module includes: The dataset unit is used to calculate fluorescence projection image data using a forward model based on the simulated fluorescence target set by the computer, so as to obtain an original dataset consisting of a three-dimensional fluorescence distribution image and a single-view fluorescence projection measurement image. The original dataset is then divided into a training set and a validation set at a ratio of 9:1. The single-view fluorescence projection unit is used to generate test sets for experiments through computer simulation. For experiments with real physical phantoms or live mice, it uses CT images to obtain real fluorescence tomographic three-dimensional images of the test sets and uses a fluorescence camera to acquire single-view fluorescence projection data.
[0007] Optionally, the model building module includes: The feature extraction unit is used to combine the conditional 2D encoding module, the first 3D decoding module, and the skip connections of the tomographic lifting representation to obtain the conditional model framework. Based on the conditional model framework, a lightweight 4-layer U-Net is used to extract features from the measured fluorescence projection data to obtain the conditional model. The basic unit of each encoding module or decoding module is set as a ResNet structure, and the ResNet structure is used to obtain local and global conditional features. An iterative denoising unit is used to combine the 3D encoding module, the second 3D decoding module, and jump-connection to obtain a diffusion model framework. Based on the diffusion model framework, a lightweight 4-layer U-Net is adopted, and time-step embedding is integrated into the ResNet unit of each encoding module or decoding module to iteratively denoise and reconstruct the fluorescence three-dimensional distribution to obtain the diffusion model. The image reconstruction unit is used to combine local guidance based on multi-feature fusion and global guidance based on difference-common transformation to obtain a dual-path conditional guidance mechanism. The local and global conditional features extracted by the conditional model are transmitted to the diffusion model. Under conditional guidance, the diffusion model generates a fluorescent molecular tomography reconstruction three-dimensional image that conforms to physical laws to obtain an initial single-view projection reconstruction model.
[0008] Optionally, the fault-lifting representation mechanism includes a first stage and a second stage; The first stage includes utilizing depthwise separable convolutions from a single view Figure 2 Cross-channel spatial features are extracted from the 3D projection, and channel transformation is performed based on the feature depth to achieve feature extraction and preliminary enhancement; wherein, the depthwise separable convolution consists of depthwise convolution and pointwise convolution; The second stage includes compressing the boosted features into channel statistics using global average pooling, converting the channel statistics into activation factors for scaling the features using a fully connected layer and Sigmoid, and then performing shape transformation on the features to obtain the final boosted volumetric feature representation.
[0009] Optionally, local guidance based on multi-feature fusion is used to utilize the fusion of spatial and channel dual branches. During the initialization stage of the diffusion model, multi-feature fusion is performed on shallow noise features, local condition features, and time step embedding features to provide fine and time-dependent local guidance for diffusion generation. Global guidance based on difference-common transformation is used to combine deep noise features with global condition features located in the middle layer of the conditional model through a difference information injection module and a pair of common information injection modules, so as to ensure the rationality of global structure reconstruction.
[0010] Optionally, the model optimization module includes: The tuning unit is used to train and update the initial single-view projection reconstruction model using the training set. During the training process, the model loss is calculated using the validation set and the loss function, and hyperparameters are set and optimized using the Adam optimizer and linear decay of the learning rate to complete the model tuning and obtain the single-view projection reconstruction model. The loss function includes a guiding term, a physical consistency term, and a diffusion prior term. The verification unit is used to verify and evaluate the single-view projection reconstruction model using a test set.
[0011] This invention discloses the following technical effects by providing a single-view projection fast fluorescence tomography reconstruction system based on a dual-path condition-guided diffusion model: 1. This invention not only fundamentally abandons the reconstruction approach based on iterative solutions, but also surpasses the reconstruction scheme based on deep learning end-to-end models. This generative reconstruction system based on a denoising diffusion probability model interferes with the real internal fluorescence distribution data through a forward diffusion process, and trains a neural network that incorporates the conditional features of fluorescence boundary measurement signals to learn inverse transformation, thereby establishing an accurate mapping between internal fluorescence and boundary fluorescence and achieving high-quality reconstruction.
[0012] 2. The fault-lifting representation mechanism can enhance the sparse two-dimensional features of single-view projection into an enhanced three-dimensional prior, overcoming the problem that the projection data is extremely sparse and each data point contains different 3D depth overlap information. It enhances the adaptive perception of potential spatial information and provides a highly adaptable three-dimensional representation for subsequent condition-guided diffusion, thus greatly alleviating the severe ill-conditioning of single-view reconstruction.
[0013] 3. The dual-path conditional guidance mechanism, through the synergistic effect of local and global guidance, not only enhances the model's ability to perceive fine details and time, but also bridges the domain differences between noise and conditional features to strengthen global relevance, ultimately obtaining fluorescence tomography 3D reconstruction images with ultra-high spatial resolution, high robustness, and high generalization.
[0014] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram of the system architecture provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of fluorescence projection image acquisition and dataset generation provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of a dual-path condition-guided diffusion model provided in an embodiment of the present invention; Figure 4 A schematic diagram of the fault lifting representation mechanism provided in an embodiment of the present invention; Figure 5 A schematic diagram of a local guidance module based on multi-feature fusion provided in an embodiment of the present invention; Figure 6 A schematic diagram of a global guidance module based on difference and common conversion provided in an embodiment of the present invention; Figure 7 The flowchart of the dual-path conditional guided diffusion model provided in the embodiments of the present invention is shown. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0019] like Figure 1 As shown, this invention provides a single-view projection fast fluorescence tomography reconstruction system based on a dual-path condition-guided diffusion model, comprising: 1. View acquisition module, used to generate a dataset consisting of a three-dimensional fluorescence distribution image and a single-view fluorescence projection measurement image, and to acquire single-view fluorescence projection data in real experiments; including: Dataset Unit: Used for acquiring fluorescence projection images, obtaining test set input data for single-view reconstruction; generating a dataset consisting of pairs of real fluorescence tomographic 3D images and fluorescence projection images.
[0020] The single-view fluorescence projection unit is used to generate a dataset for network training through computer simulation, and then divide it into training and validation sets at a ratio of 9:1; the dataset is scientifically and accurately preprocessed so that the model can accurately quantify targets with different fluorescence concentrations.
[0021] The specific details of the view acquisition module are as follows: Fluorescent projection image acquisition is used to obtain the test set input data for single-view reconstruction. A schematic diagram of fluorescent projection image acquisition and dataset generation is shown below. Figure 2 As shown.
[0022] In the computer numerical simulation experiment, a simulated fluorescent target was set up by the computer, and the fluorescence projection image data was obtained through forward model calculation. The Kirchhoff approximation (KA) based on the diffusion equation was used to establish the forward photon propagation model, which can efficiently achieve accurate mapping between boundary measurement signals and internal light sources. To simulate the shape of a mouse, a semi-cylindrical model (1.5 cm high, 3.0 cm in diameter) was constructed, with the axis 0.2 cm from the bottom surface. The fluorescent target was excited by a point light source located at (0.2, 0, 0.75) cm. The detector FOV was 160º. The detector sampling interval was 5°, and the grid resolution was 1.0 × 1.0 × 1.0 mm. 3 .
[0023] In live mouse experiments, it is necessary to acquire projection images of fluorescent targets on the mouse body surface using a fluorescence camera. The specific method is as follows: an 808 nm linear laser is used to excite an ICG fluorescent probe, and a 1050 nm long-pass emission filter is placed in front of the camera to filter out the excitation light. During data acquisition, the mouse lies prone on the scanning platform, and its abdomen is irradiated with laser light from different positions to acquire multiple fluorescence images. The purpose is to select a single fluorescence projection image with the best signal-to-noise ratio and image quality. Then, X-ray images are acquired for CT reconstruction, which is used to extract the mouse contour to establish a KA forward model and also to verify the anatomical structure of the fluorescent target later. Subsequently, the single-view fluorescence images acquired by the camera are mapped in two-dimensional and three-dimensional space to obtain projection data, completing the preprocessing before inputting into the reconstruction model.
[0024] The generated dataset consists of pairs of real fluorescence tomographic 3D images and fluorescence projection images. The dataset for network training and optimization was generated through computer simulation and then divided into training and validation sets in a 9:1 ratio. The training set was used to train the model, and the validation set was used for hyperparameter tuning and determining the optimal parameter model. The test set for the simulation experiments was generated computer-generated; for live mouse experiments, real fluorescence tomographic 3D images of the test set were obtained from CT images, and single-view fluorescence projection data were acquired using a fluorescence camera.
[0025] The training and validation sets were generated by computer simulation. First, a photon propagation model was established based on the Kirchhoff approximation (KA) of the diffusion equation to achieve accurate mapping between the surface signal and the internal source. Then, a large number of columnar fluorescent targets with different shapes (diameter: 2–5 mm), positions (X-axis: -1.2 cm to 0 cm; Y-axis: -1.2 cm to 1.2 cm), and fluorescence concentrations (1–10) were randomly generated, with a fixed height of 5 mm. The edge-to-edge distances (EEDs) of the two targets were randomly varied, with a minimum EED set to 0.3 mm. The light source-detector configuration of the data was set to match the simulation test data or real experimental data. A dataset containing 10,000 samples was generated, of which 1,000 sets were randomly selected for validation. Each sample includes a two-dimensional image of the boundary fluorescence projection. with vector and corresponding three-dimensional images of fluorescence distribution. x The obtained independent test set is used to finally evaluate the optimal model.
[0026] Scientific and precise preprocessing of the dataset enables the model to accurately quantify targets with different fluorescence concentrations. For each sample of the fluorescence projection data, the mean and standard deviation calculated from the total sample are standardized. For the three-dimensional image data of fluorescence distribution, the minimum and maximum values of the total sample are calculated first, and then each sample is normalized to the range [-1, 1] relative to the total sample value, so as to better incorporate the forward and inverse processes of the diffusion model.
[0027] 2. The model building module defines a conditional model framework for extracting conditional features from boundary fluorescence projection measurements, a diffusion model framework for iterative denoising to generate a 3D fluorescence distribution, and a dual-path conditional guidance mechanism for connecting the two model frameworks, resulting in an initial single-view projection reconstruction model. The model building module includes: The feature extraction unit is used to combine the conditional 2D encoding module, the first 3D decoding module, and the skip connections of the tomographic lifting representation to obtain the conditional model framework. Based on the conditional model framework, a lightweight 4-layer U-Net is used to extract features from the measured fluorescence projection data to obtain the conditional model. The basic unit of each encoding module or decoding module is set as a ResNet structure, and the ResNet structure is used to obtain local and global conditional features.
[0028] An iterative denoising unit is used to perform forward denoising, conditionally guided diffusion denoising model construction, and posterior probability and reverse denoising sampling on a real three-dimensional fluorescence image based on the diffusion model framework, to obtain a reconstructed three-dimensional image of the fluorescence target; specifically including: Based on the aforementioned diffusion model framework, Gaussian noise with a scheduling parameter that decreases monotonically over time is gradually added to the real three-dimensional fluorescence image to obtain a pure noise image that conforms to a Gaussian distribution, thus completing the forward noise addition. By using the boundary fluorescence projection image and time step as conditional inputs and combining them with the objective function, the mapping from boundary fluorescence to internal fluorescence is learned during the reverse denoising process, thus constructing a conditionally guided diffusion denoising network. Using the condition-guided diffusion denoising network, the pure noise image is subjected to stepwise reverse iterative denoising to obtain the reconstructed three-dimensional image of the fluorescent target.
[0029] The image reconstruction unit is used to perform local guidance based on multi-feature fusion and global guidance based on the dual-path conditional guidance mechanism, using the pure noise 3D image and the conditional features, to output accurate fluorescence molecular tomography and complete the construction of the initial projection reconstruction model; specifically including: The conditional model extracts local and global conditional features from the measured fluorescence projection data; the diffusion model integrates time-step embedding to iteratively denoise and reconstruct the three-dimensional fluorescence distribution. The tomographic enhancement representation mechanism enhances the fluorescence projection features into a volumetric representation through two stages: depthwise separable convolution and a compression excitation module, so as to adaptively guide the diffusion model for three-dimensional reconstruction. The dual-path conditional guidance mechanism provides guidance information with accurate local details and reasonable global structure for diffusion generation through local guidance based on multi-feature fusion and global guidance based on difference-common transformation.
[0030] Specifically, the model building module includes: The dual-path conditionally guided diffusion model aims to introduce a conditionally guided three-dimensional diffusion model that accurately establishes the nonlinear mapping relationship between surface fluorescence measurement and internal fluorescence source as prior knowledge into tomographic reconstruction.
[0031] The process of extracting prior knowledge based on the diffusion model begins with forward diffusion, which involves progressively adding Gaussian noise to the initial 3D fluorescence image to ultimately generate a pure noise image conforming to a Gaussian distribution N(0,I). Each step of the forward noise addition can be represented as: (1) Among them, noise scheduling parameters The noise decreases monotonically over time, causing the variance of the noise added at each step to gradually increase until... x T It becomes pure noise, and T represents the total number of steps in the diffusion process.
[0032] Subsequently, a conditionally guided diffusion model was constructed as a denoising network, incorporating the mapping relationship between surface fluorescence signals and internal fluorescence sources into the denoising mechanism to learn prior knowledge. The model parameters were optimized using the following objective function: (2) in, Indicates by A parameterized neural network whose conditional inputs are time step t and fluorescence projection image. , A three-dimensional image representing the true fluorescence distribution.
[0033] After constructing the reconstruction diffusion model, the posterior probability is calculated based on Bayes' theorem. The mean and standard deviation of this distribution are given by the following formula: (3) in, The output of the diffusion model is shown; all other variables are known. A three-dimensional image of the fluorescent target can be obtained through the following stepwise inverse denoising iterative process. Precise sampling: (4) symbol This indicates the sampling process in reverse. When the iteration converges to t=1, the final reconstruction result of the FMT inverse problem is obtained. .
[0034] The dual-path conditionally guided diffusion model mainly consists of a conditional model, a diffusion model, and a dual-path conditionally guided mechanism connecting the two, such as... Figure 3 As shown.
[0035] The conditional model employs a lightweight 4-layer U-Net structure to extract features from the fluorescence projection data at the measurement boundary, consisting of a 2D encoding module, a 3D decoding module, and skip connections for tomographic lifting representation. The basic unit of each encoding or decoding module is a ResNet structure. For each ResNet unit, the fluorescence projection input undergoes two rounds of group normalization, SiLU activation, and convolution, and is then added to the skip connections of the input. Downsampling and upsampling are achieved through average pooling and nearest-neighbor interpolation, respectively.
[0036] The diffusion model employs a U-Net structure similar to the conditional model, but differs in that it uses direct concatenation for skip connections. Furthermore, the ResNet units in both the encoding and decoding modules are 3D, and time-step embedding is integrated for iterative denoising and reconstruction of the fluorescence 3D distribution. The ResNet unit embedding the time step first undergoes group normalization, SiLU activation, and convolution for the noise feature input, followed by SiLU activation and a linear layer for the time-step feature input. The two are then added together, followed by group normalization and SiLU activation, and finally added to the skip connections of the noise feature input.
[0037] In the aforementioned fault lift representation mechanism (e.g.) Figure 4 As shown), firstly, Depthwise Separable Convolution (DSConv) is employed, incorporating both depthwise convolution and pointwise convolution, to efficiently extract cross-channel spatial features from a single-view 2D projection. When the volumetric feature depth is D, the channel size transforms from C to... .
[0038] Then, the "squeeze-excite" module adaptively re-adjusts these boosted features, emphasizing those channels more important for representing the 3D structure while suppressing noise. The boosted features are then processed using global average pooling. Compressed into channel statistics : (5) Then, the activation factor is generated through a fully connected layer and Sigmoid activation. , for features Scaling. Finally, the scaled features are shape-transformed according to depth D to obtain a volumetric feature representation. This two-stage process ensures that the improved 3D volumetric prior model is both computationally efficient and information-rich, providing a solid foundation for subsequent dual-path conditional guidance.
[0039] The local guidance mechanism based on multi-feature fusion (such as...) Figure 5 As shown in the diagram, a dual-attention mechanism (channel-space) is first used to fuse shallow noise and local conditional features, providing precise and accurate guidance for accurate error correction in the diffusion model, thereby accelerating convergence and improving detail fidelity. Channel and spatial branches are obtained by global pooling across spatial and channel dimensions, respectively, to aggregate spatial and channel information. This is achieved through convolution and... Generate channel weights and and spatial weights and The subscripts 'c' and 's' denote channels or spaces. Then, the weights of the noise and conditions are multiplied by the features and summed to effectively fuse the features, expressed as: (6) Where X and C represent noise characteristics and conditional characteristics, respectively.
[0040] Subsequently, a local guidance mechanism based on multi-feature fusion will incorporate time-step features. With features The mixture is then fused and added to the noise feature X in a skip pattern, as shown in the following equation: (7) in, , Indicates group normalization, Represents the SiLU activation function. It is a 3D convolution. X LGMF This is the output of a local guidance module based on multi-feature fusion. The fusion of time-step embeddings ensures that the conditional guidance information remains dynamically consistent with each denoising stage.
[0041] The global guidance mechanism based on difference-common transformation (such as...) Figure 5 The module (shown) consists of a Differential Information Injection Module (DIIM) and a pair of Common Information Injection Modules (CIIM), which combine deep noise features with global conditional features of the intermediate layer of the conditional model. Its formal expression is as follows: (8) in, Indicates the characteristics of deep noise. This is a global conditional feature.
[0042] DIIM first addresses the domain difference between noise and conditional features. It maps noisy features to a query Q and conditional features to keys K and values V using a linear layer. Then, it calculates the similarity matrix between Q and K using dot-product attention and multiplies it by V to obtain common information. Subsequently, the difference information is obtained by removing common information: (9) Finally, this difference information is injected into the noisy feature Q to obtain complementary information: (10) in, Representation layer normalization, It is a multilayer perceptron.
[0043] Subsequently, a pair of CIIMs alternately extract common features from the global dependencies of noise and global conditional features. This differs from DIIM in that it omits the common information elimination calculation step in equation (9). Therefore, the global guidance mechanism based on difference-common transformation effectively promotes the reconstruction process towards a globally consistent 3D topology.
[0044] 3. A model optimization module, used to train and update the initial reconstruction model using the generated dataset, and to set and optimize hyperparameters to obtain a single-view projection reconstruction model. This model is then validated and evaluated using a test set. Specifically, it includes: The tuning unit is used to train and update the initial single-view projection reconstruction model using the training set. During the training process, the model loss is calculated using the validation set and the loss function, and hyperparameters are set and optimized using the Adam optimizer and linear decay of the learning rate to complete the model tuning and obtain the single-view projection reconstruction model. The loss function includes a guiding term, a physical consistency term, and a diffusion prior term. The verification unit is used to verify and evaluate the single-view projection reconstruction model using a test set.
[0045] Specifically, the model optimization module includes: The constructed single-view projection reconstruction model based on dual-path conditional guided diffusion is trained. The model is trained and its parameters optimized using the training and validation sets. The flowchart of the final dual-path conditional guided diffusion model algorithm is shown below. Figure 7 As shown.
[0046] To enable the conditional model to extract high-quality guiding features and ensure accurate reconstruction by the diffusion model, a three-anchor diffusion loss was designed to train the model. This loss function includes a guiding term, a physical consistency term, and a prior term: (11) in, The network represents the conditional model. The fluorescence distribution vector is... The noise features from the previous layer of the diffusion model are obtained through convolution and interpolation transformation. Then, the known fluorescence projection vector is used... and system matrix Calculate physical consistency. N and M represent the number of internal grid points and detection points, respectively. Describe the features of the three-layer decoding hidden layer of the diffusion prior term in the diffusion model. Deep supervised constraints are employed to improve positioning accuracy and robustness while mitigating gradient problems.
[0047] The Adam optimizer is used to update network parameters, with the learning rate decaying linearly to improve training stability. After optimization on the validation set, the initial learning rate is set to... The training batch size was 4, and a total of 60 training rounds were performed. Finally, the optimal parameter model was determined and saved by calculating the model loss on the validation set.
[0048] The saved optimal parameter model was used for system testing and fluorescence tomography 3D image reconstruction. During the testing phase, fluorescence projection images generated through numerical simulation, or fluorescence projection images actually acquired using a fluorescence camera in live mouse experiments, were input into the saved optimal network model. The model could then quickly output the reconstructed 3D spatial distribution of the fluorescent target. Finally, in live experiments, the anatomical location of the fluorescent target glass tube obtained from CT images was used to verify and evaluate imaging quality indicators such as the accuracy of localization and spatial resolution.
[0049] Therefore, this invention provides a single-view projection rapid fluorescence tomography reconstruction system based on a dual-path condition-guided diffusion model, which solves the pathological problem exacerbated by the extremely sparse single-view data and overlapping depth information, enabling rapid dynamic in vivo three-dimensional imaging of target molecules and achieving high-quality reconstruction.
[0050] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0051] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A single-view projection rapid fluorescence tomography reconstruction system based on a dual-path condition-guided diffusion model, characterized in that, include: The view acquisition module is used to acquire the raw dataset consisting of fluorescence three-dimensional distribution images and single-view fluorescence projection measurement images, and to acquire single-view fluorescence projection data in real experiments. The model building module is used to define a conditional model framework for extracting conditional features from boundary fluorescence projection measurements, a diffusion model framework for iterative denoising to generate a three-dimensional fluorescence distribution, and a dual-path conditional guidance mechanism for connecting the two model frameworks, so as to obtain an initial single-view projection reconstruction model. The model optimization module is used to train, update, and set and optimize hyperparameters for the initial single-view projection reconstruction model using the original dataset to obtain a single-view projection reconstruction model. The single-view projection reconstruction model is then validated and evaluated using a test set.
2. The single-view projection fast fluorescence tomography reconstruction system based on a dual-path condition-guided diffusion model according to claim 1, characterized in that, The view acquisition module includes: The dataset unit is used to calculate fluorescence projection image data using a forward model based on the simulated fluorescence target set by the computer, so as to obtain an original dataset consisting of a three-dimensional fluorescence distribution image and a single-view fluorescence projection measurement image. The original dataset is then divided into a training set and a validation set at a ratio of 9:
1. The single-view fluorescence projection unit is used to generate test sets for experiments through computer simulation. For experiments with real physical phantoms or live mice, it uses CT images to obtain real fluorescence tomographic three-dimensional images of the test sets and uses a fluorescence camera to acquire single-view fluorescence projection data.
3. The single-view projection fast fluorescence tomography reconstruction system based on a dual-path condition-guided diffusion model according to claim 1, characterized in that, The model building module includes: The feature extraction unit is used to combine the conditional 2D encoding module, the first 3D decoding module, and the skip connections of the tomographic lifting representation to obtain the conditional model framework. Based on the conditional model framework, a lightweight 4-layer U-Net is used to extract features from the measured fluorescence projection data to obtain the conditional model. The basic unit of each encoding module or decoding module is set as a ResNet structure, and the ResNet structure is used to obtain local and global conditional features. An iterative denoising unit is used to combine the 3D encoding module, the second 3D decoding module, and jump-connection to obtain a diffusion model framework. Based on the diffusion model framework, a lightweight 4-layer U-Net is adopted, and time-step embedding is integrated into the ResNet unit of each encoding module or decoding module to iteratively denoise and reconstruct the fluorescence three-dimensional distribution to obtain the diffusion model. The image reconstruction unit is used to combine local guidance based on multi-feature fusion and global guidance based on difference-common transformation to obtain a dual-path conditional guidance mechanism. The local and global conditional features extracted by the conditional model are transmitted to the diffusion model. Under conditional guidance, the diffusion model generates a fluorescent molecular tomography reconstruction three-dimensional image that conforms to physical laws to obtain an initial single-view projection reconstruction model.
4. The single-view projection fast fluorescence tomography reconstruction system based on a dual-path condition-guided diffusion model according to claim 3, characterized in that, The fault-lift representation mechanism includes a first stage and a second stage; The first stage includes extracting cross-channel spatial features from a single-view two-dimensional projection using depthwise separable convolution, and performing channel transformation based on feature depth to achieve feature extraction and preliminary enhancement; wherein, the depthwise separable convolution consists of depthwise convolution and pointwise convolution; The second stage includes compressing the boosted features into channel statistics using global average pooling, converting the channel statistics into activation factors for scaling the features using a fully connected layer and Sigmoid, and then performing shape transformation on the features to obtain the final boosted volumetric feature representation.
5. The single-view projection fast fluorescence tomography reconstruction system based on a dual-path condition-guided diffusion model according to claim 3, characterized in that: Local guidance based on multi-feature fusion is used to utilize the fusion of spatial and channel dual branches. In the initialization stage of the diffusion model, multi-feature fusion is performed on shallow noise features, local condition features and time step embedding features to provide fine and time-dependent local guidance for diffusion generation. Global guidance based on difference-common transformation is used to combine deep noise features with global condition features located in the middle layer of the conditional model through a difference information injection module and a pair of common information injection modules, so as to ensure the rationality of global structure reconstruction.
6. The single-view projection fast fluorescence tomography reconstruction system based on a dual-path condition-guided diffusion model according to claim 1, characterized in that, The model optimization module includes: The tuning unit is used to train and update the initial single-view projection reconstruction model using the training set. During the training process, the model loss is calculated using the validation set and the loss function, and hyperparameters are set and optimized using the Adam optimizer and linear decay of the learning rate to complete the model tuning and obtain the single-view projection reconstruction model. The loss function includes a guiding term, a physical consistency term, and a diffusion prior term. The verification unit is used to verify and evaluate the single-view projection reconstruction model using a test set.