Training Method Based on Dual-Domain Neural Network and Photoacoustic Image Reconstruction Method
The dual-domain neural network addresses inefficiencies in light-sheet microscopy by learning sparse-to-dense image mappings, effectively reconstructing high-quality images from sparse data with improved efficiency and accuracy.
Patent Information
- Application Number
- CN202111683421.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-12-31
AI Technical Summary
In the existing photoacoustic computing tomography technology, due to the sparse viewing angle problem caused by sparse detector arrangement, serious fringe artifacts occur in the reconstruction image, reducing the readability and quantitative accuracy of the image. The traditional analytical reconstruction algorithm is not applicable. The iterative reconstruction algorithm based on deep learning is inefficient and has a large demand for GPU video memory, making it difficult to reconstruct high-resolution images.
The training method based on dual-domain neural network is adopted, including the data domain D-net network and the image domain I-net network. Through jump connection and back projection layers, a DI-net network model is constructed, and the sample data set is trained using sparse perspective photoacoustic signals and photoacoustic images to optimize the network structure, realize end-to-end training, reduce downsampling losses, and improve image quality.
Effectively suppress the striped artifacts caused by sparse viewing angles, restore more image details, improve image quality, reduce computing complexity and GPU video memory requirements, and realize efficient and high-quality photoacoustic image reconstruction.
Smart Images

Figure CN114332283B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image reconstruction, and particularly to a training method based on a dual-domain neural network and a photoacoustic image reconstruction method using the same. Background Art
[0002] Photoacoustic computed tomography (PACT) is a non-destructive biomedical imaging modality with high optical imaging contrast and ultrasonic imaging penetration depth, and has unique application value in the biomedical field. In order to obtain high-quality photoacoustic images, the signal acquisition device of the imaging system needs to perform spatial sampling with a dense view angle using a high-density array detector. However, in practical applications, due to limitations such as economic cost, process technology, and imaging time, the arrangement of detectors is often sparse, and only spatially undersampled photoacoustic signals can be obtained, which cannot meet the condition of data completeness, resulting in serious stripe artifacts in the reconstructed image and reducing the readability and quantitative accuracy of the image. Therefore, in order to improve the imaging quality, the imaging system needs to be equipped with a reconstruction algorithm that can reconstruct high-quality photoacoustic images using sparse photoacoustic data.
[0003] Traditional analytical reconstruction algorithms are derived from strict mathematical physics models. The necessary condition for traditional analytical reconstruction algorithms to achieve stable reconstruction is the completeness of data. Therefore, such algorithms are not applicable to sparse-view PACT image reconstruction.
[0004] Iterative reconstruction algorithms based on compressive sensing can reconstruct high-quality photoacoustic images using sparse photoacoustic data. However, such algorithms need to repeatedly calculate the forward and backward processes of photoacoustics many times, resulting in low reconstruction efficiency and limited application scenarios.
[0005] In recent years, researchers have begun to apply deep learning techniques to sparse-view PACT image reconstruction. In 2018, Hauptman et al. proposed an iterative reconstruction algorithm based on deep learning. The basic idea is to expand the solution process of traditional iterative algorithms into a cascaded neural network, allowing the network to automatically learn the regularization expression and optimization process of the algorithm to improve the reconstruction speed and image quality. However, similar to traditional iterative algorithms, such algorithms need to repeatedly calculate the forward and backward processes of photoacoustics many times, resulting in low reconstruction efficiency. In addition, in the end-to-end training mode, such algorithms have a large demand for GPU video memory and are not applicable to the reconstruction of high-resolution images. In 2019, Davoudi et al. proposed a PACT image post-processing technique based on a convolutional neural network (Post-Unet). The main principle is to enhance the image by allowing the U-Net network to learn the mapping relationship between low-quality images and high-quality images. The main disadvantage of such methods is that when the quality of the input image is low, details that have been lost in the image or structures obscured by artifacts are still difficult to recover. Summary of the Invention
[0006] In view of this, the main object of the present invention is to disclose a training method based on a dual-domain neural network and a photoacoustic image reconstruction method, so as to partially solve at least one of the above-mentioned technical problems.
[0007] To achieve the above object, one aspect of the present invention discloses a training method based on a dual-domain neural network, including:
[0008] Obtain a training sample data set, where the training sample data set includes photoacoustic signals and photoacoustic images;
[0009] Train the DI-net network model based on the training sample data set to obtain a trained DI-net network model, where the DI-net network model includes a data domain D-net network, an image domain I-net network, and a backprojection layer between the data domain D-net network and the image domain I-net network.
[0010] According to an embodiment of the present invention, the data domain D-net network includes a first encoder and a first decoder:
[0011] There is a skip connection between the first encoder and the first decoder; the skip connection is used to reduce data loss caused during downsampling in the data domain D-net network, and the downsampling is an operation performed in the first encoder;
[0012] Among them, the first encoder includes:
[0013] Cascaded first encoding sub-module, second encoding sub-module, third encoding sub-module, and fourth encoding sub-module, where the first encoding sub-module, the second encoding sub-module, the third encoding sub-module, and the fourth encoding sub-module each include two convolutional layers and one max pooling layer;
[0014] Among the first encoding sub-module, the second encoding sub-module, the third encoding sub-module, and the fourth encoding sub-module, the number of convolutional kernels of the two convolutional layers is the same, and the number of convolutional kernels in the cascaded first encoding sub-module, second encoding sub-module, third encoding sub-module, and fourth encoding sub-module doubles in sequence; and
[0015] The first decoder of the data domain D-net network includes:
[0016] Cascaded fourth decoding sub-module, third decoding sub-module, second decoding sub-module, and first decoding sub-module, where the fourth decoding sub-module, the third decoding sub-module, the second decoding sub-module, and the first decoding sub-module each include two convolutional layers and one max pooling layer;
[0017] Among them, the number of convolution kernels of the two convolution layers of the fourth decoding sub-module, the third decoding sub-module, the second decoding sub-module, and the first decoding sub-module is the same, and the number of convolution kernels of the cascaded fourth decoding sub-module, third decoding sub-module, second decoding sub-module, and first decoding sub-module is halved in turn;
[0018] Among them, the sub-modules with the same number of convolution kernels in the first encoder and the first decoder of the data domain D-net network are sub-modules of the same layer.
[0019] According to an embodiment of the present invention, the image domain I-net network includes a second encoder and a second decoder:
[0020] There is a skip connection between the second encoder and the second decoder; the skip connection is used to reduce data loss caused by the downsampling process in the image domain I-net network, and the downsampling is an operation performed in the second encoder;
[0021] Among them, the second encoder includes:
[0022] Cascaded fifth encoding sub-module, sixth encoding sub-module, seventh encoding sub-module, and eighth encoding sub-module, and the fifth encoding sub-module, the sixth encoding sub-module, the seventh encoding sub-module, and the eighth encoding sub-module each include two convolution layers and one max pooling layer;
[0023] Among the fifth encoding sub-module, the sixth encoding sub-module, the seventh encoding sub-module, and the eighth encoding sub-module, the number of convolution kernels of the two convolution layers is the same, and the number of convolution kernels in the cascaded fifth encoding sub-module, sixth encoding sub-module, seventh encoding sub-module, and eighth encoding sub-module doubles in turn; and
[0024] The second decoder of the image domain I-net network includes:
[0025] Cascaded eighth decoding sub-module, seventh decoding sub-module, sixth decoding sub-module, and fifth decoding sub-module, and the eighth decoding sub-module, the seventh decoding sub-module, the sixth decoding sub-module, and the fifth decoding sub-module each include two convolution layers and one max pooling layer;
[0026] Among them, the number of convolution kernels of the two convolution layers of the eighth decoding sub-module, the seventh decoding sub-module, the sixth decoding sub-module, and the fifth decoding sub-module is the same, and the number of convolution kernels in the cascaded eighth decoding sub-module, seventh decoding sub-module, sixth decoding sub-module, and fifth decoding sub-module is halved in turn;
[0027] Among them, the sub-modules with the same number of convolutional kernels in the second encoder and the second decoder of the image domain I-net network are the same-layer sub-modules.
[0028] According to an embodiment of the present invention, the data domain D-net network and the image domain I-net network further include:
[0029] Establish skip connections between the same-layer sub-modules of the data domain D-net network and the image domain I-net network and between the network and the last layer of neurons;
[0030] Establish skip connections between the first layer of neural network and the last layer of neural network of the data domain D-net network; establish skip connections between the first layer of neural network and the last layer of neural network of the image domain I-net network; and
[0031] Add a normalization layer and a non-linear activation function layer to each convolutional layer of the first encoder and the second encoder.
[0032] According to an embodiment of the present invention, the data domain D-net network and the image domain I-net network are connected by a back-projection layer, expressed in the following matrix form:
[0033] P = BY;
[0034] Wherein, P is a one-dimensional column vector containing the information of the image to be reconstructed, and after P is rearranged, it forms an image matrix to be reconstructed and is input into the image domain I-net network; B is a back-projection matrix, and the data domain D-net network and the image domain I-net network are connected through the back-projection matrix B; Y is a back-projection term matrix, which is generated based on the data domain D-net network.
[0035] According to an embodiment of the present invention, the training sample data set includes:
[0036] Spatially undersampled photoacoustic signals and spatially fully sampled photoacoustic signals, and photoacoustic images corresponding to the spatially fully sampled photoacoustic signals obtained by an image reconstruction algorithm; the spatially undersampled photoacoustic signals and the spatially fully sampled photoacoustic signals form a first data set pair, and the spatially undersampled photoacoustic signals and the photoacoustic images corresponding to the fully sampled photoacoustic signals form a second data set pair.
[0037] According to an embodiment of the present invention, the spatially undersampled photoacoustic signals and the spatially fully sampled photoacoustic signals are obtained by collecting photoacoustic signals with different numbers of channels by a detector device and processing them through a filter matrix, expressed in the following matrix form:
[0038] Y0 = FX;
[0039] Where F is the filtering matrix; X is the photoacoustic signal matrix of different channel numbers collected by the detector device; Y0 is the spatially undersampled photoacoustic signal matrix or the spatially fully sampled photoacoustic signal matrix.
[0040] According to an embodiment of the present invention, the training of the DI-net network includes:
[0041] Independently train the data domain D-net network to obtain the first optimal network structure of the data domain D-net network;
[0042] Train the image domain I-net network. After obtaining the first optimal network structure of the data domain D-net network and fixing the first weight parameters of the data domain D-net network, based on the end-to-end training of the DI-net network model, obtain the second optimal network structure of the image domain I-net network and fix the second weight parameters of the image domain I-net network;
[0043] Restore the trainability of the first weight parameters of the data domain D-net network, and perform end-to-end fine-tuning training on the DI-net network model to obtain the DI-net network model;
[0044] Among them, the fixing of the first weight parameters and the second weight parameters is monitored based on the squared error loss function.
[0045] According to an embodiment of the present invention, the Adam algorithm is used for optimization during the training process; during the training process, a learning rate decay strategy is used, and the learning rate decay strategy includes controlling the decay of the learning rate based on the decrease in the loss index of the validation set.
[0046] For another aspect of the present invention, a photoacoustic image reconstruction method is disclosed, including:
[0047] Input the sparse-view photoacoustic signal into the DI-net network model to obtain a reconstructed image; wherein, the DI-net network model is trained by the method described above.
[0048] According to the above-mentioned embodiment of the present invention, a training method based on a dual-domain neural network and a photoacoustic image reconstruction method using the same, input the sparse-view photoacoustic signal into the DI-net network model to obtain a reconstructed image. Using the trained DI-net network model for image reconstruction can suppress the stripe artifacts caused by sparse views and improve the image quality. Description of the Drawings
[0049] The above and other objects, features, and advantages of the present invention will become more apparent from the following description of the embodiments of the present invention with reference to the accompanying drawings. In the drawings:
[0050] Figure 1 (a) Schematically shows the structural diagram of the DI-net network model of an embodiment of the present invention;
[0051] Figure 1 (b) Schematically shows the structural diagram of the data domain D-net network of an embodiment of the present invention;
[0052] Figure 2 Schematically shows the flowchart of training the DI-net network of an embodiment of the present invention;
[0053] Figure 3 Schematically shows the schematic diagram of the experimental device of an embodiment of the present invention;
[0054] Figure 4 Schematically shows the reconstruction result diagram of the mouse tomographic data under 128 detection view conditions of an embodiment of the present invention;
[0055] Figure 5 Schematically shows the reconstruction result diagram of the mouse tomographic data under 256 detection view conditions of an embodiment of the present invention;
[0056] Figure 6 Schematically shows the statistical results of the quantitative evaluation of an embodiment of the present invention. Detailed Embodiments
[0057] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following further elaborates on the present invention in detail with reference to embodiments and the accompanying drawings.
[0058] The terms used herein are merely for describing embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0059] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0060] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0061] On the one hand, the present invention discloses a training method for a dual-domain neural network, including:
[0062] Obtaining a training sample data set, where the training sample data set includes photoacoustic signals and photoacoustic images;
[0063] Training the DI-net network model with the training sample data set to obtain a trained DI-net network model, and constructing the DI-net network model, where the DI-net network model includes a data domain D-net network, an image domain I-net network, and a back-projection layer between the data domain D-net network and the image domain I-net network. Figure 1 (a) Schematically shows the structural diagram of the DI-net network model according to an embodiment of the present invention.
[0064] According to an exemplary embodiment of the present invention, as Figure 1 (a) shows, the DI-net network model includes a data domain D-net network, an image domain I-net network, and a back-projection layer between the data domain D-net network and the image domain I-net network.
[0065] More specifically, during the training process, the data domain D-net network processes by extracting photoacoustic signal features from the input photoacoustic signals to obtain the mapping relationship between the sparse photoacoustic signals of the training sample data set and the dense photoacoustic signals of the training sample data set; the image domain I-net network processes by extracting photoacoustic image features from the input photoacoustic images to obtain the mapping relationship between the photoacoustic images processed by the data domain D-net network and the high-quality photoacoustic images corresponding to the dense photoacoustic signals of the training sample data set; the back-projection layer connects the data domain D-net network and the image domain I-net network to obtain the predicted high-quality photoacoustic images.
[0066] Figure 1 (b) Schematically shows the structural diagram of the data domain D-net network according to an embodiment of the present invention.
[0067] According to an exemplary embodiment of the present invention, as Figure 1(As shown in (b), the left dashed block diagram schematically shows the structure of the first encoder of the data domain D-net network. The first encoder is composed of cascaded first, second, third, and fourth encoding sub-modules. Each of the first, second, third, and fourth encoding sub-modules contains two convolutional layers and one max-pooling layer. Among them, the number of convolutional kernels of the two convolutional layers contained in each of the first, second, third, and fourth encoding sub-modules is the same, and the number of convolutional kernels in the cascaded first, second, third, and fourth encoding sub-modules doubles in sequence; the right dashed block diagram schematically shows the structure of the first decoder of the image domain I-net network. The first decoder is composed of cascaded fourth, third, second, and first decoding sub-modules. Each of the fourth, third, second, and first decoding sub-modules contains two convolutional layers and one max-pooling layer. Among them, the number of convolutional kernels of the two convolutional layers contained in each of the fourth, third, second, and first decoding sub-modules is the same, and the number of convolutional kernels in the cascaded fourth, third, second, and first decoding sub-modules halves in sequence.)
[0068] More specifically, the photoacoustic signal with a size of M×N enters the data domain D-net network and then enters the first encoding sub-module. The initial number of channels of the first encoding sub-module is k, and the value of k is determined by the number of initial filters of the data domain D-net network. After image feature extraction through the two convolutional layers in the first encoding sub-module, the method can be through a 3×3 convolutional kernel, one normalization layer, and one non-linear activation function layer, as well as downsampling by the max-pooling layer. The size of the input data is halved and the number of channels is doubled after passing through the first encoding sub-module. The operation process of the data output from the first encoding sub-module through the cascaded second, third, and fourth encoding sub-modules is the same as or similar to that of the first encoding sub-module, which will not be described here. Finally, the fourth encoding sub-module outputs data with a size of M / 16×N / 16 and 16k channels; the data with a size of M / 16×N / 16 and 16k channels is input into the first decoder and then enters the fourth decoding sub-module for upsampling processing. The size of the input data matrix doubles and the number of channels halves. The method of passing through the cascaded third, second, and first decoding sub-modules is the same as or similar to that described above, which will not be described here. Finally, the first decoding sub-module outputs an image with a size of M×N.)
[0069] According to an exemplary embodiment of the present invention, in each convolutional layer of each sub-module of the first encoder of the data domain D-net network, there is also one normalization layer and one non-linear activation function layer.)
[0070] More specifically, a normalization layer and a non-linear activation function layer included in each convolutional layer of the first encoder of the data domain D-net network improve the performance and training stability of the data domain D-net network.
[0071] According to an exemplary embodiment of the present invention, sub-modules having the same number of convolutional kernels in the first encoder and the first decoder of the data domain D-net network are sub-modules of the same layer.
[0072] More specifically, skip connections are established between sub-modules of the same layer in the data domain D-net network. When the decoding sub-module processes the photoacoustic signal, the skip connections enable the decoding sub-module to use both the features that appear after upsampling processing and the image features in the encoding sub-module of the same layer, solving the problem of possible loss of spatial resolution caused by multiple downsampling operations in the first encoding sub-module and preventing the loss of image details; skip connections are established between the first layer neural network and the last layer neural network of the data domain D-net network, and the skip connections enable the data domain D-net network to only learn the difference matrix between the input photoacoustic signal matrix and the output photoacoustic signal matrix, reducing the learning difficulty of the data domain D-net network.
[0073] According to an exemplary embodiment of the present invention, the structure diagram of the image domain I-net network and the structures of the second encoder and the second decoder in an embodiment of the present invention are the same as or similar to the structure of the above data domain D-net network. However, since the image domain I-net network needs to extract and process more detailed features of the input data, the initial number of channels of the image domain I-net network is set to twice the number of channels of the data domain D-net network, and the number of convolutional kernels corresponding to each sub-module is also twice that of the data domain D-net network. The rest will not be elaborated here.
[0074] According to an exemplary embodiment of the present invention, the data D-net network and the image domain I-net network of an embodiment of the present invention are connected by a back-projection layer, expressed in the following matrix form:
[0075] P = BY; (1)
[0076] Wherein, P is a one-dimensional column vector containing the information of the image to be reconstructed, which is rearranged to form an image matrix to be reconstructed and input into the image domain I-net network; B is a back-projection matrix, which connects the data domain D-net network and the image domain I-net network through the back-projection matrix B; Y is a back-projection term matrix, which is generated based on the data domain D-net network.
[0077] More specifically, the data domain D-net network processes by extracting the photoacoustic signal features based on the input photoacoustic signal. The image domain I-net network processes by extracting the photoacoustic image features from the input image matrix to be reconstructed. Through the connection of the back-projection matrix B in the back-projection layer, the mapping relationship between the data domain D-net network and the image domain I-net network for the input photoacoustic data is enhanced, and a predicted high-quality photoacoustic image is obtained.
[0078] According to an exemplary embodiment of the present invention, the training sample data set of an embodiment of the present invention includes:
[0079] Spatially undersampled photoacoustic signals and spatially fully sampled photoacoustic signals, and photoacoustic images corresponding to the spatially fully sampled photoacoustic signals obtained by corresponding image reconstruction algorithms. The spatially undersampled photoacoustic signals and the spatially fully sampled photoacoustic signals form a first data set pair, and the spatially fully sampled photoacoustic signals and the photoacoustic images corresponding to the spatially fully sampled photoacoustic signals form a second data set pair.
[0080] More specifically, by performing downsampling on the photoacoustic signals collected by the existing acquisition device at different multiples, undersampled photoacoustic signals with different numbers of channels can be obtained to meet different training requirements. Moreover, the spatially undersampled photoacoustic signals and the spatially fully sampled photoacoustic signals are obtained after the photoacoustic signals with different numbers of channels collected by the detector device are processed by a filtering matrix, and are expressed in the following matrix form:
[0081] Y0 = FX; (2)
[0082] where F is the filtering matrix; X is the original photoacoustic signal matrix with different numbers of channels collected by the detector device; and Y0 is the spatially undersampled photoacoustic signal matrix or the spatially fully sampled photoacoustic signal matrix.
[0083] Figure 2 Schematically shows the flowchart of training the DI-net network according to an embodiment of the present invention.
[0084] According to an exemplary embodiment of the present invention, as Figure 2 shown, training the DI-net network includes:
[0085] S1: Training the data domain D-net network alone to obtain the first optimal network structure of the data domain D-net network;
[0086] S2: Training the image domain I-net network. After obtaining the first optimal network structure of the data domain D-net network and fixing the first weight parameters of the data domain D-net network, based on the end-to-end training of the DI-net network model, the second optimal network structure of the image domain I-net network is obtained, and the second weight parameters of the image domain I-net network are fixed;
[0087] S3: Restore the trainability of the first weight parameters of the data domain D-net network, and perform end-to-end fine-tuning training on the DI-net network model to obtain the DI-net network model;
[0088] Among them, the fixing of the first weight parameters and the second weight parameters is monitored by the mean squared error loss function. When the result of the mean squared error loss function reaches the standard for the data domain D-net network and the image domain I-net network to reach the optimal network structure, the weight parameters at this time are fixed to obtain the first weight parameters and the second weight parameters.
[0089] According to an exemplary embodiment of the present invention, the Adam algorithm is used for optimization during the training process; during the training process, a learning rate decay strategy is used, and the learning rate decay strategy includes controlling the decay of the learning rate based on the decrease in the loss index of the validation set.
[0090] Another aspect of the present invention discloses a photoacoustic image reconstruction method, including:
[0091] Input the sparse-view photoacoustic imaging image into the DI-net network model to obtain a reconstructed image, where the DI-net network model is trained by the above-mentioned training method based on the dual-domain neural network. For the reconstructed image obtained by the trained DI-net network model, image artifacts are suppressed, and at the same time, more image details are restored.
[0092] Figure 3 Schematically shows a schematic diagram of an experimental device according to an embodiment of the present invention.
[0093] As Figure 3 shown, the acquisition device can use an annular detector array, which contains 512 independent detection units. The detection radius and imaging area size of the annular detector array can be similar to the sizes required for simulation for subsequent data processing.
[0094] According to an embodiment of the present invention, experimental tests show that the initial M of the data domain D-net network is set to 512 pixels, N is set to 768 pixels, the initial number of channels k is 16, each of the two convolutional layers contains a convolutional kernel with a size of 3×3 and a stride of 1, a layer of normalization layer, and a layer of non-linear activation function layer, and then enters a max-pooling layer for downsampling processing and then outputs. The max-pooling layer contains a kernel with a size of 2×2 and a stride of 2. The initial number of channels k of the image domain I-net network is 32. The above settings can effectively reduce the parameter calculation amount while ensuring the reconstruction performance of the DI-net network model.
[0095] More specifically, the encoder module of the DI-net network model can be the contracting path, and the decoder module can be the expanding path. The transposed convolutional layer in the decoder module performs the upsampling operation to restore the image resolution to the original size. The multiple downsampling operations in the encoder module may cause loss of spatial resolution. To prevent loss of image details, skip connections are introduced between the encoder and decoder modules of the D-net network in the data domain and the I-net network in the image domain. Due to the skip connections, the decoding sub-module can use both the features that appear after upsampling processing and the image features in the same-layer encoding sub-module. Additionally, skip connections are established between the first neural network and the last neural network of the D-net network in the data domain and the I-net network in the image domain. Due to the skip connections, the D-net network in the data domain and the I-net network in the image domain only need to learn the difference image between the input image and the output image, reducing the learning difficulty of the network. Normalization layers (InstanceNorm) and non-linear activation function layers (Leaky ReLu) are added to the convolutional layers of the first encoder and the second encoder of the DI-net network model, which can improve the performance and training stability of the DI-net network model.
[0096] According to an embodiment of the present invention, two sets of sparse-view photoacoustic imaging training sample data sets of live mice can be constructed. The number of detection views of the sparse-view photoacoustic imaging signals is 128 and 256 respectively. By performing 2-fold and 4-fold downsampling on the photoacoustic signals with 512 channels collected by the Figure 3 device, photoacoustic signals with 256 channels and 128 channels are obtained. The size of the above photoacoustic signals is 512×768 pixels. The photoacoustic signals with 512 channels and their corresponding reconstructed images are used as references for network training.
[0097] More specifically, tomographic imaging was performed on 14 mice, and a total of 9074 photoacoustic signals of mouse slices were obtained (the number of slices collected from different mice is not exactly the same). Among them, 12 were used for training (7500 slices), 1 was used for validation (787 slices), and the remaining 1 was used for testing (787 slices).
[0098] The reference photoacoustic image is reconstructed by the filtered back-projection algorithm. The reconstruction process of this algorithm can be expressed as:
[0099]
[0100] In the formula,
[0101]
[0102] where, p0(r s ) is the photoacoustic reconstruction image, b(r d, t) is the back-projection term, r s is the position of the photoacoustic source, r d is the position of the detector, Ω is the solid angle enclosed by the detection surface (for an infinite planar geometry, Ω = 2π; for spherical and cylindrical geometries, Ω = 4π), dΩ is the solid angle corresponding to the area dσ of a single detection unit, p(r d , t) is the photoacoustic signal obtained by the detector, t is time, c is the speed of sound, and dΩ can be expressed as:
[0103]
[0104] where n d represents the unit normal vector on the detector surface pointing to the region of interest.
[0105] The photoacoustic signals with 512 channels collected are respectively downsampled by 2 times and 4 times to obtain sparse photoacoustic signals with 128 and 256 channels; to simplify the application of the DI-Net network model and accelerate the network training process, the sparse photoacoustic signals are filtered, and the preprocessed photoacoustic signals are used as the input of the network. The preprocessing process is as follows: First, the sparse photoacoustic signals with 128 or 256 channels are interpolated into low-quality, 512-channel photoacoustic signals using the linear interpolation method; then, the interpolated sampled signals are filtered using the above formula (4) to obtain the preprocessed photoacoustic signals, and this signal is used as the input of the network. The filtering process can be expressed in the following matrix form:
[0106] Y0 = FX; (2)
[0107] where F is the filtering matrix; X is the matrix of the original photoacoustic signals with different numbers of channels collected by the detector device; the obtained Y0 is the matrix of the spatially undersampled photoacoustic signals or the matrix of the spatially fully sampled photoacoustic signals for network training.
[0108] According to an embodiment of the present invention, Y0 obtained from the above formula (2), after passing through the data domain D-net network, can obtain the back-projection term matrix Y in the above formula (1).
[0109] The data D-net network and the image domain I-net network are connected by a back-projection layer, and are expressed in the following matrix form:
[0110] P = BY; (1)
[0111] where P is a one-dimensional column vector containing the information of the image to be reconstructed, and after rearrangement, it forms the matrix of the image to be reconstructed and is input into the image domain I-net network; B is the back-projection matrix that connects the data domain D-net network and the image domain I-net network; Y is the back-projection term matrix generated based on the data domain D-net network.
[0112] More specifically, the data domain D-net network processes by extracting the photoacoustic signal features from the input photoacoustic signal. The image domain I-net network is based on the input image matrix to be reconstructed, which is obtained from the one-dimensional column vector P containing the information of the image to be reconstructed. The size of the image matrix to be reconstructed is determined according to specific circumstances and can be 256×256 pixels. It extracts the photoacoustic image features for processing, and connects the data domain D-net network and the image domain I-net network through the backprojection matrix B in the backprojection layer to obtain the predicted high-quality photoacoustic image.
[0113] According to an embodiment of the present invention, training the DI-net network includes the following steps:
[0114] S1: Train the data domain D-net network to obtain the first optimal network structure of the data domain D-net network. Through the training in this stage, the data domain D-net network can learn the mapping relationship between spatially undersampled photoacoustic data and spatially fully sampled photoacoustic data;
[0115] S2: Train the image domain I-net network. After obtaining the first optimal network structure of the data domain D-net network and fixing the first weight parameters of the data domain D-net network, perform end-to-end training based on the DI-net network model to obtain the second optimal network structure of the image domain I-net network and fix the second weight parameters of the image domain I-net network. Through the training in this stage, the image domain I-net network can map the photoacoustic image output by the backprojection layer into a high-quality photoacoustic image;
[0116] S3: Restore the trainability of the first weight parameters of the data domain D-net network, and perform end-to-end fine-tuning training on the DI-net network model to obtain the DI-net network model;
[0117] During the training process, the Adam algorithm is used for optimization, and the Batch size is 8. The initial learning rates in steps S1 and S2 are set to 5×10 -4 , and train for 100 epochs; the initial learning rate for fine-tuning training in step S3 is set to 1×10 -5 , and train for 30 epochs. During the training process, a learning rate decay strategy is used. When the loss value on the validation set does not decrease within 3 epochs, the learning rate will automatically decay to 0.8 times the original. All training is completed under the TensorFlow 2.0 framework, and the computer configuration used for training is an Inter Xeon Glod 6226R CPU, an NVIDIA RTX TITAN GPU (24GB video memory), and an Ubuntu operating system.
[0118] The image quality of the reconstructed image obtained by the DI-net network model after the above training is better than that of the reconstructed image obtained by other reconstruction methods for sparse-view photoacoustic signals.
[0119] Figure 4 Schematically shows the reconstructed result diagram of the mouse tomographic data under 128 detection view conditions of an embodiment of the present invention.
[0120] As Figure 4 shown, Figure 4 Part a of is the reference image under 128 detection view conditions, Figure 4 Part b of Figure 4 Part c of Figure 4 Part d of are the reconstructed results of the FBP algorithm, the Post-Unet algorithm, and the DI-Net algorithm under 128 detection view conditions in sequence, Figure 4 Part e of Figure 4 Part f of Figure 4 Part g of are the difference image diagrams between the reference image and the FBP reconstructed image, the Post-Unet reconstructed image, and the DI-Net reconstructed image in sequence, Figure 4 Part h of is the quantitative evaluation result of MSE, PSNR, and SSIM under 128 detection view conditions. As can be seen from the schematically shown arrows in Figure 4 Part b of Figure 4 Part c of Figure 4 Part d of , due to undersampling in the spatial domain, serious stripe artifacts appear in the image reconstructed by the FBP algorithm, and these artifacts obscure the true photoacoustic structure, resulting in the loss of a lot of details in the reconstructed image; compared with the FBP algorithm, the Post-Unet algorithm better suppresses these artifacts, but there are still problems of detail loss in the image it reconstructs; the DI-Net algorithm realizes artifact-free sparse-view photoacoustic image reconstruction, and the details in the reconstructed image are more complete and clearer; Figure 4 The difference image and quantitative evaluation result shown in part h of indicate that DI-Net discloses the best reconstruction result, and its reconstructed image has a lower MSE, a higher PSNR, and SSIM.
[0121] Figure 5 Schematically shows the reconstructed result diagram of the mouse tomographic data under 256 detection view conditions of an embodiment of the present invention.
[0122] As Figure 5 shown, Figure 5 Part a of is the reference image under 256 detection view conditions, Figure 5 Part b of Figure 5 Part of Figure 5Part d shows the reconstruction results of the FBP algorithm, the Post-Unet algorithm, and the DI-Net algorithm under 256 detection view conditions in sequence. Figure 5 Part e of Figure 5 Part f of Figure 5 Part g of Figure 5 shows the difference images between the reference image and the FBP reconstructed image, the Post-Unet reconstructed image, and the DI-Net reconstructed image in sequence. Figure 5 As can be seen from Part b of Figure 5 With the increase in the number of detection view angles, the image quality reconstructed by the FBP algorithm has been significantly improved. However, there are still obvious stripe artifacts in the image, resulting in low image contrast and poor overall visual quality. As shown in Part c of Figure 5 and Part d of
[0123] Figure 6 The Post-Unet algorithm and the DI-Net algorithm can better suppress these artifacts and obtain high-quality images without artifacts.
[0124] As Figure 6 shown, Figure 6 In Part a of Figure 6 the quantitative evaluation results of the MSE evaluation index of the FBP reconstructed image, the Post-Unet reconstructed image, and the DI-Net reconstructed image methods under 128 detection view conditions are shown from left to right in sequence; Figure 6 In Part b of Figure 6 the quantitative evaluation results of the PSNR evaluation index of the FBP reconstructed image, the Post-Unet reconstructed image, and the DI-Net reconstructed image methods under 128 detection view conditions are shown from left to right in sequence; Figure 6The quantitative evaluation results of the PSNR evaluation metrics of the FBP reconstruction image, the Post-Unet reconstruction image, and the DI-Net reconstruction image methods under 256 detection view conditions from left to right in part e; Figure 6 The quantitative evaluation results of the SSIM evaluation metrics of the FBP reconstruction image, the Post-Unet reconstruction image, and the DI-Net reconstruction image methods under 256 detection view conditions from left to right in part f. It can be seen that compared with the FBP algorithm reconstruction and the Post-Unet algorithm reconstruction, the DI-Net has higher reconstruction accuracy, the reconstructed image is closer to the reference image, with lower MSE, higher PSNR, and SSIM.
[0125] The present invention also discloses a training device, which may include:
[0126] An acquisition module, configured to acquire a training sample data set, where the training sample data set includes photoacoustic signals and photoacoustic images;
[0127] A training module, configured to train a DI-net network model based on the training sample data set acquired by the acquisition module to obtain a trained DI-net network model;
[0128] Wherein, the training module includes:
[0129] A first sub-module, configured to process the extraction of photoacoustic signal features to obtain the mapping relationship between the sparse photoacoustic signals of the training sample data set and the dense photoacoustic signals of the training sample data set;
[0130] A second sub-module, configured to process the extraction of photoacoustic image features to obtain the mapping relationship between the low-quality photoacoustic images obtained from the sparse photoacoustic signals of the training sample data set and the high-quality photoacoustic images obtained from the dense photoacoustic signals of the training sample data set;
[0131] A third sub-module, configured to connect the first sub-module and the second sub-module, enabling the algorithm to enhance photoacoustic signals and photoacoustic images simultaneously in the data domain and the image domain.
[0132] The present invention also discloses a device for reconstructing photoacoustic images, which may include:
[0133] An acquisition module, configured to acquire sparse-view photoacoustic imaging images;
[0134] A reconstructed image module, configured to acquire a reconstructed image, and the reconstructed image module is obtained by using the training device module.
[0135] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in an order different from that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features recited in the various embodiments and / or claims of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features recited in the various embodiments and / or claims of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0136] The embodiments of the present invention have been described above. However, the embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. The scope of the present invention is defined by the appended claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A training method based on a dual-domain neural network, comprising: Obtaining a training sample data set, the training sample data set including: spatially undersampled photoacoustic signals and spatially fully sampled photoacoustic signals, and photoacoustic images corresponding to the fully sampled photoacoustic signals obtained by an image reconstruction algorithm from the spatially fully sampled photoacoustic signals; the spatially undersampled photoacoustic signals and the spatially fully sampled photoacoustic signals form a first data set pair, and the spatially undersampled photoacoustic signals and the photoacoustic images corresponding to the spatially fully sampled photoacoustic signals form a second data set pair; Training a DI-net network model based on the training sample data set to obtain a trained DI-net network model, wherein the DI-net network model includes a data domain D-net network, an image domain I-net network, and a backprojection layer between the data domain D-net network and the image domain I-net network; Wherein, the training of the DI-net network model based on the training sample data set to obtain a trained DI-net network model includes: Using the first data set pair to separately train the data domain D-net network to obtain a first optimal network structure of the data domain D-net network; the first optimal network structure learns the mapping relationship between the spatially undersampled photoacoustic signals and the spatially fully sampled photoacoustic signals; Training the image domain I-net network. After obtaining the first optimal network structure of the data domain D-net network and fixing the first weight parameters of the data domain D-net network, using the second data set pair to perform end-to-end training on the DI-net network model to obtain a second optimal network structure of the image domain I-net network and obtain second weight parameters of the image domain I-net network; the second optimal network structure can map the photoacoustic image output by the backprojection layer into a photoacoustic image corresponding to the spatially fully sampled photoacoustic signal; Restoring the trainability of the first weight parameters of the data domain D-net network, and using the second data set pair to perform end-to-end fine-tuning training on the DI-net network model to obtain the DI-net network model; Wherein, the fixing of the first weight parameters and the second weight parameters is monitored based on a mean square error loss function.
2. The method according to claim 1, wherein, The data domain D-net network includes a first encoder and a first decoder: There is a skip connection between the first encoder and the first decoder; the skip connection is used to reduce data loss caused by downsampling in the data domain D-net network, and the downsampling is an operation performed in the first encoder; Wherein, the first encoder includes: Cascaded first, second, third, and fourth coding sub-modules, each of the first, second, third, and fourth coding sub-modules containing two convolutional layers and one max pooling layer; Among them, in the first encoding sub-module, the second encoding sub-module, the third encoding sub-module, and the fourth encoding sub-module, the number of convolution kernels of the two convolution layers included in each is the same, and the number of convolution kernels in the cascaded first encoding sub-module, second encoding sub-module, third encoding sub-module, and fourth encoding sub-module doubles in sequence; and The first decoder of the data domain D-net network includes: The cascaded fourth decoding sub-module, third decoding sub-module, second decoding sub-module, and first decoding sub-module, where the fourth decoding sub-module, the third decoding sub-module, the second decoding sub-module, and the first decoding sub-module each include two convolution layers and one max pooling layer; Among them, in the fourth decoding sub-module, the third decoding sub-module, the second decoding sub-module, and the first decoding sub-module, the number of convolution kernels of the two convolution layers included in each is the same, and the number of convolution kernels in the cascaded fourth decoding sub-module, third decoding sub-module, second decoding sub-module, and first decoding sub-module halves in sequence; Among them, the sub-modules with the same number of convolution kernels in the first encoder and the first decoder of the data domain D-net network are sub-modules of the same layer.
3. The method according to claim 2, wherein The image domain I-net network includes a second encoder and a second decoder: There is a skip connection between the second encoder and the second decoder; the skip connection is used to reduce the data loss caused by the downsampling process in the image domain I-net network, and the downsampling is an operation performed in the second encoder; Among them, the second encoder includes: The cascaded fifth encoding sub-module, sixth encoding sub-module, seventh encoding sub-module, and eighth encoding sub-module, where the fifth encoding sub-module, the sixth encoding sub-module, the seventh encoding sub-module, and the eighth encoding sub-module each include two convolution layers and one max pooling layer; Among them, in the fifth encoding sub-module, the sixth encoding sub-module, the seventh encoding sub-module, and the eighth encoding sub-module, the number of convolution kernels of the two convolution layers included in each is the same, and the number of convolution kernels in the cascaded fifth encoding sub-module, sixth encoding sub-module, seventh encoding sub-module, and eighth encoding sub-module doubles in sequence; and The second decoder of the image domain I-net network includes: The cascaded eighth decoding sub-module, seventh decoding sub-module, sixth decoding sub-module, and fifth decoding sub-module, where the eighth decoding sub-module, the seventh decoding sub-module, the sixth decoding sub-module, and the fifth decoding sub-module each include two convolution layers and one max pooling layer; Among them, in the eighth decoding sub-module, the seventh decoding sub-module, the sixth decoding sub-module, and the fifth decoding sub-module, the number of convolution kernels of the two convolution layers included in each is the same, and the number of convolution kernels in the cascaded eighth decoding sub-module, seventh decoding sub-module, sixth decoding sub-module, and fifth decoding sub-module halves in sequence; Among them, the sub-modules with the same number of convolutional kernels in the second encoder and the second decoder of the image domain I-net network are sub-modules of the same layer.
4. The method according to claim 3, wherein, The data domain D-net network and the image domain I-net network further include: Establishing skip connections between the sub-modules of the same layer of the first encoder and the first decoder of the data domain D-net network, and establishing skip connections between the first neural network layer and the last neural network layer of the data domain D-net network; Establishing skip connections between the sub-modules of the same layer of the second encoder and the second decoder of the image domain I-net network; establishing skip connections between the first neural network layer and the last neural network layer of the image domain I-net network; and Adding a normalization layer and a non-linear activation function layer to each convolutional layer of the first encoder and the second encoder.
5. The method according to claim 4, wherein The data domain D-net network and the image domain I-net network are connected by a back-projection layer, expressed in the following matrix form: P = BY; where P is a one-dimensional column vector containing the information of the image to be reconstructed, and after rearranging P, the image matrix to be reconstructed is input into the image domain I-net network; B is a back-projection matrix, and the data domain D-net network and the image domain I-net network are connected through the back-projection matrix B; Y is a back-projection term matrix, generated based on the data domain D-net network.
6. The method according to claim 1, wherein, The spatially undersampled photoacoustic signal and the spatially fully sampled photoacoustic signal are obtained after the photoacoustic signals of different channel numbers are collected by a detector device and processed by a filtering matrix, expressed in the following matrix form: Y0 = FX; where F is a filtering matrix; X is a photoacoustic signal matrix of different channel numbers collected by a detector device; Y0 is the spatially undersampled photoacoustic signal matrix or the spatially fully sampled photoacoustic signal matrix.
7. The method according to claim 1, wherein During the training process, the Adam algorithm is used for optimization; during the training process, a learning rate decay strategy is used, and the learning rate decay strategy includes controlling the decay of the learning rate based on the decrease of the loss index on the validation set.
8. A photoacoustic image reconstruction method, including: Inputting a sparse-view photoacoustic signal into a DI-net network model to obtain a reconstructed image; wherein, the DI-net network model is trained by using the method according to any one of claims 1-7.
Citation Information
Patent Citations
Tumor photoacoustic image rapid reconstruction method and device based on deep learning
CN110880196A
Dual-domain recursive network MR reconstruction method based on multi-module feature aggregation
CN113487507A