A Sparse Photoacoustic Image Reconstruction Method and System Combining Object Detection

By combining the object detection of Two-stage network and Transformers model and image reconstruction network based on GAN network and patchGan network, the problem of sparse image information and excessive background invalid information in photoacoustic microscopy is solved, and efficient image recovery and rapid reconstruction are achieved.

CN114332282BActive Publication Date: 2025-07-08SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111683028.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-07-08
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

When the existing photoacoustic microscopy imaging methods are used to imaging at the cell and even the nucleus level, the image information is sparse and the background is invalid, resulting in poor network robustness, difficulty in training, waste of computing overhead, and poor image recovery effect.

Method used

Combining the object detection network of the Two-stage network and the Transformers model and the image reconstruction network based on the GAN network and patchGan network, the coordinate vector data set of the region of interest is obtained through object detection, and the photoacoustic image is reconstructed using the image reconstruction network to reduce computing overhead and improve image recovery effect.

Benefits of technology

It realizes saving computing overhead in sparse photoacoustic images, good image recovery effect, fast reconstruction speed, and solves the problem of sparse image information and excessive background invalid information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332282B_ABST
    Figure CN114332282B_ABST
Patent Text Reader

Abstract

The present invention provides a sparse photoacoustic image reconstruction method combined with object detection, including: constructing a first training data set; constructing an object detection network; training the object detection network with the first training data set; processing the original photoacoustic data set with the object detection network to obtain a cropped image group data set; constructing a second training data set; constructing an image reconstruction network; training the image reconstruction network with the second training data set; processing the sparse photoacoustic image with the object detection network to obtain a coordinate vector group data set of the region of interest; and reconstructing and outputting the final photoacoustic image with the image reconstruction network. By combining an object detection network based on a Two-stage network and a Transformers model and an image reconstruction network based on a GAN network and a patchGan network, the present invention processes the sparse photoacoustic image with sparse sampling, and solves the problem of poor image restoration caused by sparse photoacoustic image information and excessive invalid background information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image reconstruction method, and in particular to a sparse photoacoustic image reconstruction method and system combined with object detection. Background Art

[0002] Photoacoustic imaging is a new type of biomedical photon imaging method. Its basic principle is the photoacoustic effect of biological tissues, that is, when biological tissues are irradiated by light, the biological tissues will absorb light energy and cause changes in internal temperature, and then generate acoustic signals due to thermal expansion. Because the light absorption coefficient and scattering coefficient of substances themselves are different, the acoustic signals generated by different substances are different. If a biological tissue is scanned and irradiated with short-pulse lasers of the same frequency, and then an ultrasonic detector is used to receive the ultrasonic waves generated by the tissue, and then an appropriate inversion algorithm is used to solve the acoustic inverse problem, the initial acoustic pressure distribution map or light absorption energy distribution map on the tissue surface can be reconstructed. Finally, the image information of the tissue can be inverted based on the distribution map. Compared with traditional medical imaging methods (CT, optical coherence tomography, MRI, ultrasonic imaging, etc.), the advantages of photoacoustic imaging are non-invasive, high penetration depth and high resolution. Currently, the main research field branches include photoacoustic tomography, photoacoustic microscopy and photoacoustic endoscopy, etc. Among them, photoacoustic microscopy (PAM) can achieve the effect of exceeding the resolution limit of optical imaging through a point excitation mode similar to optical confocal imaging. Therefore, photoacoustic microscopes are an important research focus in current biophoton medicine.

[0003] Photoacoustic microscopy generally scans an object by focusing light at a point. However, the current main problem is that the scanning method is generally realized by a mechanical three-dimensional moving platform to control the photoacoustic component or the sample, and the common scanning accuracy is at the micron level. The main problem of existing imaging is that the imaging time is too long. Among them, the method of using deep learning-based sparse sampling rate image reconstruction to reconstruct high sampling rate images has become one of the important solutions. However, the main research objects of existing deep learning image reconstruction methods are common publicly available datasets, such as DIV2K, Set5, Set14 and BSD100, etc. The images in these datasets are mostly photos of daily cities and natural landscapes, which are quite different from the objects such as biological tissues in photoacoustic microscopy. Among them, the distribution of image information is an important difference. In the target of photoacoustic microscopy, when imaging at the cell or even nuclear level, there will be a problem of overly sparse image information in the imaging of some biological tissues. Taking the monolayer (10 microns) brain cell slice of a healthy mouse as an example, the target is the brain nucleus of the mouse (with a diameter of 5-8 microns). In some imaging regions, there is a problem of sparse cell distribution, that is, the pixel ratio of the effective cell nuclei in a cropped image ranges from 10% to 20%, and the remaining pixels are the tissue fluid of the tissue slice, scanning out a large amount of invalid background information.

[0004] In the face of photoacoustic images with such sparse effective information, existing image reconstruction methods exhibit problems such as poor network robustness and difficult training, and cause a large amount of computational overhead to be wasted in the calculation of invalid backgrounds. Therefore, designing an image reconstruction network that can perform sparse sampling for sparse information images, improve the network stability while reducing the overall network computational overhead, has important research significance and broad application prospects. Summary of the Invention

[0005] To solve the problems in the prior art, the present invention provides a sparse photoacoustic image reconstruction method and system combining object detection. By combining an object detection network based on a Two-stage network and a Transformers model and an image reconstruction network based on a GAN network and a patchGan network, it processes the sparse photoacoustic images with point-by-point sampling to obtain a coordinate vector group dataset of the region of interest, and then reconstructs and outputs the final photoacoustic image, which not only saves a large amount of computational overhead, but also has good image restoration effect and fast image reconstruction speed, and solves the problem of poor image restoration effect caused by sparse photoacoustic image information and excessive invalid background information.

[0006] A sparse photoacoustic image reconstruction method combining object detection according to the present invention includes the following steps:

[0007] Step 1: Use the original photoacoustic dataset to construct a first training dataset for training the object detection network;

[0008] Step 2: Construct an object detection network based on a Two-stage network and a Transformers model;

[0009] Step 3: Train the object detection network using the first training dataset;

[0010] Step 4: Use the trained object detection network to process the original photoacoustic dataset to obtain a cropped image group dataset;

[0011] Step 5: Use the original photoacoustic dataset and the cropped image group dataset to construct a second training dataset for training the image reconstruction network;

[0012] Step 6: Construct an image reconstruction network based on a GAN network and a patchGan network;

[0013] Step 7: Train the image reconstruction network using the second training dataset;

[0014] Step 8: Use the trained object detection network to process the sparse photoacoustic images with point-by-point sampling to obtain a coordinate vector group dataset of the region of interest;

[0015] Step 9: Use the trained image reconstruction network to reconstruct and output the final photoacoustic image based on the original image and the coordinate vector set dataset of the region of interest.

[0016] The present invention is further improved. In step 1, the original photoacoustic dataset is labeled using the labeling tool labelimg to construct the first training dataset for training the object detection network.

[0017] The present invention is further improved. In step 8, the process of the object detection network processing the sparse photoacoustic image obtained by sampling at every other point to obtain the coordinate vector set dataset of the region of interest includes the following steps:

[0018] Step 801: Coarse detection, using 3 convolutional layers and 1 fully connected layer for shallow feature extraction;

[0019] Step 802: Fine detection, using the threshold method to select the grids with a confidence level greater than the set threshold, merging the selected grids with the absolute position encoding to form a new vector set, and sending it into the transformer network.

[0020] The present invention is further improved. In step 801, the size of the convolutional layer is determined according to the information distribution of the sparse photoacoustic image.

[0021] The present invention is further improved. In step 802, the absolute position encoding encodes the two-dimensional coordinates in the way of using sine and cosine functions, and the range of the output value is from -1 to 1.

[0022] The present invention is further improved. In step 802, the IOU (Intersection over Union) is used as the loss function of the object detection network to simplify the design and training speed of the object detection network.

[0023] The present invention is further improved. In step 6, the image reconstruction network includes a generator network and a discriminator network. The downsampling layer of the generator network uses dilated convolution, and the upsampling uses transposed convolution. The discriminator network consists of 5 convolutional layers, and some of the convolutional layers are densely connected.

[0024] The present invention is further improved. In step 7, during the process of training the image reconstruction network, the Dropout mechanism is added to make some feature values randomly invalid.

[0025] The present invention is further improved. In step 9, the image reconstruction network is provided with a generator loss function Loss G and a discriminator loss function Loss2. Among them, the generator loss function Loss G includes a perceptual loss function Loss p, Loss function Loss1, background suppression generation function Loss ground , where w represents the weight of the corresponding loss value, w p +w1+w g =1, Loss G =w p *Loss p +w1*Loss1+w g *Loss ground , the perceptual loss function Loss p is a feature-level loss function, the Loss1 loss function is a pixel-level error, and Loss ground is a loss function for sparse information images.

[0026] The present invention also provides a system for implementing the above-mentioned sparse photoacoustic image reconstruction method combined with object detection, including:

[0027] Training dataset construction module: used to construct a first training dataset for training the object detection network according to the original photoacoustic dataset, and used to construct a second training dataset for training the image reconstruction network according to the original photoacoustic dataset and the cropped image group dataset;

[0028] Network construction and training module: construct an object detection network based on the Two-stage network and the Transformers model, and train the object detection network using the first training dataset; construct an image reconstruction network based on the GAN network and the patchGan network, and train the image reconstruction network using the second training dataset;

[0029] Image cropping module: use the trained object detection network to process the sparsely sampled photoacoustic image at regular intervals to obtain a coordinate vector group dataset of the region of interest;

[0030] Image reconstruction module: use the trained image reconstruction network to reconstruct and output the final photoacoustic image according to the original image and the coordinate vector group dataset of the region of interest.

[0031] The beneficial effects of the present invention are: by combining an object detection network based on the Two-stage network and the Transformers model and an image reconstruction network based on the GAN network and the patchGan network, first use the trained object detection network to process the sparsely sampled photoacoustic image at regular intervals to obtain a coordinate vector group dataset of the region of interest, and then use the trained image reconstruction network to reconstruct and output the final photoacoustic image according to the original image and the coordinate vector group dataset of the region of interest, which not only saves a large amount of computing overhead, but also has good image restoration effect and fast image reconstruction speed, and solves the problem of poor image restoration effect caused by sparse photoacoustic image information and excessive background invalid information. Description of the Drawings

[0032] Figure 1 This is a flowchart of a sparse photoacoustic image reconstruction method combined with object detection according to the present invention. Detailed implementation manners

[0033] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0034] Please refer to Figure 1 , a sparse photoacoustic image reconstruction method combined with object detection according to the present invention includes the following steps:

[0035] Step 1: Use the original photoacoustic dataset to construct a first training dataset for training an object detection network;

[0036] Step 2: Construct an object detection network based on a two-stage network and a Transformers model;

[0037] Step 3: Train the object detection network using the first training dataset;

[0038] Step 4: Use the trained object detection network to process the original photoacoustic dataset to obtain a cropped image group dataset;

[0039] Step 5: Use the original photoacoustic dataset and the cropped image group dataset to construct a second training dataset for training an image reconstruction network;

[0040] Step 6: Construct an image reconstruction network based on a GAN network and a PatchGan network;

[0041] Step 7: Train the image reconstruction network using the second training dataset;

[0042] Step 8: Use the trained object detection network to process the sparsely sampled photoacoustic image at every other point to obtain a coordinate vector group dataset of the region of interest;

[0043] Step 9: Use the trained image reconstruction network to reconstruct and output the final photoacoustic image according to the original image and the coordinate vector group dataset of the region of interest.

[0044] In this embodiment, a function interface is used for communication between the target detection network and the image reconstruction network. During the training process of the networks, the target detection network is trained separately using the first training dataset, while the image reconstruction network needs the assistance of the already trained target detection network to construct the second training dataset to complete the training. The low-resolution images used in this embodiment are obtained by sampling high-resolution images at every other point, truly simulating the sampling at every other point in photoacoustic microscopy imaging. In this embodiment, the sparse photoacoustic images obtained by sampling at every other point are input into the target detection network to obtain the coordinates of the region of interest. The region of interest images are intercepted from the images according to the coordinates and packed and sent into the image reconstruction network to reconstruct and output the final photoacoustic image.

[0045] The target detection network constructed in this embodiment is based on the Two-stage network and the Transformers model. Among them, the Two-stage network mainly completes the target detection process through a convolutional neural network, and the CNN convolutional features it extracts. When training the network, it mainly trains two parts. The first step is to train the RPN network, and the second step is to train the network for detecting the target region. The accuracy of the network is high, but the speed is relatively slower than that of the One-stage network. The Transformers model is the most advanced model (BERT, GPT-2, RoBERTa, XLM, DistilBert, XLNet, CTRL...) for natural language understanding (NLU) and natural language generation (NLG) provided by the latest natural language processing libraries Transformers (formerly known as pytorch-transformers and pytorch-pretrained-bert) of TensorFlow 2.0 and PyTorch. It has more than 32 pre-trained models, supports more than 100 languages, and has deep interoperability between TensorFlow 2.0 and PyTorch. In this embodiment, the global characteristics of transformers are more suitable for the sparsity of photoacoustic images.

[0046] The image reconstruction network constructed in this embodiment is based on the GAN network and the PatchGAN network. Among them, the GAN network structure can better reconstruct the sparsely sampled photoacoustic images. PatchGAN, the difference between PatchGAN and the general GAN mainly lies in the Discriminator. In the general GAN, only a vector of true or false needs to be output, which represents the evaluation of the entire image. However, PatchGAN outputs an NxN matrix. Each element of this NxN matrix, such as a(i,j), has only two choices: True or False (the label is an NxN matrix, and each element is True or False). Such a result is often achieved through convolutional layers because the NxN matrix finally output by successive convolutional layers, each element of which actually represents a relatively large receptive field in the original image, that is, it corresponds to a Patch in the original image. Therefore, the GAN with such a structure and such an output is called Patch GAN.

[0047] Please refer to Figure 1 , in step 1, the original photoacoustic dataset is labeled using the labeling tool labelimg to construct the first training dataset for training the object detection network. Among them, Labelimg is a tool for labeling datasets used in deep network training. In this embodiment, using this tool to label and construct the dataset can better complete the training of the object detection network.

[0048] Please refer to Figure 1 , in step 8, the steps for the object detection network to process the sparsely sampled photoacoustic images with skip-point sampling to obtain the coordinate vector group dataset of the region of interest are as follows:

[0049] Step 801: Coarse detection, using 3 convolutional layers and 1 fully connected layer for shallow feature extraction, mainly aiming to find the possible regions of interest.

[0050] Step 802: Fine detection. Use the threshold method to select grids with a confidence level greater than the set threshold. Merge the selected grids with the absolute position encoding to form a new vector group and send it into the Transformer network. Since the Transformer encoding and decoder structures have a fixed number of input and output vectors, according to the median of the number of regions of interest in the dataset used by this system, this part of the Transformer network includes three networks with different vector sizes (the median of the number of regions of interest in this dataset is 25, so the input and output vector numbers are set to 15 / 25 / 35) to improve the network's operation speed. The network output of the Transformer finally outputs the target detection prediction box through the fully connected layer. The output structure is: prediction box type / confidence level / position encoding. Filter out vectors whose prediction box type is classified as the target type and whose confidence level is greater than the threshold, and obtain their coordinates through position decoding. The final output is the coordinate vector group dataset of the regions of interest.

[0051] Please refer to Figure 1 , in step 801, the size of the convolutional layer is determined according to the information distribution of the sparse photoacoustic image. In this embodiment, the size of the cut image of the photoacoustic image is 256*256, and the size of the sparsely distributed regions of interest in the image is about 32*32. Therefore, the convolutional kernel size is designed to be 5*5. Using an appropriate padding method, after passing through the fully connected layer, the size of the shallow features of the image finally extracted is 8*8. The output dimension represents how many grids the image is divided into, and the output value is the confidence level of the region of interest of the grid here. The higher it is, the richer the information here.

[0052] Please refer to Figure 1 , in step 802, the absolute position encoding encodes the two-dimensional coordinates in the way of using sine and cosine functions. It can be understood that the abscissa is mapped into the sine function and the ordinate is mapped into the cosine function. The default angle of the trigonometric function is from 0 to 2π. In the final output of the position encoding, the range of the output value is from -1 to 1.

[0053] Please refer to Figure 1 , in step 802, the Intersection over Union (IOU) is used as the loss function of the target detection network to simplify the design and training speed of the target detection network. The vector group output by the Transformer obtains the set of prediction squares after passing through the position decoder. To solve the problem of misalignment of vectors in the label set and the prediction set, the Hungarian algorithm is introduced in this embodiment to obtain the combination method of the minimum interpolation between the two sets. In the overall network, the target detection network is used as the pre-network, and its main role is to select the regions of interest. Therefore, the network only uses the IOU (Intersection over Union) as the loss function to simplify the design and training speed of the network. First, obtain the best set mapping through the Hungarian algorithm, and then calculate the IOU loss.

[0054] Please refer to Figure 1 , in step 6, the image reconstruction network includes a generator network and a discriminator network. The downsampling layer of the generator network uses dilated convolution, and the upsampling uses transposed convolution. The discriminator network consists of 5 convolutional layers, and some of the convolutional layers are densely connected. Among them, the number of channels of each layer of features of the generator network needs to be designed according to the image size and complexity, and the size of the convolutional layers in the network needs to be designed according to the size of the output feature layer and the output feature layer. Due to the autoencoder structure of the generator, its input size is the same as the output size. The input sparse sampled image needs to be filled with background color pixels to a size consistent with the target output size. For example, for a 16*16 sparse sampled image LR with a reconstruction target of a 64*64 high sampling rate image HR, 15 background color pixels need to be inserted outside each pixel of LR to obtain a 64*64 incomplete image that simulates real sparse (interval) sampling. Therefore, the input of its generator is actually an incomplete image that simulates real sparse sampling, and the output is a completed high sampling rate image. The discriminator network is a patchGAN network. The idea of patchGAN is reflected in that the output is not a value, but a matrix. In the sense of an image, it is equivalent to evenly cutting the reconstructed image and the sample image into several small pieces and evaluating the image similarity between each corresponding small piece one by one.

[0055] Please refer to Figure 1 , in step 7, during the process of training the image reconstruction network, the Dropout mechanism is added to randomly invalidate some feature values to ensure the stability of the image reconstruction network.

[0056] Please refer to Figure 1 , in step 9, the image reconstruction network is provided with a generator loss function Loss G and a discriminator loss function Loss2. Among them, the generator loss function Loss G includes a perceptual loss function Loss p , a loss function Loss1, and a background generation suppression function Loss ground . w represents the weight of the corresponding loss value, w p +w1+w g =1, Loss G =w p *Loss p +w1*Lossl+w g *Loss ground . The perceptual loss function Loss p is a feature-level loss function, the Lossl loss function is a pixel-level error, and Loss ground is a loss function for sparse information images.

[0057] Among them, the perceptual loss function Loss p is a feature-level loss function. The reconstructed image and the sample image distribution are fed into the Vgg19 network, and the features of its 16th layer are extracted in this network for comparison: Loss p = |vgg19(Y) j - vgg19(X) j | 2 / C j H j W j ; where Y represents the original sample image, X represents the generated reconstructed image, j represents which layer of features to extract, C j H j W j represents the size of the features of the jth layer. The vgg19 network uses the network structure in the vgg paper, and the vgg19 network is trained using the publicly available dataset cifar10.

[0058] The Loss1 loss function is the pixel-level error, which is directly the pixel difference between the reconstructed image and the original sample image.

[0059] Loss ground is the loss function for sparse information images. By statistically analyzing the sparse photoacoustic image group used in this embodiment, the average ratio of effective information to invalid background in its images is 2:8. Therefore, in the training of the network, Loss ground is added to constrain the ratio of generated effective information to invalid background, preventing the generator from tending to generate invalid background information to deceive the discriminator and resulting in a poor network training effect.

[0060] In the discriminator loss function calculation of Loss2, the discriminator calculates the similarity matrix Mar sH between the reconstructed image and the original sample image, and the similarity matrix Mat LH between the sparse image and the original sample image respectively. Loss G = |Mat SH - Mat LH | 2 .

[0061] Please refer to Figure 1 , the present invention also provides a system for implementing the above sparse photoacoustic image reconstruction method combined with object detection, including:

[0062] A training dataset construction module: used to construct a first training dataset for training the object detection network according to the original photoacoustic dataset, and used to construct a second training dataset for training the image reconstruction network according to the original photoacoustic dataset and the cropped image group dataset;

[0063] Network construction and training module: Construct an object detection network based on the Two-stage network and the Transformers model, and use the first training dataset to train the object detection network; construct an image reconstruction network based on the GAN network and the patchGan network, and use the second training dataset to train the image reconstruction network.

[0064] Image cropping module: Use the trained object detection network to process the sparse photoacoustic image obtained by dot-interleaved sampling, and obtain a dataset of coordinate vector groups of the region of interest.

[0065] Image reconstruction module: Use the trained image reconstruction network to reconstruct and output the final photoacoustic image according to the original image and the dataset of coordinate vector groups of the region of interest.

[0066] As can be seen from the above, the beneficial effects of the present invention are as follows: By combining an object detection network based on the Two-stage network and the Transformers model and an image reconstruction network based on the GAN network and the patchGan network, first use the trained object detection network to process the sparse photoacoustic image obtained by dot-interleaved sampling to obtain a dataset of coordinate vector groups of the region of interest, and then use the trained image reconstruction network to reconstruct and output the final photoacoustic image according to the original image and the dataset of coordinate vector groups of the region of interest. This not only saves a large amount of computational overhead, but also has good image restoration effect and fast image reconstruction speed, and solves the problem of poor image restoration effect caused by sparse photoacoustic image information and excessive background invalid information.

[0067] The above specific implementation manners are the preferred implementation manners of the present invention, and do not limit the specific implementation scope of the present invention. The scope of the present invention includes but is not limited to this specific implementation manner. All equivalent changes made in accordance with the present invention are within the protection scope of the present invention.

Claims

1. A sparse photoacoustic image reconstruction method combined with object detection, characterized in that Including the following steps: Step 1: Use the original photoacoustic dataset to construct a first training dataset for training the object detection network; Step 2: Construct an object detection network based on the Two-stage network and the Transformers model; Step 3: Train the object detection network using the first training dataset; Step 4: Use the trained object detection network to process the original photoacoustic dataset to obtain a cropped image group dataset; Step 5: Use the original photoacoustic dataset and the cropped image group dataset to construct a second training dataset for training the image reconstruction network; Step 6: Construct an image reconstruction network based on the GAN network and the patchGan network; Step 7: Train the image reconstruction network using the second training dataset; Step 8: Use the trained object detection network to process the sparsely sampled photoacoustic image at every other point to obtain a coordinate vector group dataset of the region of interest; Step 9: Use the trained image reconstruction network to reconstruct and output the final photoacoustic image according to the original image and the coordinate vector group dataset of the region of interest.

2. The sparse photoacoustic image reconstruction method combined with target detection according to claim 1, characterized in that, In the said Step 1, use the annotation tool labelimg to annotate the original photoacoustic dataset and construct the first training dataset for training the object detection network.

3. The sparse photoacoustic image reconstruction method combined with object detection according to claim 2, characterized in that, In the said Step 8, the object detection network processes the sparsely sampled photoacoustic image at every other point to obtain a coordinate vector group dataset of the region of interest, including the following steps: Step 801: Coarse detection, use 3 convolutional layers and 1 fully connected layer for shallow feature extraction; Step 802: Fine detection, use the threshold method to select the grids with a confidence level greater than the set threshold, merge the selected grids with absolute position encoding to form a new vector group, and send it into the transformer network.

4. The sparse photoacoustic image reconstruction method combined with target detection according to claim 3, characterized in that, In the said Step 801, the size of the convolutional layer is determined according to the information distribution of the sparsely sampled photoacoustic image.

5. The sparse photoacoustic image reconstruction method combined with target detection according to claim 3, wherein, In the said Step 802, the absolute position encoding encodes the two-dimensional coordinates in the way of using sine and cosine functions, and the output value range is from -1 to 1.

6. The sparse photoacoustic image reconstruction method combined with object detection according to claim 5, characterized in that, In the said Step 802, use IOU (Intersection over Union) as the loss function of the object detection network to simplify the design and training speed of the object detection network.

7. The sparse photoacoustic image reconstruction method combined with target detection according to claim 6, characterized in that In the said Step 6, the image reconstruction network includes a generator network and a discriminator network. The downsampling layer of the generator network uses dilated convolution, and the upsampling uses transposed convolution. The discriminator network consists of 5 convolutional layers, and there are partial dense connections between the convolutional layers.

8. The sparse photoacoustic image reconstruction method combined with target detection according to claim 7, characterized in that In the said Step 7, during the process of training the image reconstruction network, add the Dropout mechanism to randomly invalidate some feature values.

9. The sparse photoacoustic image reconstruction method combined with target detection according to claim 8, wherein, In the step 9, a generator loss function Loss is provided in the image reconstruction network G and a discriminator loss function Loss2. Among them, the generator loss function Loss G includes a perceptual loss function Loss p , a loss function Loss1, and a background generation suppression function Loss ground . w represents the weight of the corresponding loss value. w p +w1+w g =1. Loss G =w p *Loss p +w1*Loss1+w g *Loss ground . The perceptual loss function Loss p is a feature-level loss function. The Loss1 loss function is a pixel-level error. Loss ground is a loss function for sparse information images.

10. A system for implementing the sparse photoacoustic image reconstruction method combined with object detection according to any one of claims 1-9, characterized in that, Including: A training dataset construction module: used to construct a first training dataset for training the object detection network according to the original photoacoustic dataset, and used to construct a second training dataset for training the image reconstruction network according to the original photoacoustic dataset and the cropped image group dataset; Network construction and training module: Construct an object detection network based on the Two-stage network and the Transformers model, and train the object detection network using the first training dataset; construct an image reconstruction network based on the GAN network and the patchGan network, and train the image reconstruction network using the second training dataset; Image cropping module: Use the trained object detection network to process the sparse photoacoustic images obtained by dot-sampling, and obtain a dataset of coordinate vector groups of the regions of interest; Image reconstruction module: Use the trained image reconstruction network to reconstruct and output the final photoacoustic image according to the original image and the dataset of coordinate vector groups of the regions of interest.

Citation Information

Patent Citations

  • Remote sensing image target detection method for constructing convolutional neural network based on pruning strategy

    CN107609525A

  • Tumor photoacoustic image rapid reconstruction method and device based on deep learning

    CN110880196A