Photoacoustic microscopic imaging noise reduction method and system based on deep learning
By using a deep learning-based 3D U-Net network and a spatial consistency attention module, the problem of noise influence in photoacoustic microscopy was solved, generating high-quality photoacoustic microscopic images and improving the signal-to-noise ratio and vascular continuity.
Patent Information
- Application Number
- CN202511767034.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-03-17
Smart Images

Figure CN121685304A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a photoacoustic microscopy imaging noise reduction method and system, and in particular to a photoacoustic microscopy imaging noise reduction method and system based on deep learning. Background Technology
[0002] This section provides only background information relevant to this disclosure and is not necessarily prior art.
[0003] Photoacoustic microscopy (PAM) is an emerging biomedical imaging modality that combines high optical contrast with high ultrasound resolution. It is widely used in fields such as microvascular imaging, tumor boundary identification, and brain function research. PAM reconstructs a three-dimensional image reflecting optical absorption characteristics by detecting the ultrasound signal generated after tissue absorbs pulsed laser light. It can achieve functional imaging with micron-level resolution without the need for external labeling, and has significant scientific and clinical value.
[0004] However, due to limitations such as laser energy safety thresholds, ultrasound transducer sensitivity, and tissue scattering noise, raw PAM images often suffer from low signal-to-noise ratios, structural blurring, and peak location drift. This is especially true in deep tissues or low-perfusion areas, where noise severely disrupts vascular continuity and spatial structural integrity, affecting the accuracy of subsequent quantitative analysis and clinical diagnosis. Traditional denoising methods, such as median filtering, wavelet transform, and nonlocal means, can suppress noise to some extent, but often at the cost of spatial detail, making it difficult to balance denoising intensity and structural fidelity. In particular, they cannot restore the spatial correlation of the A-line signal—that is, the physical continuity of adjacent pixels in peak position and waveform morphology—which is disrupted by noise. In recent years, deep learning-based image denoising methods have made significant progress in natural and medical imaging. Two-dimensional convolutional neural networks (CNNs) and U-Net architectures have been used for PAM image enhancement, but most schemes only process single frames or maximum projection images, ignoring the spatial contextual information of three-dimensional volume data.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] Purpose of the invention: The technical problem to be solved by the present invention is to provide a method and system for noise reduction of photoacoustic microscopy imaging based on deep learning, which addresses the shortcomings of the existing technology.
[0007] To address the aforementioned technical problems, this invention discloses a deep learning-based photoacoustic microscopy noise reduction method and system, the method comprising the following steps:
[0008] Step 1: Perform block preprocessing on the input raw photoacoustic microscopic imaging data;
[0009] Step 2: Construct and train the denoising network, namely a 3D U-Net network with spatial structure awareness. After the denoising network is trained, input the preprocessed sub-block to be denoised and output the denoising result.
[0010] Step 3: Input the noise reduction result into the data fusion projection layer, perform weighted stitching and boundary fusion on all noise reduction sub-blocks, and perform maximum intensity projection along the depth direction to generate the final high-quality photoacoustic microscopic image.
[0011] Furthermore, the preprocessing described in step 1 includes:
[0012] Step 1-1: Normalize the amplitude of the raw volume data to eliminate signal strength differences between acquisition devices;
[0013] Steps 1-2 involve dividing the normalized volumetric data into three-dimensional sub-blocks with overlapping regions. The sub-block sequence is represented as follows:
[0014] {B1, B2, ..., B K}
[0015] Where K is the total number of sub-blocks, This represents the k-th 3D sub-block, where N is the horizontal dimension and M is the depth dimension.
[0016] Furthermore, the 3D U-Net network described in step 2 includes:
[0017] Encoder, decoder, skip connections, and spatial consistency attention module; among which,
[0018] The encoder extracts deep semantic features through multi-level 3D convolution and downsampling;
[0019] The decoder gradually restores spatial resolution through 3D deconvolution and upsampling;
[0020] The skip connection concatenates and fuses the encoder's corresponding level features with the decoder's upsampled features.
[0021] The spatial consistency attention module, embedded after the skip connection, is used to model the spatial continuity of the A-line signals of adjacent pixels in terms of peak position and intensity.
[0022] Furthermore, the spatial consistency attention module includes:
[0023] Query branch, key branch, value branch, and residual output unit; among which,
[0024] The query branch and the key branch are respectively compressed through 1×1×1 convolution and global average pooling along the depth direction to obtain the XY plane semantic response map;
[0025] The value branch retains the original three-dimensional features for weighted aggregation;
[0026] The residual output unit calculates the Query-Key similarity matrix, normalizes it using SoftMax, applies it to the Value branch, and outputs it in residual form.
[0027] Y = γ·Attention(X) + X
[0028] Where γ is the learnable scaling parameter, initialized to 0, X is the input feature, and Y is the output feature.
[0029] Furthermore, the calculation process of Attention(X) includes:
[0030] First, flatten the two-dimensional spatial feature maps output by the query branch and the key branch into matrix form, so that each spatial location corresponds to a feature vector;
[0031] Second, calculate the dot product similarity between all spatial locations and normalize it using the SoffMax function to obtain the spatial attention weight matrix, which represents the semantic relevance of each location to other locations.
[0032] Third, rearrange the three-dimensional features output by the value branch according to their spatial positions, and perform weighted aggregation of the features at each position according to the attention weight matrix to generate enhanced spatial perception features.
[0033] Fourth, the enhanced features are restored to the original 3D structure and added to the input features to form the final output with residual connections.
[0034] Furthermore, the training of the 3D U-Net network in step 2 employs a composite loss function, defined as:
[0035]
[0036] in, For pixel-level fidelity loss, For structural similarity loss, λ1 and λ2 are adjustable weight coefficients for spatial consistency regularization.
[0037] The pixel-level fidelity loss Using the L1 norm, it is represented as follows:
[0038]
[0039] Among them, I out Igt is the output volume data of the network, N is the clear reference volume data, and N is the total number of voxels.
[0040] Structural similarity loss Defined as:
[0041]
[0042] Where d, h, and w represent the discrete spatial coordinate indices of the 3D volume data in the depth, height, and width directions, respectively; I represents the denoised volume data output by the network, and J represents the corresponding sharp reference volume data; μ I μ J The local mean of I and J is given within a local 3D Gaussian weighted window centered at position (d, h, w). For the corresponding local variance; σ IJ Let C1 be the covariance of I and J within this local window; C1 = (0.01L) 2 C2 = (0.03L) 2 L is the dynamic range of the image (e.g., L = 1 after normalization); M is the total number of center points of the sliding window;
[0043] Spatial consistency regularization Defined as:
[0044]
[0045] Where z(x,y) represents the peak depth of the A-line signal at position (x,y), and Ω is the effective pixel region participating in the calculation, which is determined in the following way: the volume data is projected along the depth direction with maximum intensity to obtain a two-dimensional intensity map and normalized, and the set of pixels above the preset threshold is selected to form Ω.
[0046] Furthermore, the weighted stitching and boundary fusion described in step 3 includes:
[0047] Step 3-1: For adjacent 3D sub-blocks in the overlapping region, a weighted average fusion is performed using spatial attenuation weights. The weighting function is defined as:
[0048]
[0049] Where (x0, y0, z0) are the coordinates of the current sub-block center, and σ is the standard deviation parameter that controls the decay rate.
[0050] Step 3-2: Apply gradient domain smoothing filter to the block boundary region of the stitched complete 3D volume data. The goal is to minimize the gradient difference between the output image and the ideal continuous image in the boundary neighborhood, so as to further eliminate residual stitching traces.
[0051] This invention also proposes a deep learning-based three-dimensional photoacoustic microscopy noise reduction system for implementing the aforementioned method, comprising:
[0052] The data consists of a preprocessing layer, a network noise reduction layer, and a data fusion and projection layer; among which...
[0053] The data preprocessing layer described above preprocesses the original three-dimensional photoacoustic volume data;
[0054] The aforementioned network noise reduction layer performs noise reduction prediction on the preprocessed photoacoustic data;
[0055] The data fusion projection layer stitches and fuses the noise-reduced sub-blocks to generate a maximum intensity projection (MAP) image.
[0056] Beneficial effects:
[0057] 1. This invention significantly improves the signal-to-noise ratio of PAM imaging by introducing a 3D U-Net and a block training mechanism to learn the spatial continuity of A-line signals of adjacent pixels in photoacoustic imaging.
[0058] 2. This invention enhances the network's ability to learn the distribution of background noise through a spatial consistency attention mechanism, reduces the interference of random noise on local structure reconstruction in a single scan, and the output image is superior to traditional noise reduction methods in terms of vascular continuity, boundary sharpness and peak localization accuracy, making it suitable for various tissue types and clinical equipment deployment scenarios. Attached Figure Description
[0059] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.
[0060] Figure 1 This is a schematic diagram of the overall architecture of the present invention.
[0061] Figure 2 This is a schematic diagram of the image noise reduction model structure of the present invention. Detailed Implementation
[0062] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0063] This invention discloses a method and system for noise reduction in photoacoustic microscopy imaging based on deep learning.
[0064] like Figure 1 As shown, the system proposed in this invention adopts a layered architecture design, including a data preprocessing layer, a network noise reduction layer, and a data fusion projection layer.
[0065] like Figure 2 As shown, the three-dimensional photoacoustic microscopy imaging denoising model proposed in this invention (hereinafter referred to as the model) is an image denoising method applied to photoacoustic microscopy, comprising the following steps:
[0066] Step 1: Perform block preprocessing on the input raw photoacoustic microscopic imaging data;
[0067] Step 2: Construct and train the denoising network, namely a 3D U-Net network with spatial structure awareness. After the denoising network is trained, input the preprocessed sub-block to be denoised and output the denoising result.
[0068] Step 3: Input the noise reduction result into the data fusion projection layer, perform weighted stitching and boundary fusion on all noise reduction sub-blocks, and perform maximum intensity projection along the depth direction to generate the final high-quality photoacoustic microscopic image.
[0069] In the deep learning-based photoacoustic microscopy noise reduction method and system described in this embodiment, step 1 includes:
[0070] Step 1-1: Normalize the amplitude of the original volume data to [0, 1] to eliminate the signal strength differences between the acquisition devices;
[0071] Steps 1-2 involve dividing the normalized volumetric data into three-dimensional sub-blocks with overlapping regions. The sub-block sequence is represented as follows:
[0072] {B1, B 2,.. ., B K}
[0073] Where K is the total number of sub-blocks, This represents the k-th 3D sub-block, where N is the horizontal dimension and M is the depth dimension. The total number of sub-blocks is determined by the original volume data dimension and the overlap ratio (50%).
[0074] In the deep learning-based photoacoustic microscopy noise reduction method and system described in this embodiment, step 2 includes:
[0075] like Figure 2 As shown, a 3D U-Net network with spatial structure awareness is constructed and trained. The 3D U-Net network mainly consists of an encoder, a decoder, skip connections, and a spatial consistency attention module.
[0076] The encoder receives preprocessed sub-block tensors of shape (batch_size, channels, height, width, depth), which are then expanded to include three levels of 3D convolutional and pooling layers (1→32→64→128). The encoder includes ReLU activation function and batch normalization layer, which can effectively improve the model's expressive power and prevent overfitting.
[0077] The decoder reduces the number of feature map channels (128→64→32) through two levels of three-dimensional deconvolution layers and upsampling layers, and then maps the 32 channels back to a single channel output by a 1×1×1 convolution, without activation functions to preserve the original amplitude range;
[0078] In the skip connection part, the feature map output by each level of the encoder (32 and 64 channels) is spliced and fused with the feature map after upsampling of the corresponding level of the decoder along the channel dimension to form a multi-scale feature combination, which effectively makes up for the loss of spatial details caused by downsampling, so that the network can retain the original structural information while restoring the resolution.
[0079] The spatial consistency attention module, embedded after each skip connection stitching operation and before the decoder convolutional block processing, is used to explicitly model the spatial continuity of the A-line signals of adjacent pixels in terms of peak position and intensity. This module receives the stitched fused features (e.g., 128 or 64 channels) and achieves spatial awareness enhancement through the following structure:
[0080] Query branch, key branch, value branch, and residual output unit; among which,
[0081] The query branch and the key branch are respectively compressed through 1×1×1 convolution and global average pooling along the depth direction to obtain the XY plane semantic response map;
[0082] The value branch retains the original three-dimensional features for weighted aggregation;
[0083] The residual output unit calculates the Query-Key similarity matrix, normalizes it using SoftMax, applies it to the Value branch, and outputs it in residual form.
[0084] Y = γ·Attention(X) + X
[0085] Where γ is a learnable scaling parameter, initialized to 0, X is the input feature, and Y is the output feature. The specific calculation process of Attention(X) includes:
[0086] First, flatten the two-dimensional spatial feature maps output by the query branch and the key branch into matrix form, so that each spatial location corresponds to a feature vector;
[0087] Second, calculate the dot product similarity between all spatial locations and normalize it using the SoftMax function to obtain the spatial attention weight matrix, which represents the semantic relevance of each location to other locations.
[0088] Third, rearrange the three-dimensional features output by the value branch according to their spatial positions, and perform weighted aggregation of the features at each position according to the attention weight matrix to generate enhanced spatial perception features.
[0089] Fourth, the enhanced features are restored to the original 3D structure and added to the input features to form the final output with residual connections.
[0090] The training of the 3D U-Net network uses a composite loss function, defined as:
[0091]
[0092] in, This is a pixel-level fidelity loss used to constrain the point-by-point consistency of the noise reduction output with the reference image in voxel intensity, ensuring the fidelity of the underlying signal. Structural similarity loss is used to maintain the consistency of local texture, contrast, and structural distribution, thereby improving visual perception quality. λ1 and λ2 are spatial consistency regularization terms used to force adjacent pixel A-line signals to maintain spatial continuity at peak depth, suppressing artifacts such as blood vessel rupture or depth jump caused by noise; λ1 and λ2 are adjustable weighting coefficients (default values are 0.5 and 0.3, respectively).
[0093] The pixel-level fidelity loss Using the L1 norm, it is represented as follows:
[0094]
[0095] Among them, I out For network output volume data, I gt For clear reference data, N is the total number of voxels;
[0096] Structural similarity loss Defined as:
[0097]
[0098] Where d, h, and w represent the discrete spatial coordinate indices of the 3D volume data in the depth, height, and width directions, respectively; I represents the denoised volume data output by the network, and J represents the corresponding sharp reference volume data; μ I μ J The local mean of I and J is given within a local 3D Gaussian weighted window centered at position (d, h, w). For the corresponding local variance; σ IJ Let C1 be the covariance of I and J within this local window; C1 = (0.01L) 2 C2 = (0.03L) 2 L is the dynamic range of the image (e.g., L = 1 after normalization); M is the total number of center points of the sliding window;
[0099] Spatial consistency regularization Defined as:
[0100]
[0101] Where z(x, y) represents the peak depth of the A-line signal at position (x, y), and Ω is the effective pixel region participating in the calculation, which is determined in the following way: the volume data is projected along the depth direction with maximum intensity to obtain a two-dimensional intensity map and normalize it, and the set of pixels above the preset threshold (0.3) is selected to form Ω.
[0102] In the deep learning-based photoacoustic microscopy noise reduction method and system described in this embodiment, step 3 includes:
[0103] Step 3-1: For adjacent 3D sub-blocks in the overlapping region, a weighted average fusion is performed using spatial attenuation weights. The weighting function is defined as:
[0104]
[0105] Where (x0, y0, z0) are the coordinates of the current sub-block center, and σ is the standard deviation parameter that controls the decay rate (default is 5).
[0106] Step 3-2: Apply gradient domain smoothing filter to the block boundary region of the stitched complete 3D volume data. The goal is to minimize the gradient difference between the output image and the ideal continuous image in the boundary neighborhood, so as to further eliminate residual stitching traces.
[0107] After stitching, the volume data is subjected to maximum intensity projection (MAP) along the depth direction to generate the final high-quality photoacoustic micrograph.
[0108] In its specific implementation, this invention provides a computer storage medium and a corresponding data processing unit. The computer storage medium is capable of storing a computer program, which, when executed by the data processing unit, can run the invention's content regarding a deep learning-based photoacoustic microscopy noise reduction method and system, as well as some or all of the steps in various embodiments. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0109] Those skilled in the art will clearly understand that the technical solutions in the embodiments of the present invention can be implemented using computer programs and their corresponding general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of computer programs, i.e., software products. These computer program software products can be stored in a storage medium and include several instructions to cause a device containing a data processing unit (which may be a personal computer, server, microcontroller, MCU, or network device, etc.) to execute the methods described in various embodiments or certain parts of the embodiments of the present invention.
[0110] This invention provides a method and system for noise reduction in photoacoustic microscopy imaging based on deep learning. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.
Claims
1. A deep learning based denoising method for three-dimensional photoacoustic microscopic imaging, characterized in that, The method comprises the following steps: Step 1, preprocessing the input raw photoacoustic microscopic imaging volume data by blocking; Step 2, constructing and training a denoising network, i.e. a 3D U-Net network with spatial structure perception capability, after the training of the denoising network, inputting the preprocessed sub-blocks to be denoised and outputting denoising results; Step 3, inputting the denoising results into a data fusion projection layer, performing weighted splicing and boundary fusion on all denoising sub-blocks, and performing maximum intensity projection along the depth direction to generate a final high-quality photoacoustic microscopic image.
2. The three-dimensional photoacoustic microscopic imaging denoising method based on deep learning according to claim 1, characterized in that, The preprocessing in step 1 comprises: Step 1-1, amplitude normalization of the original volume data to eliminate signal intensity differences between acquisition devices; Step 1-2, dividing the normalized volume data into three-dimensional sub-blocks with overlapping areas, and the sub-block sequence is represented as: {B1, B2,..., B K} where K is the total number of sub-blocks, denotes the k-th three-dimensional sub-block, N is the lateral dimension, and M is the depth dimension.
3. The three-dimensional photoacoustic microscopic imaging denoising method based on deep learning according to claim 1, characterized in that, The 3D U-Net network in step 2 comprises: An encoder, a decoder, a skip connection and a spatial consistency attention module; wherein, The encoder extracts deep semantic features through multi-level 3D convolution and downsampling; The decoder gradually restores the spatial resolution through 3D deconvolution and upsampling; The skip connection splices and fuses the corresponding level features of the encoder and the upsampling features of the decoder; The spatial consistency attention module is embedded after the skip connection and is used to model the spatial continuity of adjacent pixel A-line signals in peak position and intensity.
4. The three-dimensional photoacoustic microscopic imaging denoising method based on deep learning according to claim 3, characterized in that, The spatial consistency attention module comprises: A query branch, a key branch, a value branch and a residual output unit; wherein, The query branch and the key branch respectively compress the channels through 1x1x1 convolution and globally average pool along the depth direction to obtain XY plane semantic response maps; The value branch retains the original three-dimensional features for weighted aggregation; The residual output unit calculates the Query-Key similarity matrix, which is normalized by SoftMax and then acts on the Value branch, and outputs in the form of residual: Y=y·Attention(X)+X Wherein, γ is a learnable scaling parameter, initialized to 0, X is the input feature, and Y is the output feature.
5. The three-dimensional photoacoustic microscopic imaging denoising method based on deep learning according to claim 4, characterized in that, The calculation process of Attention(X) comprises: Firstly, flatten the two-dimensional spatial feature maps output by the query branch and the key branch into matrix form respectively, so that each spatial position corresponds to a feature vector; Secondly, calculate the dot product similarity between all spatial positions, and normalize it by the SoftMax function to obtain a spatial attention weight matrix, which represents the semantic correlation of each position to other positions; Thirdly, rearrange the three-dimensional features output by the value branch according to the spatial position, and weight and aggregate the features of each position according to the attention weight matrix to generate enhanced spatial perception features; Fourthly, restore the enhanced features to the original three-dimensional structure and add them to the input features to form the final output with residual connection.
6. The three-dimensional photoacoustic microscopic imaging denoising method based on deep learning according to claim 1, characterized in that, The training of the 3D U-Net network in step 2 adopts a composite loss function, which is defined as: wherein, is a pixel-level fidelity loss, is a structural similarity loss, is a spatial consistency regularization term, and λ1, λ2are adjustable weight coefficients. The pixel-level fidelity loss With L1 norm, expressed as follows: where I out is the network output volume data, I gt is the ground truth volume data, and N is the total number of voxels. Structural similarity loss is defined as: where d, h, w represent the discrete spatial coordinate indexes of the three-dimensional data in the depth direction, height direction, and width direction, respectively; I represents the denoised volume data output by the network; J represents the corresponding clear reference volume data; μ I , μ J is the local mean of I and J in the local 3D Gaussian weighted window centered at position (d, h, w); is the corresponding local variance; σ IJ is the covariance of I and J in the local window; C1 = (0.01L) 2 , C2 = (0.03L) 2 , L is the image dynamic range (such as L = 1 after normalization); M is the total number of sliding window center points; spatial coherence regularizer is defined as: Wherein z(x, y) represents the peak depth of the A-line signal at position (x, y), and Ω is an effective pixel region participating in the calculation, which is determined by the following method: the body data is maximum intensity projected along the depth direction to obtain a two-dimensional intensity map and normalized, and a pixel set higher than a preset threshold is selected to form Ω.
7. The three-dimensional photoacoustic microscopic imaging denoising method based on deep learning according to claim 1, characterized in that, The weighted splicing and boundary fusion in step 3 comprises: Step 3-1, the spatial attenuation weight is used for weighted average fusion of adjacent three-dimensional sub-blocks in the overlapping area, and the weight function is defined as: Wherein (x0, y0, z0) is the center coordinate of the current sub-block, and σ is a standard deviation parameter for controlling the attenuation speed. Step 3-2, gradient domain smoothing filtering is applied to the block boundary region of the complete three-dimensional body data after splicing, and the target is to minimize the gradient difference between the output image and the ideal continuous image in the boundary neighborhood, so as to further eliminate the residual splicing traces.
8. A deep learning based three-dimensional photoacoustic microscopic imaging denoising system, characterized in that, For implementing the method of claims 1-7, comprising: A data preprocessing layer, a network noise reduction layer, and a data fusion projection layer; wherein The data preprocessing layer is used for preprocessing the original three-dimensional photoacoustic body data; The network noise reduction layer is used for noise reduction prediction of the preprocessed photoacoustic body data; The data fusion projection layer is used for splicing and fusion of the noise reduction sub-blocks, and generates a maximum intensity projection (MAP) image.