Systems and methods for photoacoustic imaging in real-time

Deep learning enhances photoacoustic imaging by denoising and restoring tissue oxygenation information in real-time, addressing the limitations of LED arrays in delivering high fluence outputs and enabling efficient real-time imaging.

WO2026039549A1PCT designated stage Publication Date: 2026-02-19TRUSTEES OF TUFTS COLLEGE
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/041851
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-13
Filing Date
2025-08-13
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Conventional photoacoustic imaging systems using LED arrays face challenges in delivering high fluence outputs, leading to prolonged acquisition times and compromising real-time imaging capabilities, especially in dynamic biological processes.

Method used

Employing a deep learning approach to denoise PA images and restore physiologically accurate tissue oxygenation information, enabling real-time imaging without extensive averaging or external contrast agents.

Benefits of technology

The deep learning method achieves a signal-to-noise ratio comparable to high frame averaging, significantly improving acquisition speed and image quality in real-time photoacoustic imaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025041851_19022026_PF_FP_ABST
    Figure US2025041851_19022026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods relying on deep learning (DL) approaches are used to denoise photoacoustic images (PAI), including under extreme noise, as well as preserve and restore physiologically accurate tissue oxygenation information across pre-clinical and clinical settings. U-Net-based DL models are used to improve the signal-to-noise ratio (SNR) in real-time, including for low frame averaged PA data, matching the image quality of high frame averaged PA data.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR PHOTOACOUSTIC IMAGING IN REAL-TIME Cross Reference to Related Applications

[0001] The present application is based on, claims priority to, and incorporates herein by reference in its entirety for all purposes, US Provisional Application Serial No.63 / 682,335, filed August 13, 2024. Statement of Government Support

[0002] N / A Background

[0003] Photoacoustic imaging (PAI) is a hybrid imaging modality combining optical illumination and ultrasound (US) reception with good optical contrast and spatial resolution. Conventionally, PAI employs nanosecond pulse laser systems irradiating tissues at specific wavelengths tailored to their optical properties. However, the traditional reliance on costly Nd- YAG pumped optical parametric oscillator (OPO) lasers has posed challenges regarding mobility and cost-effectiveness. Recent strides have been made towards mitigating these limitations through the emergence of laser diode or light emitting diode (LED)-based illumination systems. Despite their advantages in terms of cost-effectiveness and portability, LED arrays face constraints (~ 400 ^^^^) in delivering high fluence outputs compared to lasers (~ 40 ^^^^), necessitating compensatory strategies such as high number of frame averaging. However, this leads to prolonged acquisition times, impeding real-time imaging crucial for understanding dynamic biological processes in vivo.

[0004] Therefore, there is a need to improve the acquisition speed in LED and other low excitation energy-based PA imaging systems without compromising the signal-to-noise ratio (SNR). Summary

[0005] The present disclosure provides systems and methods that overcome the aforementioned drawbacks by employing deep learning (DL) approach that not only denoises PA images under extreme noise but also preserves and restores physiologically accurate tissue oxygenationinformation across pre-clinical and clinical settings. The DL approach’s ability to function with real-time constraints, without reliance on extensive averaging or external contrast agents, provides a significant advancement in PAI, with broad implications for functional imaging, cancer diagnostics, vascular assessment, and more.

[0006] In one aspect of the present disclosure, a system for improving signal-to-noise ratio (SNR) in real-time in photoacoustic imaging (PAI) is described. The system comprises an excitation probe configured to deliver light to a tissue, an acoustic detection probe configured to receive photoacoustic signals at a first frame acquisition rate from the tissue in response to excitation of the tissue by the light and a processor. The processor is configured to receive the photoacoustic signals, reconstruct the photoacoustic signals by performing a low number frame averaging (LA) on the photoacoustic signals to produce LA image data, process the LA image data in real-time using a deep learning (DL) model, wherein the SNR of the processed LA image data is statistically identical to a SNR of a high number frame averaging (HA) of the photoacoustic signals acquired at a second frame acquisition rate, and generate one or more output images from the processed LA image data. The system further comprises a display configured to display the one or more output images.

[0007] In one aspect of the present disclosure, a system for monitoring oxygen saturation (StO2) in real-time using PAI is described. The system comprises an excitation probe configured to deliver light to a tissue, wherein the light includes two or more wavelengths, an acoustic detection probe configured to receive photoacoustic signals at a first frame acquisition rate from the tissue in response to excitation of the tissue by each of the two or more wavelengths, and a processor. The processor is configured to receive the photoacoustic signals from each of the two or more wavelengths, reconstruct the photoacoustic signals from each of the two or more wavelengths by performing a low number frame averaging (LA) on the photoacoustic signals to produce LA image data, process the LA image data from each of the two or more wavelengths in real-time using a deep learning (DL) model, wherein a SNR of the reconstructed images is statistically identical to a SNR of a high number frame averaging (HA) of the photoacoustic signals for each of the two or more wavelengths acquired at a second frame acquisition rate, generate one or more output images from the processed LA image data for each of the two or more wavelengths, perform a linear unmixing on the one or more output images of each of thetwo or more wavelengths, and generate a StO2 map from the linear unmixing of the output images. The system further comprises a display configured to display the StO2 map.

[0008] In one aspect of the present disclosure, a method for improving SNR in real-time in PAI is described. The method comprises delivering light to a tissue using an excitation probe, detecting photoacoustic signals using a detection probe at a first frame acquisition rate from the tissue in response to excitation of the tissue by the light, receiving the photoacoustic signals via a processor, reconstructing the photoacoustic signals by performing a low number frame averaging (LA) on the photoacoustic signals to produce LA image data, processing the LA image data in real-time using a deep learning (DL) model, wherein the SNR of the processed LA image data is statistically identical to a SNR of a high number frame averaging (HA) of the photoacoustic signals acquired at a second frame acquisition rate, generating one or more output images from the processed LA image data, and displaying, via a display, the one or more output images.

[0009] These aspects are nonlimiting. Other aspects and features of the systems and methods described herein will be provided below. Brief Description of the Drawings

[0010] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0011] The foregoing features of embodiments will be more readily understood by reference to the following detailed description, taken with reference to the accompanying drawings. The application file contains at least one drawing executed in color. Copies of the patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0012] FIG.1 is a non-limiting example schematic of a system for improving SNR in real-time in PAI, according to aspects of the present disclosure.

[0013] FIG.2 is a non-limiting example schematic of a system for monitoring oxygen saturation (StO2) in real-time using PAI, according to aspects of the present disclosure.

[0014] FIG.3A is a workflow of the Pix2Pix network-based conditional generative adversarial network (cGAN) denoising model (Den-P2P), according to aspects of the present disclosure.

[0015] FIG.3B is a schematic of the network architecture of the generator of the Den-P2P deep learning (DL) model.

[0016] FIG.3C is a schematic of the network architecture of the discriminator of the Den-P2P deep learning model.

[0017] FIG.4 is a schematic of the network architecture of the Conjoint Frequency & Intensity Regulated UNet (CONFIG-UNet) DL model, according to aspects of the present disclosure.

[0018] FIG.5 is a block diagram of a system for improving SNR in PAI in real-time in accordance with some implementations of the disclosed subject matter.

[0019] FIG.6 is a block diagram showing an example of hardware that can be used to implement the system of FIG.5 in accordance with some implementations of the disclosed subject matter.

[0020] FIG.7 is a method of improving SNR in real-time, according to aspects of the present disclosure.

[0021] FIG.8 shows US image (grayscale) and low number of frame average PA images of a mouse subcutaneous tumor and corresponding plots of SNR, Peak SNR (PSNR), and contrast-to- noise ratio (CNR). One-way ANOVA with Tukey's post hoc test (p < 0.05) been used for the statistical analysis of all the data. ****p < 0.0001 and ns: not significant.

[0022] FIG.9A is a comparison between non-learning traditional denoising algorithms with the cGAN DL (DenP2P) model.

[0023] FIG.9B is a comparison between two DL models: U-Net and cGAN denoising model. Images of the carbon fiber generated with one pulse illumination and its denoising with the U- Net and the cGAN model, respectively, are shown.

[0024] FIG.9C is a plot of the average PA amplitude along the axial direction bounded by the white dotted region for one pulse illumination (dotted line), U-Net (blue), and DenP2P (green), respectively. The corresponding inset is the enlarged version of the peak.

[0025] FIG.10 shows ultrasound images (column 1), AR-PAM image with one pulse illumination (column 2), post-processed PA images with a non-learning BM3D algorithm (column 3), a U-Net model (column 4), and cGAN (DenP2P) model (column 5) of a mouse kidney. Each row specifies a different spatial location of the kidney. Blue box: ROI; Yellow box: background.

[0026] FIG.11 shows ultrasound images of tumor (column 1), AR-PAM images with one pulse illumination (column 2), Post-processed PA images with a non-learning BM3D algorithm(column 3), a U-Net model (column 4), and the cGAN (DenP2P) model (column 5). Each row specifies a different spatial location of the tumor.

[0027] FIG.12A is a plot of a percentage of photobleaching at different scanning positions for one pulse, 30 pulse, and cGAN denoising. Location 0.1 mm shows no photobleaching as it was the reference starting position.

[0028] FIG.12B is a volumetric comparison between one pulse and 30 pulse illuminated PA images.

[0029] FIG.12C is a comparison of PA amplitude at three scan positions for one pulse, 30 pulse, and cGAN (DenP2P) denoising. The left most column depicts the ultrasound images of the corresponding scanning positions of the tube.

[0030] FIG.13 is a schematic of StO2 imaging with LA, HA, and DL methods.

[0031] FIG.14 shows plot of training and validation loss curves for the CONFIG-UNet model trained on dual-wavelength (750 nm and 850 nm) PA data. The network was trained over 40 epochs with a 10% validation split. The plot demonstrates consistent convergence of the model, with validation loss closely tracking the training loss, indicating effective generalization and minimal overfitting during training.

[0032] FIG.15 shows plots of training loss curves for all the deep learning models trained on dual-wavelength (750 nm and 850 nm) PA data. The network was trained over 40 epochs. The plot demonstrates consistent convergence of the model.

[0033] FIG.16 shows plots evaluation the reconstruction quality using both conventional and advanced perceptual metrics for CONFIG-UNet trained on dual-wavelength (750 nm and 850 nm) photoacoustic data. Alongside standard Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM), the Chaos-Aware Structural Similarity Metric (CASSIM) and Turbulence-Weighted Image Quality (TWIQ) were employed to assess the perceptual fidelity and physiological relevance of the reconstructed outputs. All four metrics were integrated into the model’s compile function under the metrics parameter, enabling real-time monitoring of quality progression during training for 40 epochs.

[0034] FIG.17A shows denoised outputs from various DL models using 750^nm and 850^nm LA captured inputs show clear differences in low-intensity signal recovery (white arrows).

[0035] FIG.17B shows plots of vertical line profiles (blue and white lines in FIG.17A) across ROIs highlight CONFIG-UNet’s superior fidelity in preserving true signal structures, especially under heavy noise at 850^nm.

[0036] FIG.17C shows plots of zoomed-in insets further comparing outputs from top- performing models against LA inputs and HA ground truths.

[0037] FIG.17D shows StO2maps reconstructed by each model using known HbO2and Hb tube phantoms; white circles / arrows mark low-fluence regions. Our model reliably recovers accurate StO2 values even in poorly illuminated zones.

[0038] FIG.17E shows confusing matrices of each model, illustrating CONFIG-UNet’s unique ability to distinguish biologically plausible saturation levels in low-signal regions, underscoring its denoising and generalization.

[0039] FIG.18A shows US co-registered StO2 maps of mouse thigh muscle under alternating 100% and 21% oxygen breathing, captured using LA, HA, and the DL model (CONFIG-UNet), demonstrating the ability of the DL approach to recover physiologically plausible oxygenation levels in real time, closely matching HA GT outputs.

[0040] FIG.18B is a plot of a quantitative comparison of mean StO2across mice and breathing cycles confirming that DL matches HA ground truth with low variability, while LA overestimates oxygenation.

[0041] FIG.18C shows StO2maps of subcutaneous tumors from two mice showing that the DL method (CONFIG-UNet) accurately captures spatially heterogeneous oxygenation patterns— including hypoxic and well-perfused regions (white and green arrows, respectively)—in agreement with HA references and outperforming noisy LA reconstructions.

[0042] FIG.19A shows US co-registered StO2maps of the human palm (I–III) and wrist (IV– VI), illustrating that the DL method (CONFIG-UNet) closely matches the HA ground truth in localizing and quantifying oxygenation in superficial vessels, including common palmar digital arteries and distinguishable arteries and veins (white arrows). LA inputs remain noisy and overestimate StO2 across regions.

[0043] FIG.19B shows StO2 imaging of the dorsal foot (I–III: 2D, IV–VI: 3D) demonstrates that DL and HA clearly delineate arcuate artery branches (white arrow), in contrast to noisy and overestimated LA outputs. These results confirm the DL method’s reliability in capturing physiologically meaningful oxygenation patterns in complex human vascular anatomy.

[0044] FIG.20A shows plots of frequency flow analysis of the CONFIG-UNet denoising network.

[0045] FIG.20B shows plots of Eigenvalue spectrum and retained energy of the CONFIG-UNet denoising network.

[0046] FIG.20C shows plots of the training dynamics and representation compression of the CONFIG-UNet denoising network.

[0047] FIG.20D shows plots of Wasserstein distance and sparsity trends of the CONFIG-UNet denoising network.

[0048] FIG.20E shows plots of wavelet-domain feature decomposition of the CONFIG-UNet denoising network.

[0049] FIG.20F shows saliency-based signal localization.

[0050] FIG.20G shows plots of loss landscape evolution of the CONFIG-UNet denoising network.

[0051] FIG.20H shows plots of lie algebra analysis for network stability of the CONFIG-UNet denoising network.

[0052] FIG.20I shows plots of UMAP trajectory across depth with neighborhood size as 5 of the CONFIG-UNet denoising network.

[0053] FIG.20J shows plots of epoch-wise feature convergence of the CONFIG-UNet denoising network. Detailed Description

[0054] Real-time diagnosis plays a crucial role in clinical settings as critical decisions often rely on the dynamic functionalities of organs or tissues. Therefore, fast imaging is essential to minimize corruption caused by motion artifacts resulting from natural breathing or heartbeats. Furthermore, fast imaging can reduce photobleaching when using contrast agents due to decreased exposure. These challenges can be mitigated by accelerating the imaging process without compromising the quality of signal amplitude and structural reconstruction. The systems and methods described herein enable denoising of low SNR images to achieve similar outcomes as high SNR images, increase imaging speeds and subsequent diagnosis by a user, and can also reduce the impact of photobleaching or photodamage caused by exogenous contrast agents due to a smaller number of photon interactions.

[0055] Referring to FIG.1, a system 100 for improving SNR in real-time in PAI is shown. The system 100 includes a PAI imaging system for interrogation of a sample 104. In a non-limiting example, the sample 104 is a biological tissue that may be in vivo or ex vivo. The PAI system 102 includes an excitation probe 106 such as a laser source for emitting non-ionizing light 108 toward the sample 104. In a non-limiting example the wavelength of the light may range from 400 – 2000 nm. In a non-limiting example, the wavelength may be chosen based on a target structure 110 and its depth 111 within the sample 104 relative to the surface of the sample 104.

[0056] The sample 104 absorbs the energy from the light 108 and is converted to heat. The sample 104 and any target objects or regions of interest (ROI) 110 therein undergo transient thermoelastic expansion resulting in photoacoustic signals 112. In a non-limiting example, the target object 110 includes, but is not limited to, vasculature, tumors, muscle, bone, lymphatic vessels, implants, or other natural or foreign objects. The PAI system 102 further includes a detection probe 114 for detecting the photoacoustic signals 112 at a first frame acquisition rate. In a non-limiting example, the detection probe 114 may be an ultrasound probe. The PAI system 102 may further include a signal processing unit 116 to convert the photoacoustic signals 112 into one or more image frames 118. Furthermore, the detection probe 114 is adjustable for frame acquisition rate. In a non-limiting example, the frame acquisition rate of the detection probe 114 may be adjusted via the signal processing unit 116. As used herein, “frame acquisition rate” refers to rate at which consecutive frames of photoacoustic signals 112 are captured by the detection probe 114. In a non-limiting example, the frame acquisition rate may be between 0.15 and 30 Hz. The signal processing unit 116 may also control the excitation probe 106 to emit light at a desired wavelength.

[0057] The system 100 further includes a processor 120. In a non-limiting example, the processor 120 includes a frame averaging unit 122 which receives the photoacoustic image frames 118. In a non-limiting example, the signal processing unit 116 may be integrated into the processor 120, or its functions carried out by a photoacoustic signal processing unit within processor 120.

[0058] The frame averaging unit 122 reconstructs the photoacoustic image frames by performing frame averaging. As used herein, “frame averaging” refers to the number of frames that are averaged into a composite image. In a non-limiting example, the averaging unit 122 performs a low number frame averaging (LA) to produce LA image data 124. For example, the averagingunit 122 may average 128 frames. Alternatively, the averaging unit 122 may perform a high number frame averaging (HA) to produce HA image data. For example, the averaging unit 122 may average 256,000 frames. In a non-limiting example, the frame averaging may be performed on less than 128 frames, between 128 frames and 256,000 frames, or more than 256,000 frames.

[0059] The processor 120 may further include a DL model unit 126. In a non-limiting example, the DL model unit 126 receives the LA image data 124 as input and processes the LA image data 124 in real-time using a DL model. The DL model may process the LA image data 124 such that after processing, the SNR of the LA image data 124 is statistically identical to the SNR of a high number frame averaging (HA) of the photoacoustic signals 112 acquired at a second frame acquisition rate. In a non-limiting example, the first frame acquisition rate is larger than the second acquisition frame rate. In another non-limiting example, the first frame acquisition rate may be 200 times greater than the second acquisition frame rate or the first frame acquisition rate may be anywhere in a range of from 100 to 300 times greater than the second acquisition frame rate.

[0060] In a non-limiting example, the SNR of the processed LA image data is statistically identical to the SNR of HA of the photoacoustic signal at a probability value (p-value) less than 0.05. In a non-limiting example, the p-value may be less than 0.01. See Example 1 below for further details. In a non-limiting example, the low number frame averaging is smaller than the high number frame averaging. In another non-limiting example, the LA is 200 times smaller than the HA or the LA may be anywhere in a range of from 100 to 300 times smaller than the LA.

[0061] The DL model unit 126 may further generate one or more output images 128 from the processed LA image data. Furthermore, the SNR of the output image may be at least three times greater than the SNR of the LA image data or the SNR of the output image may be anywhere in a range of from three to five times greater than the SNR of the LA image data.

[0062] In a non-limiting example, the system 100 may further include a display 130 configured to display the one or more output images 128.

[0063] Referring now to FIG.2, a system 200 for monitoring oxygen saturation (StO2) in real- time using PAI is shown. Descriptions may BE omitted for structures with the same reference numerals of FIG.1 as they represent the same structure already described. In system 200, the excitation probe 206 may emit light 208 at different wavelengths. In a non-limiting example, the excitation probe 206 may be controlled to emit two or more wavelengths sequentially orsimultaneously. In one example, the excitation probe 206 may emit a first wavelength and a second wavelength. The first wavelength and the second wavelength may be different wavelengths. In a non-limiting example, the first wavelength may be 750 nm and the second wavelength may be 850 nm, or vice versa. In another non-limiting example, the two or more wavelengths may be anywhere in a range of from 400 – 2000 nm. In a non-limiting example, the wavelengths may be chosen based on the target structure 110 and its depth within in the sample 104 relative to the surface of the sample 104. As described previously in FIG.1, photoacoustic signal processor 116 may also control the excitation probe 206 to emit light at a desired two or more wavelengths.

[0064] The PAI system 202 converts the photoacoustic signals 112 of each of the two or more wavelengths in one or more image frames 118, 218. As shown in FIG.2, image frames 118 correspond to the first wavelength and image frames 218 correspond to the second wavelength. However, this example is not intended to limiting and the system 200 can process greater than two sets of image frames.

[0065] The system 200 further includes a processor 220. In a non-limiting example, the processor 220 includes a frame averaging unit 122 which receives the photoacoustic image frames 118 and 218. In a non-limiting example, the signal processing unit 116 may be integrated into the processor 220, or its functions carried out by a photoacoustic signal processing unit within processor 220.

[0066] The frame averaging unit 122 reconstructs the photoacoustic image frames 118 and 218 by performing frame averaging. In a non-limiting example, the averaging unit 122 performs a low number frame averaging (LA) to produce LA image data 124 and 224 for each of tge two or more wavelengths. For example, the averaging unit 122 may average 128 frames. Alternatively, the averaging unit 122 may perform a high number frame averaging (HA) to produce HA image data. For example, the averaging unit 122 may average 256,000 frames.

[0067] The processor 220 may further include a DL model unit 126. In a non-limiting example, the DL model unit 126 receives the LA image data 124 and 224 as inputs and processes the LA image data 124 and 224 in real-time using a DL model. The DL model may process the LA image data 124 and 224 such that after processing, the SNR of the LA image data 124 and 224 is statistically identical to the SNR of a high number frame averaging (HA) of the photoacoustic signals 112 for each of the two or more wavelengths acquired at a second frame acquisition rate.In a non-limiting example, the first frame acquisition rate may be 200 times greater than the second acquisition frame rate.

[0068] In a non-limiting example, the SNR of the processed LA image data is statistically identical to the SNR of HA of the photoacoustic signal at a probability value (p-value) less than 0.05. In a non-limiting example, the p-value may be less than 0.01. In another non-limiting example, the LA is 200 times smaller than the HA.

[0069] The DL model unit 126 may further generate one or more output images 128 and 228 for each of the two or more wavelengths from the processed LA image data.

[0070] In a non-limiting example, the processor 220 further includes a linear unmixing unit 229. Output images 128 and 228 are input into the linear unmixing unit, which performs linear unmixing on the output images of each of the two or more wavelengths 128, 228.

[0071] The linear unmixing, as will be described in further detail below, is defined by^^ఒభ ∗ ^^ఒమ − ^^ఒమ ∗ ^ ఒభ^^^^^^ ு^ ^ ு^ ^^ଶ =^^where ^^ఒௗ^^^ = ^^ఒு^ − ^^ఒு^ைమ ,coefficients for deoxygenated(Hb) and oxygenated hemoglobin (HbO2), respectively, at the two or morewavelengths ^^^, and ^^^ఒdenotes the absorption coefficient at the two or more wavelengths ^^^, and ^^ is the numbera wavelength of the two or more wavelengths In a non-limiting example,^^^ is 750 nm and ^^ଶ is 850 nm.

[0072] Furthermore, in a non-limiting example, the linear unmixing unit 229 may generate an StO2 map from the linear unmixing of the output images.

[0073] In a non-limiting example, the system 200 may further include a display 130 configured to display the StO2map. The display 130 may additionally or alternatively display the one or more output images 128, 228.

[0074] The DL model unit 126 in FIGS.1 and 2 may be a cGAN or U-Net based convolutional neural network (CNN). An example workflow 300 utilizing the DenP2P model as described herein is shown in FIG.3A. DenP2P is a cGAN and includes a generator 302 and a discriminator 308. In a non-limiting example, the generator 302 is a U-Net based model. In another non- limiting example, the discriminator 308 includes a PatchGAN classifier. For example, the PatchGAN classifier may focus on classifying a 70x70 portion of the input image based on a 30x30 output image patch. As shown in FIG.3B, the generator comprises a plurality of blocks ineach of an encoding pathway 304 and a decoding pathway 306. Furthermore, the encoder (down- sampling) 304 and decoder (up-sampling) 306 operations include corresponding residual connections or skip connections which help to recover local information lost during the down- sampling process.

[0075] In a non-limiting example, each block in the down-sampling pathway consists of a series of convolution layer, batch normalizations, and rectified linear unit (^^^^^^^^) activations. For example, the first block includes a convolution layer and a ReLU activation. A second block and third block include a convolution layer, a batch normalization, and a ReLU activation. A remainder of blocks in the encoding pathway may each include a convolution layer, a batch normalization, a ReLU activation, and a dropout layer. In a non-limiting example, all dropout layers have a probability of 0.5 for discarding random nodes.

[0076] The decoding pathway 304 may also include a plurality of blocks, whereby each block includes a transposed convolution layer, a batch normalization, and a ReLU activation. Additionally, each block except for the final three blocks include a dropout layer.

[0077] In a non-limiting example, the weight initialization for all layers may be additionally performed using ‘he_normal’ which draws weights from a truncated normal distribution. This initialization strategy works well with ^^^^^^^^ activation due to its controlled weight initialization, facilitating the rapid convergence to a global minimum of the loss function.

[0078] Referring now to FIG.3C, the discriminator 308 is presented. In a non-limiting example, the discriminator 308 includes a plurality of blocks. A first block includes a convolution layer and a ReLU activation. Subsequent layers between the first block and a final block include a convolution layer, a batch normalization, and a ReLU activation. The final block may include a convolution layer, a batch normalization, and a leakyReLU activation.

[0079] The DenP2P DL model will be described in further detail in Example 2 below

[0080] In an alternative example, the DL model may be a U-Net generator including attention- dense residual blocks in the encoding pathway and decoding pathway, referred to hereinafter as CONFIG-UNET. As will be described in further detail in Example 3 below, the CONFIG-UNet DL model is built on the U-Net framework, and further includes residual connections, dense blocks, and attention mechanisms.

[0081] In a non-limiting example, the CONFIG-UNet DL model a plurality of layers as shown in FIG.4. A first encoder layer may include a residual convolution block, a batch normalization, aReLU activation, a dense block a skip connection, and downward max pooling. A second encoder layer may include a residual convolution block and a dense block, a skip connection, and downward max pooling. A third and fourth encoder layer may each include a dense block, a skip connection, and downward max pooling. The CONFIG-UNet DL model may further comprise a bottleneck layer including a residual convolution block, a skip connection, and downward max pooling. The CONFIG-UNet DL model may further comprise a first, second, and third decoder layer, each including an attention gate and upward deconvolution, and a fourth decoder layer including an attention gate and a convolution layer. In a non-limiting example, each decoder layer includes a convolution layer, a dense block, batch normalization, and ReLU activation. Furthermore, each residual convolution block and each dense block of each decoder layer may include a plurality of convolution sublayers, batch normalizations, and ReLU activations.

[0082] In a non-limiting example, the CONFIG-UNet model is trained using a loss function including an intensity aware loss (IAL) component, a frequency loss (FL) component, and amean squared error (MSE) component, wherein the los function is defined asℒ௧^௧^^ = ^^^^௧^^^^௧௬ ∙ ℒ^^௧^^^^௧௬ + ^^^^^^ ∙ ℒ^^^^,where ℒ is the lossthe IAL, and ^^^^^^[0,1] is a weighting factor for the FL.

[0083] The IAL is defined as1^ ^ ௧^௨^ ^^^ௗ ଶℒ^^௧^^^^௧௬ =^^^ ^ 1 − ^^ ∙ ^^^^ + ^^൧ ∙ ൫^^^^ − ^^^^ ൯ ,where ^^ ∈ [0,1]and aimage intensity at pixel location (^^, ^^), respectively, and ^^ ^^^ = ^ା^௬^^ೠ^ .^ೕ ^

[0084] The FL is defined as1ℒ = ∙ ^^ ^௧^௨^ ^^^^ௗ^^^^ ห^^ ห − ห^^ ห^ ,where|∙|denotes a magnitudeTransform (FFT) at spatialfrequency (^^, ^^), ^^ is a number of frequency bins, ห^^^௧^௨^ ௧^௨^ ^^^ௗ௨௩ ห is ^^^^^^2(^^ ), and ห^^^௨௩ ห is^^^^^^2(^^^^^ௗ).

[0085] In a non-limiting example, the CONFIG-UNet DL model is evaluated using evaluation metrics including a Chaos-Aware Structural Similarity (CASSIM) metric, a Turbulence- Weighted Image Quality (TWIQ) metric, a Peak Signal to Noise Ratio (PSNR), and a Structural Similarity Index (SSIM).

[0086] The CASSIM metric may be defined as^^^^^^^^^^^^(^^௧^௨^ , ^^^^^ௗ) = ^^ିఒ ∙ ^^^^^^^^(^^௧^௨^ , ^^^^^ௗ),where ^^ is a2൫^^^^௫ + ^^^^௬൯,|]|],^^^^^^^^(^^௧^௨^ , ^^^^^ௗ) = ห^^^^^ೠ^ − ^^^^^^^ห,where average spectral ^^ (^^ ) | |ଶ^ , ^^ = ℱ{^^}(^^, ^^) ,and ℱ{^^} is a Fourier transform

[0088] The PSNR metric may be defined as^^^^^^^^ = ୫ୟ^(^ୟ୰^^^)where max(Target) is a maximumof interest (ROI) in the output image and ^^^^^^^^^௨^ௗis a standard deviation of a selected background region of the output image.

[0089] The SSIM metric may be defined as (2^^ ^^ + ^^ )(2^^ + ^^ )SSIM = ^ ^ ^ ^^ ଶ ,where ^^^is a sampleimage, A, ^^^ଶis a of A, ^^^^ is a covariance of A and a second image array B of the ROI, ^^^ is defined as[^^^^^]ଶ and ^^ଶ is defined as [^^ଶ^^]ଶ, where ^^^ and ^^ଶ are 0.01 and 0.03, respectively, and L is adynamic range of the ROI.

[0090] Referring now to FIG.5, an example of a system 500 for improving SNR in real-time PAI in accordance with some embodiments of the systems and methods described in the present disclosure is shown. As shown in FIG.5, a computing device 550 can receive one or more types of data (e.g., PAI frames) from data source 502. In some embodiments, computing device 550 can execute at least a portion of a PAI processing system 504 to process data received from the data source 502.

[0091] Additionally or alternatively, in some embodiments, the computing device 550 can communicate information about data received from the data source 502 to a server 552 over a communication network 554, which can execute at least a portion of the PAI processing system 504. In such embodiments, the server 552 can return information to the computing device 550 (and / or any other suitable computing device) indicative of an output of the PAI processing system 504.

[0092] In some embodiments, computing device 550 and / or server 552 can be any suitable computing device or combination of devices, such as a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable computer, a server computer, a virtual machine being executed by a physical computing device, and so on. The computing device 550 and / or server 552 can also reconstruct images from the data.

[0093] In some embodiments, data source 502 can be any suitable source of data (e.g., measurement data, images reconstructed from measurement data, processed image data), another computing device (e.g., a server storing measurement data, images reconstructed from measurement data, processed image data), and so on. In some embodiments, data source 502 can be local to computing device 550. For example, data source 502 can be incorporated with computing device 550 (e.g., computing device 550 can be configured as part of a device for measuring, recording, estimating, acquiring, or otherwise collecting or storing data). As another example, data source 502 can be connected to computing device 550 by a cable, a direct wireless link, and so on. Additionally or alternatively, in some embodiments, data source 502 can be located locally and / or remotely from computing device 550, and can communicate data to computing device 550 (and / or server 552) via a communication network (e.g., communication network 554).

[0094] In some embodiments, communication network 554 can be any suitable communication network or combination of communication networks. For example, communication network 554 can include a Wi-Fi network (which can include one or more wireless routers, one or more switches, etc.), a peer-to-peer network (e.g., a Bluetooth network), a cellular network (e.g., a 3G network, a 4G network, etc., complying with any suitable standard, such as CDMA, GSM, LTE, LTE Advanced, WiMAX, etc.), other types of wireless network, a wired network, and so on. In some embodiments, communication network 554 can be a local area network, a wide area network, a public network (e.g., the Internet), a private or semi-private network (e.g., a corporate or university intranet), any other suitable type of network, or any suitable combination of networks. Communications links shown in FIG.5 can each be any suitable communications link or combination of communications links, such as wired links, fiber optic links, Wi-Fi links, Bluetooth links, cellular links, and so on.

[0095] Referring now to FIG.6, an example of hardware 600 that can be used to implement data source 502, computing device 550, and server 552 in accordance with some embodiments of the systems and methods described in the present disclosure is shown.

[0096] As shown in FIG.6, in some embodiments, computing device 550 can include a processor 602, a display 604, one or more inputs 606, one or more communication systems 608, and / or memory 610. In some embodiments, processor 602 can be any suitable hardware processor or combination of processors, such as a central processing unit (“CPU”), a graphics processing unit (“GPU”), and so on. In some embodiments, display 604 can include any suitable display devices, such as a liquid crystal display (“LCD”) screen, a light-emitting diode (“LED”) display, an organic LED (“OLED”) display, an electrophoretic display (e.g., an “e-ink” display), a computer monitor, a touchscreen, a television, and so on. In some embodiments, inputs 606 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.

[0097] In some embodiments, communications systems 608 can include any suitable hardware, firmware, and / or software for communicating information over communication network 554 and / or any other suitable communication networks. For example, communications systems 608 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 608 can include hardware, firmware,and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.

[0098] In some embodiments, memory 610 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 602 to present content using display 604, to communicate with server 552 via communications system(s) 608, and so on. Memory 610 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 610 can include random-access memory (“RAM”), read-only memory (“ROM”), electrically programmable ROM (“EPROM”), electrically erasable ROM (“EEPROM”), other forms of volatile memory, other forms of non-volatile memory, one or more forms of semi- volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 610 can have encoded thereon, or otherwise stored therein, a computer program for controlling operation of computing device 550. In such embodiments, processor 602 can execute at least a portion of the computer program to present content (e.g., images, user interfaces, graphics, tables), receive content from server 552, transmit information to server 552, and so on. For example, the processor 602 and the memory 610 can be configured to perform the methods described herein.

[0099] In some embodiments, server 552 can include a processor 612, a display 614, one or more inputs 616, one or more communications systems 618, and / or memory 620. In some embodiments, processor 612 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some embodiments, display 614 can include any suitable display devices, such as an LCD screen, LED display, OLED display, electrophoretic display, a computer monitor, a touchscreen, a television, and so on. In some embodiments, inputs 616 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.

[0100] In some embodiments, communications systems 618 can include any suitable hardware, firmware, and / or software for communicating information over communication network 554 and / or any other suitable communication networks. For example, communications systems 618 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 618 can include hardware, firmware,and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.

[0101] In some embodiments, memory 620 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 612 to present content using display 614, to communicate with one or more computing devices 550, and so on. Memory 620 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 620 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other types of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 620 can have encoded thereon a server program for controlling operation of server 552. In such embodiments, processor 612 can execute at least a portion of the server program to transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 550, receive information and / or content from one or more computing devices 550, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone), and so on.

[0102] In some embodiments, the server 552 is configured to perform the methods described in the present disclosure. For example, the processor 612 and memory 620 can be configured to perform the methods described herein.

[0103] In some embodiments, data source 502 can include a processor 622, one or more data acquisition systems 624, one or more communications systems 626, and / or memory 628. In some embodiments, processor 622 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some embodiments, the one or more data acquisition systems 624 are generally configured to acquire data, images, or both, and can include photoacoustic signals. Additionally or alternatively, in some embodiments, the one or more data acquisition systems 624 can include any suitable hardware, firmware, and / or software for coupling to and / or controlling operations of PAI system, such as PAI systems 102, 202. In some embodiments, one or more portions of the data acquisition system(s) 624 can be removable and / or replaceable.

[0104] Note that, although not shown, data source 502 can include any suitable inputs and / or outputs. For example, data source 502 can include input devices and / or sensors that can be usedto receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, a trackpad, a trackball, and so on. As another example, data source 502 can include any suitable display devices, such as an LCD screen, an LED display, an OLED display, an electrophoretic display, a computer monitor, a touchscreen, a television, etc., one or more speakers, and so on.

[0105] In some embodiments, communications systems 626 can include any suitable hardware, firmware, and / or software for communicating information to computing device 550 (and, in some embodiments, over communication network 554 and / or any other suitable communication networks). For example, communications systems 626 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 626 can include hardware, firmware, and / or software that can be used to establish a wired connection using any suitable port and / or communication standard (e.g., VGA, DVI video, USB, RS-232, etc.), Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.

[0106] In some embodiments, memory 628 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 622 to control the one or more data acquisition systems 624, and / or receive data from the one or more data acquisition systems 624; to generate images from data; present content (e.g., data, images, a user interface) using a display; communicate with one or more computing devices 550; and so on. Memory 628 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 628 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other types of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some embodiments, memory 628 can have encoded thereon, or otherwise stored therein, a program for controlling operation of data source 502. In such embodiments, processor 622 can execute at least a portion of the program to generate images, transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 550, receive information and / or content from one or more computing devices 550, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone, etc.), and so on.

[0107] In some embodiments, any suitable computer-readable media can be used for storing instructions for performing the functions and / or processes described herein. For example, insome embodiments, computer-readable media can be transitory or non-transitory. For example, non-transitory computer-readable media can include media such as magnetic media (e.g., hard disks, floppy disks), optical media (e.g., compact discs, digital video discs, Blu-ray discs), semiconductor media (e.g., RAM, flash memory, EPROM, EEPROM), any suitable media that is not fleeting or devoid of any semblance of permanence during transmission, and / or any suitable tangible media. As another example, transitory computer-readable media can include signals on networks, in wires, conductors, optical fibers, circuits, or any suitable media that is fleeting and devoid of any semblance of permanence during transmission, and / or any suitable intangible media.

[0108] As used herein in the context of computer implementation, unless otherwise specified or limited, the terms “component,” “system,” “module,” “framework,” and the like are intended to encompass part or all of computer-related systems that include hardware, software, a combination of hardware and software, or software in execution. For example, a component may be, but is not limited to being, a processor device, a process being executed (or executable) by a processor device, an object, an executable, a thread of execution, a computer program, or a computer. By way of illustration, both an application running on a computer and the computer can be a component. One or more components (or system, module, and so on) may reside within a process or thread of execution, may be localized on one computer, may be distributed between two or more computers or other processor devices, or may be included within another component (or system, module, and so on).

[0109] In some implementations, devices or systems disclosed herein can be utilized or installed using methods embodying aspects of the disclosure. Correspondingly, description herein of particular features, capabilities, or intended purposes of a device or system is generally intended to inherently include disclosure of a method of using such features for the intended purposes, a method of implementing such capabilities, and a method of installing disclosed (or otherwise known) components to support these purposes or capabilities. Similarly, unless otherwise indicated or limited, discussion herein of any method of manufacturing or using a particular device or system, including installing the device or system, is intended to inherently include disclosure, as embodiments of the disclosure, of the utilized features and implemented capabilities of such device or system.

[0110] Referring now to FIG.7, a method 700 for improving SNR in real-time in PAI is shown. In a non-limiting example, the method 700 may be performed using any of the systems described previously. At step 702, light is delivered to a tissue and detecting PA signals at a first frame acquisition rate from the tissue in response to excitation of the tissue by the light. In a non- limiting example, the light may be delivered using an excitation probe and PA signals may be detected using a detection probe, such as excitation probes 106, 206 and detection probe 114 of FIGS.1-2. Next at step 704, a processor may receive the PA signals, such as processor 116 in FIGS.1-2. At step 706, the processor may reconstruct the PA signals by performing low number frame averaging (LA) on the PA signals to produce LA image data. Next, at step 708, the processor may further process the LA image data in real-time using a DL model. The LA image data is processed such that the SNR of the processed LA image data is statistically identical to the SNR of a high number frame averaging (HA) of the PA signals acquired at a second frame acquisition rate. At step 710, the processor may further generate one or more output images from the processed LA image data. Next at step 712, a display, such as display 130 in FIGS.1-2, may output the one or more output images.

[0111] Additionally or alternatively, where the light delivered to the tissue in step 702 includes two or more wavelengths, the method 700 may also be used to measure StO2. After step 710, the processor may optionally perform linear unmixing of the one or more output images from each of the two (or more) wavelengths at step 714. Further, the processor may further generate an oxygen saturation (StO2) map from the linear unmixing of the one or more output images from each of the two (or more) wavelengths at step 716. The StO2 map also be display on the display at step 712.

[0112] The following examples provide further details and applications of the systems and DL models described above. Example 1

[0113] Using a four-layer U-Net architecture, noise can be removed from LA image data to be quantitatively similar to HA image data of the same PA signals. FIG.8 various quantitative metrics were used to measure the enhancement of image quality in the U-Net images over the low number of frame averaged image inputs. Specifically, SNR, PSNR, and CNR were used as image quality metrics. FIG.8 illustrates data calculated from 24 mouse tumor images (8 mice × 3 cross-sections / mouse). All the metricswere calculated considering various tumor ROIs andbackground zones. There is approximately a five-fold increase (4.27 ± 0.87) in SNR between LA and the U-Net outcomes. No significant difference in the image quality metrics were observed between the U-Net and the HA images. Example 2

[0114] 2.1 Methodology

[0115] 2.1.1 Microscopic imaging system

[0116] To perform PA imaging, the AR-PAM system employed an OPO (Phocus HE Benchtop, OPOTEK) to generate pulsed laser energy. The subsequent acoustic signal generated was collected using a 1-inch focused transducer (V324-SU, Olympus) with a center frequency of 25 MHz. This signal was amplified by a pulser / receiver (DPR500, JSR Ultrasonics) which additionally facilitated signal digitization by a data acquisition card (CSE161G2, GaGe).

[0117] To transmit the laser energy to the sample, the system utilized an achromatic doublet (AC254-030-B-ML, Thorlabs) and an aspheric condenser (ACL2520U-B, Thorlabs) to couple the laser beam from the OPO fiber into a secondary fiber bundle. This secondary fiber bundle consisted of seven multimode fibers (FT1000EMT, Thorlabs) epoxied together towards the proximal end for coupling. The distal end of the fiber bundle was designed in a fanned configuration and connected to a 3D-printed transducer-and-fiber mount which oriented the fibers to mimic an optical condenser.

[0118] 2.1.2 Non-learning traditional models

[0119] Several non-learning noise removal algorithms were implemented, such as Savitzky- Golay (SG), Wiener, and BM3D for comparison to the DenP2P model. In the SG algorithm, a 3rd order SG filter was employed (^^^^^^^^^^^^^^^^^^^^ in MATLAB®) with a frame length of 9 for convolution-based smoothing. The Wiener algorithm was implemented using a pixel-wise adaptive low-pass filter (^^^^^^^^^^^^2 function in MATLAB®) with a neighborhood size of 5 ^^ 5 to estimate the local image mean and standard deviation. In BM3D, the standard deviation (^^) was initially estimated using a multi-channel approach and subsequently applied Python's ^^^^3^^ function with the estimated ^^ and ^^^^3^^^^^^^^^^^^^^.^^^^^^ ^^^^^^^^^^^^ configuration. This configuration, which includes Wiener filtering in addition to hard thresholding, provides slightly improved performance for the system-specific noise compared to other ^^^^3^^^^^^^^^^^^^^.

[0120] 2.1.3 Deep learning architecture

[0121] 2.1.3.1 U-Net

[0122] Within the implementation, two DL networks were employed, one of which was the original U-Net architecture. the ‘mean absolute error’ loss function and the ‘Adam’ optimizer with an initial learning rate of 0.0001 were utilized. The U-Net is distinguished by its inclusion of residual skip connections which contribute to the network's notable effectiveness.

[0123] 2.1.3.2 Conditional GAN (DenP2P)

[0124] A Pix2Pix network-based cGAN denoising model was devised, referred to herein as DenP2P, as previously depicted in FIGS.3A-3C. The general workflow of DenP2P is presented in FIG.3A, while FIGS.3B-C showcase the architecture of the generator and the PatchGAN discriminator, respectively.

[0125] Generator: Instead of using an autoencoder, a U-Net-based model was employed as the generator. The down-sampling and up-sampling operations are accompanied by corresponding residual connections which help recover local information lost during the down-sampling process.

[0126] Each block in the down-sampling pathway consists of a series of convolution layers, batch normalizations, and rectified linear unit (^^e^^^^) activations. The first block does not include a batch normalization layer and the first three blocks do not have a dropout layer. Similarly, each block in the decoder section is comprised of a sequence of transposed convolution layers, batch normalizations, and ^^e^^^^ activations, with the last three blocks excluding a dropout layer.

[0127] All dropout layers have a probability of 0.5 for discarding the random nodes. Additionally, the weight initialization for all layers was performed using ‘he_normal’ whichdraws weights from a truncated normal distribution. This initialization strategy works well with^^^^^^^^ activation due to its controlled weight initialization, facilitating the rapid convergence to aglobal minimum of the loss function.

[0128] Discriminator: Adhering to the Pix2Pix discriminator architecture, a convolutional PatchGAN classifier was employed. This classifier focuses on classifying a 70^^70 portion of the input image based on a 30^^30 output image patch. Each block of the discriminator consists of a convolution layer, batch normalizations, and ^^e^^^^ activations. However, the first block does notinclude a normalization layer and the last layer has a ^^^^^^^^^^^^^^^^^^ activation instead of a regular^^^^^^^^ activation.

[0129] The cGAN computational architecture attempts to find a mapping from the observed image (noisy input) x and random noise vector z to label (ground truth) y, ^^: {^^, ^^} → ^^. The generator of the cGAN learns a structured loss by minimizing the difference between the network outputs and the corresponding targeted images. The cost function of a cGAN isnormally expressed asℒ^ீ^ே(^^, ^^) = ^^௫,௬[^^^^^^^^(^^, ^^)] + ^^௫,௭[log(1 − ^^(^^,^^(^^,^^)]which ^^ and discriminator ^^ where ^^ tries tomaximize the objective function is expressedas ^^∗ = arg^^^^^^ீ^^^^^^^ ℒ^ீ^ே(^^, ^^). Keeping the loss function of the discriminator unchangedfrom the DenP2P model, a Huber loss an additional structural similarity (^^^^^^^^) loss may be used for the generator. As a result, ^^ attempts to fool ^^ as it tries to produce images that are not distant from the real data following the Huber loss which considers both ^^1and ^^2, norm based on a ^^ value and are structurally similar to the ground truth based on the ^^^^^^^^ loss. The Huber loss is defined as 1(^^ − ^^(^^, ^^))ଶ |^^ − ^^(^^, ^^)| ≤ ^^^and SSIM lossℒௌௌூெ(^^) = 1^^^1 − ^^^^^^^^൫^^,^^(^^, ^^)൯,where SSIM(A, B) is(2^^ ^^ + ^ )( )= ^ ^ ^^ 2^^^^ + ^^SSIM(^^, ^^) ଶ^^ ^ ,where ^^^is a sample(ROI) of the outputimage, A, ^^ଶ^ is a variance of A, ^^^^ is a covariance of A and a second image array B of the ROI,^^ ଶ ଶ^ is defined as [^^^^^] and ^^ଶas [^^ଶ^^] , where ^^^ and ^^ଶ are 0.01 and 0.03,respectively, and L is a dynamic range of the ROI.

[0130] These additional losses help the cGAN resist the vanishing gradient problem because the loss metrics can measure how far the generated data distribution is from real ground truth, and the generator learns from discriminator losses. The combined generator loss functions give us the final objective function of ^^ as:^^∗ = arg^^^^^^ீವ^^ುమು^^^^^^^ವ^^ುమು ℒ^ீ^ே(^^^^^^ଶ^, ^^^^^^ଶ^) + ^^^^ℒு(^^^^^^ଶ^)+whereIn a non-limiting example, ^^^^is 100.

[0131] 2.1.4 Database

[0132] Training data: In an imaging system, the number of frame averages directly affects the quality of the image. A low number of frame averages results in a noisy image with low SNR, whereas a high number of frame averages produces a high SNR image.

[0133] The LED-based Acoustic-X PA imaging system was utilized to capture the training inputs (low SNR) and their corresponding paired labels or ground truths (high SNR). For 800 paired training images, the low frame averaging scenario averaged 128 frames, while for the high frame averaging scenario averaged 25600 frames. There was no spatial misalignment between the training inputs and their corresponding labels as the transducer was not moved while performing the low and high frame averaging. Additionally, computational noise was not introduced; rather, training was performed with inherent system-generated electronic noise, such as shot noise and thermal noise. Note that the training data originated from a completely different imaging platform which typically operates with macro-level objects compared to the AR-PAM testing system which mainly deals with microscopic data.

[0134] Test data: Test data was collected using a custom-built AR-PAM system. For imaging samples, carbon fiber, mouse kidney, and mouse tumor specimens were utilized. To investigate the photobleaching phenomenon, an experiment was conducted using a polyethylene tube filled with a mixture of 0.5 μM ICG and deionized water, which was securely positioned at both ends of a box. To assess the plausibility of inducing photobleaching of ICG by laser pulses, the tube was illuminated at a single point for 15 minutes in three separate trials using both deionized water and dimethyl sulfoxide (DMSO). In all cases, we observed a consistent decrease in the PA amplitude over time, indicating the occurrence of photobleaching.

[0135] Subsequently, a 2D scanning of the tube in the ^^ − ^^ direction was performed with a step size of 100 μm. Two laser pulse illuminations were employed: single-pulse illumination (no averaging) and 30-pulse illumination.

[0136] 2.1.5 Network parameters

[0137] The Adam optimizer was utilized for both the generator and discriminator, with an initial learning rate of 0.0001. The parameters ^^+ and ^^, were set to 0.5 and 0.999, respectively, and ∈ was set to 1034. During training, a batch size of 1 was employed and the process consisted of40,000 steps. Within the combined generator objective function, the regularization parameter^^*2 was set to 100. Additionally, the ^^ parameter for the Huber loss was set to 1.

[0138] All deep learning code was implemented using TensorFlow (version 2.9.2) with the Keras backend and executed on the Google Colab Pro platform. The computations were performed on a Tesla T4 GPU (CUDA version: 11.2) with 25.45 GB of RAM. The training process took approximately 93-94 seconds per 1000 steps.

[0139] For other data preprocessing tasks, such as generating images from the A-scans and calculating image quality metrics, MATLAB® (version 9.10.0.2015706, R2021a, update 7) was utilized.

[0140] 2.1.6 Image Quality Metrics

[0141] To check the quality of the cGAN generated output images, two reference image quality metrics were used.

[0142] SNR: SNR is a no reference quality metric measured in dB scale to accommodate signalswith a wide dynamic range. It is defined as^^^^^^^^ = 20 ∙ log ோைூ^^ ^^ ,^ீwhere ^^ோைூis the mean signal amplitude of the target region of interest (ROI), and ^^^ீis the standard deviation of the background noisy region.

[0143] CNR: Contrast to noise ratio (CNR)

[0053] is a no reference image quality metric which compares the standard deviation of the target ROI with respect to the background noisy region. Itis defined as^^^^^^ = ^^ோைூ^^ ,where ^^ோைூis the standard deviation ofand ^^^ீis the standard deviation of the background noisy region.

[0144] 2.2 Results

[0145] 2.2.1 Carbon fiber experiment

[0146] A carbon fiber was imaged with two laser pulse illuminations: one pulse (input) and 20 pulse averaging (ground truth). In Fig.9A, cross-section images of the carbon fiber are shownfor one pulse ((a)) and 20 pulses ((b)). Noise removal was applied to the one pulse images using traditional non-learning algorithms (SG: (c), Wiener: (d), and BM3D: (e)), U-Net ((f)) and DenP2P ((g)). FIG.9B compares the outcomes of the U-Net-based model with DenP2P (h-j) using the mean PA amplitude along the axial direction (as shown by the white dotted vertical lines in (h)). FIG 9C.) shows the line profile of the one pulse illumination (dotted line), the U- Net generated image (blue), and DenP2P generated image (green) for the carbon fiber along the axial direction ((k)). The inset in (k) shows the enlarged display of the PA amplitude.

[0147] 2.2.2 Ex-vivo organ imaging

[0148] Ex-vivo organ imaging was performed on mouse kidney using the AR-PAM system. In FIG.10, each row represents an image of the kidney captured at different cross-sections. The second column displays the kidney image obtained with one pulse illumination. The subsequent three columns compare the denoising results of the one pulse illumination using a traditional non-learning algorithm (BM3D), the U-Net-based model, and DenP2P. The non-learning algorithm BM3D was displayed as it performed the best compared to SG and Wiener. Ultrasound images are provided alongside for morphological confirmation of the organs in the first column. Within the three rows, different spatial positions of the kidney are depicted. For calculating image quality metrics, the blue square was chosen as the ROI, while the yellow box served as the background.

[0149] 2.2.3 Tumor imaging

[0150] Ex-vivo tumor imaging was performed on subcutaneous skin carcinoma models FIG.11. Similar analysis was performed as the mouse kidney and depicted in FIG.10. Organization of FIG.11 is similar to FIG.10, but with different spatial statistics. The blue and yellow boxes in Fig.11 (b), (g), (l), and (q) represent the ROI and background, respectively.

[0151] 2.2.4 Photobleaching

[0152] For the photobleaching experiment, the tube filled with ICG in DMSO / deionized water was initially illuminated at a specific scanning location for three trials, each for 15 minutes. ICG in DMSO showed little photobleaching as expected. For all trials of ICG in water, the PA signal amplitude reduced uniformly by approximately 30%. This indicated that ICG with deionized water may undergo photobleaching, as evidenced by the consistent decrease in signal strength.

[0153] Next, a 2D raster scan of tubes filled with ICG and deionized water was conducted with a step size of 100 μm. FIG.12A illustrates the calculated percentage of photobleaching orphotodamage of ICG within the tube at four distinct crosssections spaced 0.7 mm apart. The percentage of photobleaching was determined by comparing the PA amplitude values at corresponding locations to the initial starting position. Note that the initial starting position was considered to be at 0.1 mm and therefore shows no photobleaching.

[0154] FIG.12B showcases volumetric images obtained using one pulse and 30 pulse illumination. Individual B-scans were obtained in the Y direction, orthogonal to the tube's diameter. A 2 x 3 mm2 area of the tube was scanned, consisting of 30 cross-sections spaced 100 μm apart. FIG.12C presents three of those cross-sections at different locations (start, middle, and end). Each row corresponds to a specific scanning location, and the columns display the outcomes for one pulse illumination, 30 pulse illumination, and the resulting DenP2P denoising.

[0155] 2.3 Discussions

[0156] Although AR-PAM is desirable due to its optical contrast and ultrasound morphological detailing, its speed is limited by the low pulse repetition frequency (~10 Hz) of the laser. Multiple laser pulse illuminations are required for each cross-section to average out inconsistent noise and generate a high SNR image. Conversely, rapid scanning with fewer or only one laser pulse illumination results in a low SNR image. To address this challenge, our approach involves denoising the low SNR image to achieve a similar outcome as the high SNR image. This not only enables high-quality imaging, but also reduces the impact of photobleaching or photodamage caused by exogenous contrast agents due to a smaller number of photon interactions.

[0157] FIG.9A-9C revealed that traditional non-learning denoising algorithms (SG, Wiener, and BM3D) did not effectively reduce noise. Their performance in terms of the image quality metrics SNR and CNR was unsatisfactory as presented in Table 1. Among the many extensively explored deep learning networks, U-Net has gained traction in the biomedical community for tasks such as segmentation and denoising. To address the issue of low signal strength generated by U-Net (as depicted in FIG.9B (i)), another Denp2P was developed. For both spatial positions, the maximum signal amplitude within the white dotted region was lower for U-Net compared to the original noisy image and the DenP2P-generated image. While DenP2P did not yield a significant improvement in axial resolution, the method demonstrated the ability to preserve high levels of contrast after denoising, distinguishing it from U-Net. Additionally, when comparingthe image quality metrics of traditional non-learning methods and U-Net to DenP2P, the model demonstrated significant improvement as presented in Table 1. IMAGE QUALITY ASSESSMENT METRIC FOR TUMOR Imaging SNR CNR One pulse 73.75 ± 6.32 4.12 ± 2.16 SG 77.01 ± 6.17 5.43 ± 1.47 Wiener 75.4 ± 4.74 5.62 ± 2.54 BM3D 77.93 ± 8.37 6.32 ± 2.88 U-Net 84.52 ± 11.82 6.82 ± 2.12 DenP2P 92.77 ± 10.74 9.9 ± 4.41Imaging SNR CNR One pulse 72.76 ± 9.25 6.19 ± 4.36 SG 78.09 ± 3.62 7.29 ± 3.79 Wiener 81.34 ± 6.49 6.41 ± 2.46 BM3D 78.81 ± 7.89 6.98 ± 3.44 U-Net 86.35 ± 3.97 7.59 ± 0.82 DenP2P 93.54 ± 6.07 11.82 ± 4.42 Table 1 compares image quality metrics (SNR and CNR) among different non-learning and deep learning computational methods along with the one pulse input images of tumors and kidneys at various cross-sections. Each of the corresponding table entries is depicted as ^^^^^^^^ ± ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^.

[0158] To validate the performance of our denoising deep learning model on biological models that exhibit spatial and morphological variations compared to the training images, experiments on ex-vivo mouse kidneys and subcutaneous skin carcinoma tumors were conducted. FIGS.10 and 11, and the corresponding entries in Table 1 provide evidence that DenP2P outperforms non- learning methods and the UNet-based algorithm. While non-learning methods partially reduced noise, U-Net successfully removed background noise at the expense of the original signal amplitude. In contrast, DenP2P achieved complete denoising of the background while maintaining satisfactory signal amplitude.

[0159] It is worth noting that our training data exclusively consisted of metal frame / wire phantom PA images obtained from the LED-based Acoustic-X PA imaging system. In contrast, all testing was conducted using images generated from our ARPAM imaging system. Therefore, the test data exhibited morphological dissimilarities compared to the training objects.

[0160] The speed improvement in AR-PAM is accomplished by reducing or eliminating the averaging of laser pulse-illuminated A-scans, which also leads to a supplementary benefit of reducing photobleaching. While endogenous contrast agents are typically not affected by optical bleaching, exogenous contrast agents are prone to it over time. In the tube experiment, the 15- minute period laser pulse illumination resulted in a decrease in PA amplitude signal strength of approximately 30%. The long exposure possibly caused photodamage to the ICG and deionized water which attributed to the photobleaching.

[0161] To validate this claim, a 2D raster scan of the ICG tube at 30 cross-sections was conducted spaced 100 μm apart using both single and 30 laser pulse illuminations. The 3D volumetric image generated from the single pulse illumination (FIG.12B left column) exhibited high background noise, but also high PA signal strength, indicating low photobleaching. In contrast, the volumetric image generated from the averaged 30 pulse illumination (FIG.10B right column) exhibited a gradual reduction in PA signal strength along subsequent B-scans (FIG.12B X direction) due to photobleaching.

[0162] DenP2P applied to one pulse illumination images successfully generated a high SNR image without background noise (FIG.12C 3rd column). This approach significantly reduced the scanning time by approximately 25-30 times when compared to scans with averaging. Consequently, the impact of photobleaching was mitigated and we obtained an image with satisfactory SNR and CNR. Example 3

[0163] 3.1 Introduction

[0164] Continuous monitoring of tissue oxygenation represents a critical challenge in modern clinical practices. Despite significant advancements in medical imaging and monitoring technologies, the real-time visualization of oxygen saturation (StO2) distributions across tissues and organs remains an unmet need in numerous clinical scenarios. Current standard monitoring approaches such as pulse oximetry provide only point measurements at peripheral locations, failing to capture the spatial heterogeneity of oxygen distribution that is crucial for early detection of tissue hypoxia and ischemia. The absence of real-time StO2 mapping technologies creates substantial clinical limitations with potentially severe consequences. Tissue hypoxia, when undetected, can rapidly progress to cellular dysfunction, metabolic acidosis, and ultimately irreversible organ damage. This is particularly problematic in critical care settings, whereperfusion heterogeneity in conditions such as sepsis, shock, and acute respiratory distress syndrome necessitates spatial monitoring of oxygen delivery at the tissue level. Similarly, during surgical procedures, the inability to visualize real-time tissue oxygenation may lead to undetected ischemic regions, compromising surgical outcomes and increasing postoperative complications.

[0165] Current gold standard methods for measuring tissue oxygenation, such as blood gas analysis and invasive probes, are either intermittent, provide only localized measurements, or involve significant procedural risks. While advanced imaging modalities including BOLD-MRI and PET can provide oxygen-related metrics, their clinical application is limited by cost, accessibility, and temporal resolution constraints. These limitations highlight the pressing need for a non-invasive, real-time imaging modality capable of mapping StO2 across tissues with high spatial and temporal resolution. Photoacoustic imaging (PAI) has emerged as a promising hybrid technique that combines the spectral selectivity of optical excitation with the spatial resolution of ultrasound detection. Traditional PAI systems employing laser sources have demonstrated remarkable capabilities in visualizing StO2 based on the distinct optical absorption spectra of oxygenated and deoxygenated hemoglobin. However, the widespread clinical adoption of PAI has been hindered by the size, cost, and safety concerns associated with laser-based implementations. LED-based PAI has emerged as a transformative modality for StO2 mapping due to its cost-effectiveness, safety profile, and operational simplicity compared to traditional laser-based systems. However, the inherently low pulse energy of LED sources (300-500 μJ range versus 40-90 mJ in lasers) results in significantly reduced signal-to-noise ratios (SNR), making accurate quantification of StO2particularly challenging. This SNR limitation is especially pronounced when imaging deep-seated structures or attempting to resolve subtle differences in oxygen concentration within tissues as the signal intensities are approximately at the same level as the noise floor. As StO2 is typically estimated via linear spectral unmixing or model-based inversion using dual-wavelength illumination, noise in low-intensity regions affects relative absorption measurements. The unmixing becomes unstable when the denominator of spectral ratios tends to zero, leading to spectral crosstalk.

[0166] To address these fundamental challenges, we present a novel deep learning (DL)-based computational framework Conjoint Frequency & Intensity Regulated UNet (CONFIG-UNet) centered around an attention-dense-residual U-Net architecture specifically engineered for thereconstruction and enhancement of low-intensity PA signals. DL has shown proficiency in enhancing SNR, handling under-sampled scenarios, tackling limited view problems, and reconstructing images from raw data, bypassing conventional beamforming routes. In the context of PA imaging, DL approaches have demonstrated the ability to improve image quality and extract meaningful information from low-quality raw data, making them particularly suitable for addressing the challenges posed by LED-based systems. By leveraging multiple wavelength illuminations (750 nm and 850 nm) targeting the absorption differentials between oxyhemoglobin (HbO2) and deoxyhemoglobin (Hb), the network enables the quantitative reconstruction of StO2maps. Our approach diverges from conventional reconstruction methodologies by integrating spatial attention mechanisms that adaptively prioritize relevant signal features while suppressing noise artifacts. A key innovation in our methodology is the development of an intensity-aware loss function that dynamically adjusts its weighting based on signal intensity distributions, thereby preventing the network from overlooking critical low- intensity regions that often correspond to physiologically significant hypoxic zones. It increases gradient contributions from dim regions during backpropagation helping the network to retain faint vessels, low-contrast tumor regions, etc. This loss function is complemented by a frequency-domain loss component that enforces domain regularization, spectral fidelity across multiple scales, preserving essential frequency characteristics that encode hemoglobin absorption differentials between HbO2and Hb. The network architecture incorporates dense connectivity patterns to maximize information flow throughout the encoding-decoding pathway, enabling robust feature extraction even from severely degraded PA signals. For robust evaluation of reconstruction quality, we implement Chaos-Aware Structural Similarity Metric (CASSIM) and Turbulence-Weighted Image Quality (TWIQ) assessment frameworks into the compile program of our CONFIG-UNet model for training and test monitoring. These advanced metrics transcend traditional quality measures by accounting for the non-linear signal characteristics and inherent instabilities in PA signal propagation, providing more physiologically relevant quality assessments than conventional Peak SNR (PSNR) or Structural similarity index (SSIM) metrics.

[0167] Moreover, our comprehensive technical analysis of the network encompasses multiple dimensions: complexity analysis quantifies the effective capacity and computational efficiency of the model; distribution analysis characterizes the statistical properties of network activations; saliency mapping visualizes critical decision regions; loss landscape analysis ensuresoptimization stability; eigen analysis examines the neural representation space; frequency and wavelet decomposition evaluates multi-scale signal reconstruction fidelity; Lie algebra analysis characterizes the transformation invariance properties; chaos analysis quantifies system stability under perturbations; and Uniform Manifold Approximation and Projection (UMAP) visualization reveals the emergent latent space organization.

[0168] The system's effectiveness has been validated across diverse experimental scenarios with progressively increasing complexity: controlled phantom studies utilizing tubes filled with precisely titrated oxygenated and deoxygenated blood providing ground-truth calibration; in vivo mouse muscle tissue under normoxic (21% O2) and hyperoxic (100% O2) respiratory conditions demonstrating dynamic oxygen response capabilities; murine tumor models exhibiting heterogeneous StO2 distributions characteristic of neoplastic microenvironment; and human extremity imaging of palmar and plantar vascular networks demonstrating clinical translational potential. The proposed algorithm integrates advanced noise suppression techniques, spectral unmixing methodologies, and accelerated reconstruction strategies optimized for the unique characteristics of LED-PAI signal acquisition.

[0169] The clinical implementation of such technology holds transformative potential across numerous medical domains. In critical care, it would enable continuous monitoring of tissue oxygenation in conditions such as septic shock, facilitating early intervention before cellular damage occurs. For surgical applications, real-time visualization of tissue oxygenation could guide procedural decisions and verify tissue viability. In wound care management, objective assessment of tissue oxygenation could predict healing outcomes and guide therapeutic interventions. Furthermore, in cerebrovascular monitoring, mapping of brain tissue oxygenation could provide crucial information in conditions like stroke, traumatic brain injury, and during neurosurgical procedures. By addressing the current limitations in StO2 monitoring, this LED- based PAI approach aims to establish a new paradigm in clinical oxygen assessment, ultimately improving patient outcomes through earlier detection of hypoxic conditions and more precise therapeutic interventions guided by real-time physiological feedback.

[0170] 3.2 Methods

[0171] 3.2.1 Imaging Platform

[0172] All imaging was conducted using Cyberdyne Inc.'s Acoustic X LED-based PA imaging system. The system comprises a 128-element linear array transducer with a 7 MHz centerfrequency and -6 dB bandwidth of 80%. The illumination source is comprised of LEDs emitting at 750 and 850 nm wavelengths, delivering pulses of 70 nanoseconds width with a PRF of 4 kHz, controlled via AcousticX software for multiwavelength oximetry. To compensate for the lower energy of 750 nm, energy compensation weighting is preset to 1.63 times more for the 750 nm signal compared to the 850 nm signal. Consistent gain settings were maintained throughout the experiment, ranging from 60 to 67 dB depending on the sample type (in vitro or in vivo), with the dynamic range set to 19 dB for mouse models and 19-25 dB for metal phantoms. Two types of acquisitions were performed: low-frame averaging acquisition (LA) at 30 Hz, averaging 128 image frames, and high-frame averaging acquisition (HA) at 0.15 Hz, averaging 25,600 frames – ground truth (GT). Post-processing involved converting HA images to binary masks for DL segmentation, applying non-linear median filtering to remove localized and multiplicative granular speckle noise, and performing linear unmixing to generate the StO2 maps. All post- processing and quantitative analyses were conducted using custom MATLAB scripts (R2024a).

[0173] 3.2.2 Testing models: In vitro phantoms, Pre-clinical and Clinical Biology

[0174] For in vitro tube phantom experiments, reconstituted blood from a bovine source was used. This was prepared by dissolving powdered bovine Hb in 1x phosphate-buffered saline to achieve a concentration of 2.5 mM. The prepared blood solution was 100% oxygenated and confirmed using an oxygen sensor (Oxylite Pro, Oxford Optronix).0% oxygenated blood (or deoxygenated blood) was created by adding sodium dithionite (Sigma Aldrich) to the blood at a 2.5 mg / g ratio of blood and gently mixing it to prevent cellular damage. Sodium dithionite dissociates in solution, and the dithionite ion reacts with oxygen and converts oxyhemoglobin to deoxyhemoglobin. The oxygen level of deoxygenated blood was further confirmed with the oxygen sensor.

[0175] The in vivo oxygen saturation imaging was performed on both preclinical and clinical models. Two pre-clinical models were employed. For oxygen saturation tracking in nude mice, the animals were anesthetized with isoflurane and immersed them in a warm water bath. The transducer was immersed in water to maintain acoustic coupling. The flank region closer to the thigh was aligned to the imaging plane using US imaging. The animal was first allowed to breathe medical air (21% oxygen) for two minutes, and then oxygen saturation images were acquired. Following this, the breathing gas was switched to 100% oxygen, and the animal was allowed to stabilize for two minutes, and then oxygen saturation images were saved. Thisconstituted one cycle, and we performed a total of three cycles. The US and PA raw RF files were saved for the entire duration.

[0176] For tumor oxygen saturation imaging, we used a human hypopharyngeal squamous cell carcinoma (FaDu) xenograft model, as these are highly vascularized. Briefly, FaDu cells were suspended in a 1:1 (v / v) mixture of Matrigel (BD Bioscience) and phosphate-buffered saline at a density of 1 million cells / 100 µL and injected subcutaneously in the flank region. Once the tumor reaches a diameter of 8-10 mm (7-10 days post inoculation) presenting a heterogeneous microenvironment, tumor oxygen saturation imaging was performed. The mice were anesthetized and immersed in a warm water bath until the tumor was submerged, and oxygen saturation imaging of different tumor cross-sections was performed.

[0177] In the clinical imaging, the upper left palms of a healthy volunteer were imaged, focusing on the palmar digital arteries and veins to obtain the 2D cross-sectional oxygen saturation images. Additionally, a detailed 3D volumetric scan of the foot’s arcuate arteries were performed, covering a length of 2.25 cm with 0.125 mm step size to ensure high spatial resolution data acquisition. The 3D reconstruction of the foot imaging was performed using Amira (Thermo Fischer Scientific).

[0178] 3.2.3 Process Workflow

[0179] In general, the linear unmixing process (without employing any additional intelligent learning algorithm) involves two wavelengths (750 nm and 850 nm) data neighboring the isosbestic point (~800 nm) constituted by the superposition of the absorption spectra of Oxygenated- (HbO2) and Deoxyoxygenated-hemoglobin (Hb). Based on the optical absorption (cm-1) and molar extinction coefficient (M-1cm-1) values of HbO2and Hb, we can calculate theStO2 value from the following equation:ళఱబ ^ೌఱబ ^ఱబ ళೌఱ^^^^^^ = ఌಹ್ ∗ఓ ିఌಹ್ ∗ఓ బ, where ^^ఒ ఒ ఒଶ ^ఱబ ళೌఱబ ళఱబ ^ೌఱబ ௗ^^^ = −^^ு^ − ^^for Hb and HbO , respectively,denotes the absorption coefficient at the wavelength ^^. Due to LED’s energy outputif the LA dataset is directly use for calculating the StO2map, it would result in a very noisy and incorrect outcome, as shown in the top left portion of FIG.13. However, the mapping can be achieved in real-time. To correctly produce a StO2 map, a highnumber of frame averaging is performed to reconstruct the clean PA images before unmixing them.

[0181] In the CONFIG-UNet DL-enhanced linear unmixing process, data preprocessing is initiated using a Dense-UNet-based CNN DL model to enhance SNR by generating high-quality images from low-energy LA inputs. Subsequently, both sets of denoised wavelength data undergo processing through the DL model described in the subsequent subsection. The resultant outcomes are then subjected to linear unmixing to derive a StO2map. This streamlined workflow removes the need for high frame averaging, as the DL model can produce high SNR images in real-time, typically within the order of ~0.1-0.2 milliseconds. Consequently, the entire process can be seamlessly executed in real-time, offering significant efficiency and practical applicability advantages. The bottom left StO2 map is one of the representatives of our DL-enhanced linear unmixing.

[0182] 3.2.4 CONFIG-UNet details

[0183] A custom DL architecture (CONFIG-UNet) was implemented primarily built on the U- Net framework, but with major enhancements: residual connections, dense blocks, and attention mechanisms (FIG.13). It is designed for multi-wavelength image fusion. The model is a hybrid encoder-decoder architecture with dual input modality fusion, Residual connections for stable training, Dense blocks for efficient feature reuse, Attention gates to suppress irrelevant features during decoding, and multi-scale skip connections for spatial precision. The residual convolution blocks (RCB) at first two network depth-levels and bottleneck tackled vanishing gradient problem, helped in learning identity mapping. The shortcuts ensured feature propagation even if convolutional layers are ineffective at the beginning of training. They helped in keeping low- level features alive by improving convergence and preventing degradation in deep layers. Along with RCB, dense blocks were also sequentially added for local dense feature aggregation. As each layer of the dense blocks receives input from all previous layers, the architecture promoted feature reuse and mitigated overfitting ensuring information flow across layers. They are more parameter-efficient than wide convolutional stacks and facilitate stronger gradients through short paths from early to later layers. Essentially, they acted as local feature extractors before and after wavelength modality fusion. The attention gates suppress background features and highlight salient regions as vanilla skip connections copy all features without prioritizing the feature relevance. The attention blocks helped in segmenting small objects or low-contrast regions byfocusing the decoder on relevant spatial information (gating signal ensured semantic guidance for spatial selection). The BatchActivate layer combined BatchNormalization for normalization across batch to stabilize gradients and ReLU activation for introducing non-linearity in decision boundaries. The layer was used everywhere to promote regularization, prevent covariate shift and encourage faster convergence with clean and consistent activation patterns. The blocks are described below for better clarity:

[0184] 3.2.4.1 Residual convolution block inspired by ResNet

[0185] First, two 3 × 3 convolutions (conv) with optional ^^^^ + ^^^^^^^^.
A shortcut connection(1×1 Conv) is added to skip the main path. The final output is the element-wise sum followed byReLU:
^^ = ^^^^^^^^(^^^^^^^^(^^) + ^^ℎ^^^^^^^^^^^^(^^))

[0186] 3.2.4.2 DenseBlock implements a dense connectivity pattern (like in DenseNet)

[0187] Four sequential Conv blocks (1×1 followed by 3×3), each taking as input the concatenation of all previous outputs. It results in very expressive and information-rich features.Each DenseBlock performs something like:^^1 = ^^^^^^^^ଷ×ଷ(^^^^(^^))^^2 = ^^^^^^^^ଷ×ଷ(^^^^([^^, ^^1 ]))^^3 = ^^^^^^^^ଷ×ଷ(^^^^([^^, ^^1 , ^^2 ]))…

[0188] 3.2.4.3 Gating signal from the decoder features

[0189] A 1 × 1 ^^^^^^^^ + ^^^^ + ^^^^^^^^ that adjusts the depth to match the skip connection. Thisis a pre-processing step for attention.

[0190] 3.2.4.4 A self-learned soft-attention mechanism, adapted from Attention U-Net

[0191] Downsample encoder features ^^௫and align decoder gating features (∅^).

[0192] Add them and then add ^^^^^^^^ → 1 × 1 ^^^^^^^^ → ^^^^^^^^^^^^^^ → ^^^^^^^^^^^^^^^^ to encoder’soriginal size.

[0193] Multiply attention map with encoder features. Project it back to original channel size via1 × 1 ^^^^^^^^.

[0194] This block selectively weights spatial features to focus on relevant regions—useful when merging encoder and decoder features.

[0195] 3.2.4.5 Key Features

[0196] Two Inputs: input_tensor1 (750 nm) and input_tensor2 (850 nm).

[0197] Each input goes through a residual + dense feature extractor.

[0198] Outputs are concatenated early (like early fusion).

[0199] Follows U-Net encoder-decoder structure, with: a) Residual + Dense blocks in the encoder. b) Skip connections passed through attention gates. c) Conv2DTranspose (deconvolution) for upsampling.

[0200] At the bottleneck: another residual fusion block combines both modalities.

[0201] 3.2.4.6 Final Architecture Design

[0202] This is a dual-input U-Net-like architecture with dense, residual, and attention components.4 encoder levels (with downsampling via max pooling) → Bottleneck layer with RCB → decoder levels (with upsampling + attention + DenseBlocks). The DL model outputs a regression map.

[0203] 3.2.5 Other Network Architectures

[0204] Dense-UNet, Attention-Res UNet, Attention UNet, Res UNet, UNet, and cGAN

[0205] 3.2.6 Loss function

[0206] A custom composite loss function was designed to improve the performance of a DL network by addressing both pixel-level intensity characteristics and frequency-domain fidelity. The loss function has two major components - intensity aware loss (IAL) and frequency loss (FL). IAL’s purpose was to emphasize accurate reconstruction in low-intensity regions, which are often more prone to noise or signal loss. As preserving faint signal is more critical, low- intensity pixels were penalized more than high-intensity ones by calculating an inverse dynamic weight map, increasing gradient contributions from dim regions during backpropagation. The loss values were equally balanced with Mean Squared Error (MSE) per pixel to provide global stability. This combination loss function avoids learning priors biased towards the influence of high-intensity regions and succeeds in generalizing to low-intensity textures so that subtle features do not vanish. FL ensured that structural patterns and textures (which appear in frequency space) are preserved, even if pixel-wise MSE might miss them. The magnitude spectra in the frequency domain are compared ignoring any phase differences to focus only on the strength of the frequencies. This helps maintain textures, edges, and periodic patterns that MSE may neglect. It is known that noise tends to populate high-frequency bands stochastically, and true signals, even faint ones, often have structural coherence in frequency space. So, matchingfrequency magnitude helps retain textural and vascular fidelity even in dim regions. The total loss is a combined strategy with tunable focus for balanced trade-off between pixel accuracy and structure. Training was provided and validation loss curves for both wavelength’s captured inputs in FIG.14. Training loss plots are also depicted for all other considered DL models in FIG.15.

[0207] Total loss function is defined as:ℒ௧^௧^^ = ^^^^௧^^^^௧௬ ∙ ℒ^^௧^^^^௧௬ + ^^^^^^ ∙ ℒ^^^^Where:ℒ is the loss function.^^^^௧^^^^௧௬ ∈ [0,1] is the weighting factor for the intensity-aware loss.^^^^^^ ∈ [0,1] is the weighting factor for the frequency loss.For our case, ^^^^௧^^^^௧௬ = 0.99, ^^^^^^ = 0.01

[0208] IAL is defined as:1ℒ = ^ [(1 − ^^) ∙ ^^ + ^^ ௧^௨^ ^^^ௗ ଶ^^௧^^^^௧௬ ^^ ^^ ] ∙ (^^^^ − ^^^^ )Where:^^ ∈ [0,1] be the proportion of unweighted MSE (e.g., 0.5).^^௧^௨^^^ , ^^^^^ௗ^^ be the true and predicted image intensities at pixel location (^^, ^^), respectively.^^ 1^^ =ห

[0209] FL is defined as1ℒ = ∙ ^^ห^^^௧^௨^ห − ^^^^ௗ^^^^ ^^ ห^^ ห^Where:|∙| denotes magnitude of the complex valued FFT at spatial frequency (^^, ^^).^^ is the number of frequency bins.ห^^^௧^௨^ห i ௧^௨^௨௩ s ^^^^^^2(^^ ))can be depicted as:1∙^^^ [(1 − ^^) ∙ ^^ + ^^] ௧ ^^^ௗ 1ℒ = ^^ ^௨^ ଶ௧^௧^^ ^^௧^^^^௧௬ ^^ ∙ (^^^^ − ^^^^ ) + ^^^^^^ ∙^,^ ^^

[0211]

[0212] of extreme signal degradation, two novel and complementary evaluation metrics were introduce: Chaos- Aware Structural Similarity (CASSIM) and Turbulence-Weighted Image Quality (TWIQ). These metrics are specifically designed to capture complex perceptual and spectral distortions that are often overlooked by traditional scalar image quality indicators such as PSNR and SSIM, particularly in low-SNR regimes like LED-based PAI. By integrating these two metrics in training and testing phase, a multi-domain, physics-aware evaluation strategy tailored for LED- based PAI is enabled. These four image quality metrics evaluations were integrated into the model.compile part of CONFIG-UNet and plotted in FIG.16.

[0213] CASSIM: In low-intensity PAI, noise manifests not only as additive distortion but often exhibits chaotic spatial patterns and spatiotemporal instability due to physiological motion, LED fluctuations, and nonlinear reconstruction artifacts. Standard metrics like SSIM assume locally stationary distortions, and hence may fail to penalize spatial chaos or irregularity effectively. To address this, SSIM was augmented with a chaos-penalizing factor derived from a Lyapunov Exponent-inspired measure in CASSIM that quantifies local sensitivity to small perturbations across spatial dimensions. By exponentially penalizing SSIM in proportion to Lyapunov exponent, CASSIM provides a more sensitive gauge of structured noise, particularly beneficial incapturing failure modes in low-intensity signals.^^^^^^^^^^^^(^^௧^௨^ , ^^^^^ௗ) = ^^ିఒ ∙ ^^^^^^^^(^^௧^௨^ , ^^^^^ௗ)^^ is the Lyapunov exponent, defined as: ^ଶ (^^^^௫ + ^^^^௬); ^^^^௫ = ^^^,^[|^^(^^ + 1, ^^) − ^^(^^, ^^)|];^^^^௬ = ^^^,^[|^^(^^, ^^ + 1) − ^^(^^, ^^)|]; ^^(^^, ^^) = |^^௧^௨^(^^, ^^) − ^^^^^ௗ(^^, ^^)|. The spatial sensitivity iserror gradients across the spatial axes.

[0214] TWIQ: Traditional error metrics (e.g., MSE) operate in spatial domain, often failing to account for the spectral degradation of structural information due to turbulence-like distortions caused by low fluence. In LED-PAI, the frequency content of the image—generally mid-to-high frequency components—correlates with fine spatial features like vessel boundaries and oxygengradients. TWIQ quantifies deviation in the energy spectrum between the ground truth and the prediction, offering a frequency-domain sensitivity metric. A lower TWIQ indicates higher fidelity in spectral energy preservation, implying better retention of spatial textures and fine features. It is particularly informative when evaluating subtle structural recovery in low-intensityzones, where local textures can be disproportionately affected.^^^^^^^^(^^௧^௨^ , ^^^^^ௗ) = ห^^^^^ೠ^ − ^^^^^^^หWhere average spectral ^^ ^ spectrum is: ^^^(^^, ^^) =|ℱ{^^}(^^, ^^)|ଶ, ℱ{^^} is the Fourier transform of the image space.

[0215] 3.2.8 Network analysis

[0216] 3.2.8.1 Frequency Analysis of Feature Maps Across Convolutional Layers

[0217] To understand how different frequency components propagate through the convolutional layers of our DL model - CONFIG-UNet, a frequency analysis was performed on feature maps extracted from a trained model at various depths, including both convolutional layers and skip connections. This analysis was carried out over multiple training epochs to observe the evolution of frequency characteristics as the model learned.

[0218] After each epoch, the test dataset was passed through the model to evaluate frequency characteristics of intermediate activations. For each test sample, intermediate feature maps were extracted from all convolutional layers and skip connections. Each feature map was processed using a 2D Fast Fourier Transform (FFT) to obtain its frequency spectrum. The FFT results were shifted to center the zero-frequency component. The transformed spectrum was segmented into high- and low-frequency regions. Low frequencies were taken from the bottom-left quadrant and the high frequencies were extracted from the top-right quadrant. The relative energy (as a fraction of the total spectrum energy) in each band was computed. Average high- and low- frequency energy values were recorded for each convolutional layer and each skip connection. These values were stored for comparison across layers and across epochs.

[0219] 3.2.8.2 Eigenvalue Spectrum and Energy Retention Analysis of Noisy vs. Denoised Features

[0220] To evaluate the signal richness and information compression characteristics of the denoised outputs from our CONFIG-UNet model, a principal component spectrum analysis was conducted using eigenvalue decomposition of the covariance matrices for both noisy and denoised image representations.

[0221] First, for each test sample, noisy PA image was normalized and passed through the model to obtain its denoised counterpart. Both noisy and denoised images were converted to grayscale and cast to double precision. Covariance matrices were computed along columns (i.e., pixel-wise correlation across rows). Eigenvalue decomposition was performed on these covariance matrices to obtain eigenvalues representing the variance captured by each principal component (i.e., each eigenvector). Eigenvalues were normalized to the [0, 1] range to enable comparative analysis ^^^^^ೞ^ss samples. The energy retention ratio was computed as:∑^ ^ ^acro^^^^^ೞ^where ^^^represents the ^^- ^ th eigenvalue.

[0222] 3.2.8.3 Training dynamics and representation complexity

[0223] To investigate how the internal complexity of our DL model evolves during progressive denoising training, two complementary metrics were systematically tracked over training epochs: compressed model size and weight entropy. These measures captured the informational richness and compressibility of the learned parameters, respectively, serving as proxies for network capacity utilization and representational complexity. When combined with performance metrics such as SSIM, PSNR or any other task-specific losses, it offers a deeper understanding of whether improved accuracy is being achieved through genuinely richer internal representations or by overfitting / redundant parameterization.

[0224] Compressed model size: It serves as a practical proxy for its Kolmogorov complexity— the minimal description length needed to represent the network’s parameters. Specifically, the saved models were compressed after each epoch using GZIP to approximate the entropy-coded storage. Smaller compressed sizes indicate simpler representations and reduction of parameter redundancy. If the model retains high performance along with high compressibility, then it generally indicates efficient learning.

[0225] Weight entropy: This metric computes the Shannon entropy of the model’s learnable parameters. After flattening and concatenating all weight tensors (excluding biases and scalars), a histogram with 256 bins over normalized weight values was generated. Finally, we calculatedthe entropy using ^^ = −∑^^(^^) logଶ(^^(^^)), where ^^(^^) describes the empirical distribution acrossweight magnitudes. Low entropy in the weights suggests sparse, structured, and compressible parameter distributions. It generally indicates that the model has converged to a stable solution with less redundancy. This is especially valuable in low-resource deployment, e.g., edge devices,real-time systems. Overparameterized networks often memorize noise, inflating entropy. Lower entropy might signal better generalization, especially if test performance improves as entropy decreases. Also, models with lower entropy are often easier to prune, quantize, or compress without much loss in performance.

[0226] 3.2.8.4 Layer-Wise Feature Distribution Analysis via Wasserstein Distance and Activation Sparsity

[0227] To evaluate how feature distributions evolve across convolutional layers in the trained CONFIG-UNet model, a distributional similarity and sparsity analysis was performed on a test dataset using intermediate feature maps.

[0228] For each pair of adjacent convolutional layers, activations were averaged across channels to reduce dimensionality and highlight global spatial differences. The Wasserstein Distance or Earth Mover's Distance was computed between the flattened distributions of two consecutive layer outputs which quantifies how much the feature distribution shifts from one layer to the next. It quantifies the distributional shift between feature maps of adjacent layers by measuring the minimal effort required to transform one distribution into another. The activation sparsity metrics (^^1 & ^^2) were used to assess how compact or spread out the activations are in each layer. ^^1 sparsity denotes the mean absolute value of activations per layer and is sensitive to zero-centeredness. ^^2 sparsity is the root mean squared energy of activations, penalizing high- amplitude activations.

[0229] 3.2.8.5 Wavelet-Based Feature Decomposition Analysis Across Convolutional Layers

[0230] To further investigate the evolution of spatial-frequency characteristics in feature maps throughout the CONFIG-UNet architecture, a detailed wavelet decomposition analysis was conducted using the Haar wavelet basis. This analysis aimed to identify how spatial features such as edges, textures, and global structures evolve and propagate, particularly in relation to skip connections.

[0231] For each test image and each convolutional layer, individual feature maps (channels) were isolated and passed through a 2D discrete wavelet transform (DWT) using the Haar basis. For each feature map, we extracted the wavelet coefficients: cA - Approximation (low-pass), cH - Horizontal, cV - Vertical and cD - Diagonal details. The absolute values of these coefficients were averaged across all feature maps for each layer to produce scalar indicators of averagespatial-frequency content. These per-layer averages were aggregated across all the test images for statistical consistency

[0232] 3.2.8.6 Saliency-Based Analysis for Signal Localization

[0233] To interpret and validate the attention of our trained DL-based denoising model two wavelength-captured image inputs, a two-pronged saliency analysis strategy was employed combining Gradient-weighted Class Activation Mapping (Grad-CAM) and Occlusion sensitivity mapping. These techniques enabled visualization of spatial regions contributing most to the network’s output, helping distinguish true signal content from noise.

[0234] Grad-CAM computation: For each input pair of PA and US images, Grad-CAM was applied by constructing a sub-model targeting a deep convolutional layer. Using TensorFlow’s automatic differentiation (GradientTape), gradients of the model's output with respect to the target feature maps were computed and spatially pooled. These pooled gradients were then weighted with the feature maps to generate a coarse localization map (heatmap), highlighting regions most influential to the model’s output. The heatmap was normalized to enhance interpretability.

[0235] Occlusion sensitivity mapping: In parallel, an occlusion sensitivity map162-164 was generated by systematically masking localized patches of both wavelengths-captured PA inputs and observing the change in the model's output. The change in prediction (mean squared difference from the unoccluded output) was used to quantify the importance of each region. The occlusion map was normalized for visualization.

[0236] Hybrid saliency fusion: To improve spatial precision and robustness of the interpretation, the Grad-CAM and occlusion maps were linearly combined with a weighting factor (α=0.6) to produce hybrid saliency maps. This fusion preserved both the semantic relevance captured by Grad-CAM and the perturbation sensitivity measured by occlusion mapping.

[0237] 3.2.8.7 Visualization of Loss Landscape in Weight Space

[0238] To assess the smoothness, curvature, and overall geometry of the optimization landscape of our trained denoising network, a 3D loss landscape analysis was performed by visualizing the loss surface in a plane defined by two random, orthonormal directions in parameter space. This approach helps in evaluating the stability and generalization properties of the model, particularly across different training epochs.

[0239] The trainable parameters of CONFIG-UNet were flattened into a single high-dimensional weight vector, denoted as ^^∗. This vector serves as the central point of analysis in the parameter space. Two perturbation directions, ^^^and ^^ଶ, were randomly sampled from a standard Gaussiandistribution and normalized to unit length where ^^ ^^^ = ‖^^‖, ^^^ is length parameter in some ^^ −thdirection. These directions span a 2D plane in thedimensional weight space centered at ^^∗. The loss surface was then probed across this plane. We evaluated the model's performance over amesh grid of perturbation scalars ^^, ^^ ∈ [−500,500] using ^^ఈ,ఉ = ^^∗ + ^^^^^ + ^^^^ଶ.

[0240] At each (^^,^^) point, the model weights were updated accordingly and computed the model output on a held-out two wavelengths-captured image pair. The loss was then evaluated using a custom-designed loss function between the predicted and ground truth outputs. Theresulting 2D grid of loss values defines the local geometry of the optimization landscape around^^∗.

[0241] The collected loss values were visualized as a 3D surface using Matplotlib, where the x- and y-axes correspond to the perturbation directions and the z-axis represents the computed loss. The landscape is indicative of the flatness, sharpness, or multi-modal structure around the current solution

[0242] 3.2.8.8 Lie Algebra-Based Stability Analysis of Feature Representations in CONFIG- UNet

[0243] To evaluate the intrinsic structural behavior of feature representations across different layers of our trained DL model, a Lie bracket analysis was conducted on features extracted from early, bottleneck, and late convolutional layers. The goal was to examine the geometric and algebraic stability of features in response to clean and noisy inputs, leveraging tools from Lie algebra and Frobeniüs norm-based metrics.

[0244] Feature maps were extracted from each of the three layers (Early convolutional layer, Bottleneck, and last convolutional layer) for both clean and noisy input image pairs, as well as denoised outputs obtained from the model. Due to the high dimensionality of convolutional feature maps, Principal Component Analysis (PCA)) was applied to reduce the feature dimensions to 500 while preserving the majority of variance in the data. This allowed efficient computation of pairwise Lie brackets between the reduced features. Given two reduced feature vectors ^^ & ^^ corresponding to clean and noisy (or denoised) inputs respectively, the Liebracket were computed as: [^^, ^^] = ^^^^் − ^^^^். This bracket captures the non-commutativenature of transformations in feature space. The Frobeniüs norm of the resulting bracket matrix serves as a stability indicator, measuring how aligned or misaligned the representations are.

[0245] 3.2.8.9 UMAP-Based Feature Space Analysis Across Network Depth

[0246] To gain insight into the feature evolution within our denoising network, Uniform Manifold Approximation and Projection (UMAP) was applied to visualize and quantify the progression of learned representations at three key stages of the network: an early convolutional layer, the bottleneck layer, and the final convolutional layer. For each of these layers, feature maps were extracted from the network when processing three categories of input images: noisy, denoised (model output), and clean (ground truth).

[0247] For feature extraction, a custom intermediate-layer Model was defined using TensorFlow / Keras to output activations from a chosen convolutional layer. Feature maps were flattened into 2D arrays before being passed into UMAP. UMAP was employed with neighborhood sizes of 5 and 10 and used Euclidean distance in a 2D and 3D projection space.

[0248] 3.2.8.10 Evolution of Feature Representations Across Network Depth and Training Epochs

[0249] To analyze the evolution of feature representations in our DL network, intermediate features were extracted and compared from the CONFIG-UNet model at three distinct stages: early, bottleneck, and late convolutional layers. For this, a feature extractor was constructed from the trained model to access the output of a specific convolutional layer.

[0250] For each training epoch, the model's checkpoint weights were loaded, and corresponding noisy and clean image pairs were passed through the network as primary and auxiliary inputs. Feature maps were extracted separately for both noisy and clean inputs from the selected layer. These features were flattened, and their mean intensity values were computed.

[0251] 3.3 Results

[0252] 3.3.1 Evaluation of DL outcomes with tube phantoms

[0253] 3.3.1.1 Comparative study of denoising performance

[0254] To rigorously evaluate the real-time denoising efficacy of our proposed model, a controlled experiment was conducted using optically tunable tube phantoms filled with known concentrations of oxygenated (HbO2– 100%) and deoxygenated hemoglobin (Hb – 0%). This setup allowed us to generate repeatable, well-characterized PA signals under known conditions. Dual wavelength low number of frame averaging (LA) captured images, acquired at 750^nm and850^nm, were used as input to all DL models under investigation. The tubes were oriented transversely to the receiving surface of the transducer array, ensuring that both tubes were imaged simultaneously with similar noise profiles. As shown in FIG.17A (I–XVV), the model outputs were visually compared, with low-intensity signal regions of interest (ROIs) clearly marked by white arrows. These ROIs, representing the challenging zones with signal levels close to the noise floor, served as critical evaluation zones for each model’s denoising ability. For a more precise comparison, intensity line profiles were extracted along two vertical axes—denoted by blue and white lines—across the low-signal regions for both wavelengths and presented the results in FIG.17B. To enhance visibility, zoomed-in insets of the ROIs were included for our model and other top-performing networks (such as Attention-Res-7), alongside their corresponding LA inputs and high number of frame averaging (HA) ground truth (GT) images. At 750^nm, which yielded relatively cleaner inputs, most DL models performed adequately; however, at 850^nm—where signal degradation and stochastic noise were more pronounced— many models exhibited excessive false activations or failed to recover the low-intensity structures accurately. In contrast, the CONFIG-UNet model demonstrated robustness across both wavelengths, consistently recovering true signal structures while minimizing background noise. This qualitative assessment was further substantiated through quantitative metrics (FIG.17C): CONFIG-UNet achieved the most consistently high PSNR and SSIM scores, particularly in the low-signal zones, where many other networks either over-smoothed or amplified noise artifacts. These results underscore the superior generalization and signal-preserving capabilities of our model in denoising challenging PA inputs. With this foundation established, the downstream effect of denoising performance on StO2estimation was next evaluated, an essential clinical parameter derived from dual-wavelength PA images.

[0255] 3.3.1.2 Comparative study of StO2 estimation

[0256] To assess the accuracy and robustness of StO2 estimation across DL models, the controlled phantom experiment was leveraged, wherein the precise concentrations of HbO2and Hb were known a priori. This allowed computation of the ground-truth StO₂ values for benchmarking purposes. FIG.17D (I–X) presents a comprehensive comparison of StO₂ maps obtained from: (i) LA captured input reconstructions, (ii) HA captured GT images, and (iii) various DL-generated outputs based on the LA inputs. In these visualizations, white dotted circles highlight cross-sectional regions of the tubes—specifically the low-illumination zones,which are further emphasized by white arrows. The vertical illumination setup resulted in non- uniform photon fluence, with the upper surfaces of the tubes receiving more photons and consequently yielding stronger PA signals, while the lower regions suffered from photon attenuation and reduced signal intensity. This spatially varying illumination created intra-tube heterogeneity in StO2 levels, thereby offering a critical testbed to evaluate the DL models' capability to infer accurate saturation levels from partially degraded inputs.

[0257] Notably, the proposed model demonstrated a clear advantage by accurately reconstructing StO2 maps in both well-illuminated and poorly illuminated regions. This superior performance can be attributed to its robust low-intensity signal recovery capability, as established in the prior denoising evaluation. In contrast, other models tended to either over- smooth or misclassify the low-signal zones, leading to biologically implausible saturation estimates. While qualitative inspection provides initial insight, we further substantiated our findings through a rigorous quantitative framework. Specifically, confusion matrices were constructed for each method by discretizing the StO2 range into five classes: (1) top of HbO2- filled tubes (90–100% StO2), (2) bottom of HbO2-filled tubes (70–85%), (3) top of Hb-filled tubes (0–15%), (4) bottom of Hb-filled tubes (20–35%), and (5) other background or noise regions. The distinction between top and bottom zones arises from differential light fluence and optical absorption properties within the tubes. As shown in the resulting confusion matrices, while most DL models performed adequately in high-signal regions (e.g., top of the tubes), only CONFIG-UNet accurately captured the subtle saturation variations in the low-signal bottom regions. These results reinforce the critical importance of denoising fidelity, especially in challenging imaging scenarios, for accurate physiological parameter estimation.

[0258] 3.3.2 StO2evaluation for pre-clinical biology with mouse models

[0259] 3.3.2.1 Oxygen challenge in mouse muscle

[0260] Building upon the promising results with tube phantoms, the DL approach was next validated in a more complex and physiologically relevant setting. The ability to accurately measure tissue oxygenation in living systems is crucial for understanding various physiological and pathological processes, from muscle metabolism to tumor hypoxia. Therefore, in vivo StO2 imaging experiments were conducted on mouse thigh muscles under controlled oxygen challenges. This step is vital in bridging the gap between controlled phantom studies and potential clinical applications, allowing us to assess the performance of our DL model in adynamic biological system. FIGS.18A-C presents the results of our in vivo StO2 imaging on mouse thigh muscles under alternating oxygen concentrations. The mice were subjected to cycles of normal air (21% oxygen) and pure oxygen (100% oxygen) breathing, mimicking physiological changes in tissue oxygenation that might occur in various clinical scenarios. This experimental design allows us to evaluate the accuracy of our imaging methods in detecting and quantifying real-time changes in tissue oxygenation.

[0261] The first and third columns of FIG.18A (I, V, IX, III, VII, XI) showcase US co- registered StO2 maps generated using LA, HA, and our DL method when the mouse was breathing 100% oxygen. The second and fourth columns (II, VI, X, IV, VIII, XII) illustrate the same imaging methods applied when the mouse was breathing 21% oxygen (normal air). While we can reasonably infer that vascular oxygenation status will follow the administered oxygen concentrations, the relationship may not be strictly linear due to various physiological factors. Nevertheless, these repeated cycles provide us with a good estimate of the actual oxygenation status of the blood vessels in the mice. To quantitatively assess the performance of each method across multiple animals and cycles, FIG.18B presents a scatter plot of the oxygenation statuses observed with LA, HA, and DL processes. Each data point represents the mean oxygenation status, with error bars indicating the standard deviation. Three distinct shapes were used to represent datasets from different mice, allowing us to account for inter-individual variability. The StO2quantification reveals a consistent trend of falsely elevated oxygenation status in the LA imaging. This overestimation aligns with the observations from the tube phantom experiments, further highlighting the limitations of LA in accurately representing tissue oxygenation in real- time imaging scenarios. In contrast, both the HA and DL methods demonstrate more consistent and physiologically plausible oxygenation measurements. The DL method shows remarkable agreement with the gold-standard HA method but with the crucial advantage of real-time processing capability. This agreement is evident across different mice and oxygen challenge cycles, underscoring the robustness and reliability of our DL approach in dynamic in vivo settings.

[0262] 3.3.2.2 Subcutaneous mouse tumor

[0263] Building upon our findings from the oxygen challenge experiments in healthy mouse muscle, the DL imaging techniques were next applied to another clinically relevant scenario: tumor oxygenation. Accurate assessment of tumor oxygenation is crucial in cancer research andtreatment planning, as it can provide valuable insights into tumor physiology, predict treatment response, and guide therapeutic strategies. Tumors often exhibit heterogeneous oxygenation patterns due to their abnormal vasculature and metabolic demands, making them an ideal test case for evaluating the spatial accuracy of our StO2 mapping techniques. FIG.18C presents StO2 maps of subcutaneous tumors in two different mice, showcasing the capabilities of LA, HA, and DL approaches in capturing the complex oxygenation landscape of these tumors. This comparison allows us to assess how well our DL model performs in a setting where rapid, high- resolution imaging is particularly valuable for understanding tumor biology and potentially guiding interventions. The first row of FIG.18C (I - III) displays US co-registered StO2maps obtained using the three approaches. This side-by-side comparison allows for a direct assessment of the spatial resolution, contrast, and overall quality of StO2 mapping achieved by each method within the tumor microenvironment. To demonstrate the reproducibility of our findings across different subjects and tumor locations, the second row (IV - VI) of FIG.18C shows a similar set of images for a different cross-section of another mouse's tumor. All images are overlaid with US data to validate the spatial alignment of PA signals with the underlying tissue morphology as visualized by the US to localize regions of interest within the tumor and correlating StO2values with specific anatomical features. While providing real-time capabilities, the LA captured images show a noisier and less accurate representation of tumor oxygenation. In contrast, the HA GT images reveal a more detailed and nuanced picture of the tumor's oxygenation landscape, capturing areas of hypoxia and regions of higher oxygenation with greater clarity. Similarly, CONFIG-UNet produces StO2 maps that closely resemble those generated by the HA method. The denoised StO2maps produced by both HA and DL methods reveal intricate oxygenation patterns within the tumors. We can observe regions of heterogeneity, including well-oxygenated areas likely corresponding to functional blood vessels (indicated by green arrows in figures), as well as hypoxic zones that may indicate areas of poor perfusion or high metabolic activity (indicated by white arrows). These detailed oxygenation maps could provide valuable information for real-time understanding of tumor biology. The real-time imaging capability of LED-based PA systems is precious for the rapid assessment of tumor oxygenation during image- guided interventions and for monitoring acute changes in tumor physiology in response to treatments.

[0264] 3.3.3 StO2 evaluation for in vivo clinical biology

[0265] 3.3.3.1 Human palm & wrist

[0266] Having demonstrated the efficacy of our DL approach in pre-clinical models, the performance of CONFIG-UNet was evaluated in a human subject. The human palm, with its complex network of superficial blood vessels and relatively thin overlying tissue, presents an ideal model for assessing the translational potential of our method. Moreover, accurate StO2 mapping of peripheral circulation could have significant implications for diagnosing and monitoring various vascular conditions, making this a particularly relevant test case for the technology. FIG.19A (I-III) presents the results of StO2 imaging performed on the left palm of a human subject, focusing on the superficial palmar vessels. The common palmar digital arteries extending towards the middle, ring, and little fingers were specifically targeted, areas that are readily accessible for imaging and represent important conduits in the hand's vascular network. The inset schematic in first column panel provides anatomical context, with light illumination around our regions of interest. This top row displays US overlayed StO2maps of the space between the little and ring fingers, generated using LA, HA, and our DL processes, respectively. The arrows indicate skin and vessel cross-section structures that visually correspond to the common palmar digital arteries depicted in the anatomical schematic. This correlation between our imaging results and known vascular anatomy provides a crucial validation of the method's ability to localize and characterize blood vessels in human tissue accurately. Consistent with our observations in pre-clinical models, the LA captured inputs tend to be noisy and falsely elevates StO2levels across the entire region. This overestimation could lead to potential misinterpretations in a clinical setting, highlighting the limitations of current real-time imaging approaches. In contrast, both the HA GT and DL outcomes clearly delineate the anatomical vessel structures, providing a more nuanced and accurate representation of tissue oxygenation. The second row presents a similar imaging setup for the human palmar wrist region to demonstrate the reproducibility of the findings herein across different regions of the hand. Apart from skin, we clearly observed artery and veins (as indicated by the arrows in panel V) in both HA GT and DL outcomes. The consistency of results across different palm and wrist regions underscores the reliability of the CONFIG-UNet DL approach in capturing the oxygenation patterns in human tissue. By accurately capturing the oxygenation status of superficial blood vessels, our method could provide valuable insights into peripheral circulation, potentially aidingin the diagnosis and monitoring of various vascular conditions, such as peripheral artery disease, diabetic neuropathy, or Raynaud's phenomenon.

[0267] 3.3.3.2 Human foot

[0268] Next, the technique was tested by applying it to a different complex anatomical structure: the human foot. Accurate mapping of foot vasculature and oxygenation have profound implications for diagnosing and monitoring various conditions, including peripheral artery disease, diabetic foot complications, and wound healing processes. To this end, an in vivo 3D scanning of the upper right foot of a healthy volunteer was conducted, focusing on the arcuate arteries and their subsidiaries. This region, depicted in the schematic of FIG.19B (panel I), is of particular interest due to its complex vascular architecture and critical role in foot perfusion. Specifically, we aimed to image the branches of the arcuate artery and its subsidiaries before they merge into the dorsal metatarsal artery system, providing a comprehensive view of foot vasculature that could be valuable in both research and clinical settings. The first row of FIG. 19B (I-III) presents a comparative analysis of the 2D US co-registered StO2 maps (blood vessels and skin) generated by the three methods (LA, HA, and CONFIG-UNet). This co-registration is crucial for providing anatomical context to the oxygenation data, allowing for more accurate localization of vascular structures within the foot's complex anatomy. The second row of FIG. 19B (IV-VI) showcases snapshots of the 3D views produced by each method, offering a direct comparison of their capabilities in capturing the complex vascular structure of the foot. Consistent with our previous observations, the LA method produces images that are noisy and falsely over-oxygenated, obscuring the underlying vascular architecture. In stark contrast, both the HA and DL methods clearly delineate the arcuate artery, as indicated by the white arrow in panel V. The ability to pinpoint specific vascular structures is a key advantage of these methods, offering the potential for more precise diagnoses and treatment planning in clinical scenarios. In 3D visualizations, vein structures are prominently observable indicating areas of lower oxygenation. These veins vary in positioning, with some located superficially and others embedded in deeper tissue layers. The 3D rotational videos provide a comprehensive view of the entire vascular network, offering a more complete perspective compared to a single 2D cross- sectional frame. The striking differences in image quality and vessel visibility between the LA captured input, HA GT, and DL outcomes underscore our DL approach's significant advancement. By achieving image quality comparable to the gold-standard HA GT whilemaintaining real-time processing capabilities, our DL technique opens new possibilities for rapid, detailed assessment of foot vasculature and oxygenation. This real-time 3D StO2 imaging capability can be used in clinical scenarios requiring quick decision-making, such as assessing tissue viability in diabetic foot ulcers or monitoring reperfusion following revascularization procedures. Moreover, it could enhance our understanding of dynamic changes in foot perfusion during various activities or in response to treatments, potentially informing both research endeavors and clinical practice across a range of vascular conditions.

[0269] 3.3.4 Network analysis

[0270] 3.3.4.1 Frequency Analysis of Feature Maps Across Convolutional Layers

[0271] As the depth of the network increased, low-frequency components tended to accumulate, suggesting that deeper layers emphasized coarse, global features. Conversely, high-frequency components were progressively suppressed, indicating the removal or attenuation of fine-grained details or noise. In contrast to the general convolution layers, skip connections exhibited a sharp increase in high-frequency energy and a steep drop in low-frequency energy (FIG.20A(I-IV)). This suggests that skip connections reintroduce high-frequency information — such as fine edges and textures — that might have been lost in the downsampling path. This behavior aligns with the functional role of skip connections in architectures like U-Net, where they serve to preserve spatial resolution and fine detail during the reconstruction phase. Overall, the analysis reveals a frequency flow pattern in convolutional neural networks with skip connections. Downsampling (or contracting) path emphasizes low-frequency, global features, possibly for better abstraction and generalization. The skip connections restore high-frequency, local details, which are crucial for accurate reconstruction or segmentation. The upsampling path integrates both low- and high- frequency features to generate the final output. Such frequency-aware insights not only will help interpret model behavior but also guide architectural improvements — for example, in designing better skip connections or attention mechanisms that explicitly manage frequency content.

[0272] 3.3.4.2 Eigenvalue Spectrum and Energy Retention Analysis of Noisy vs. Denoise Features

[0273] Across qualifying samples, the retained energy in the denoised output was consistently below 50%, confirming that the denoising network compresses or filters a significant portion of the high-variance (often noisy) content (FIG.20B (I-IV)). The eigenvalue decay over increasing factor numbers (principal component indices) revealed distinct behaviors. For noisy inputs,eigenvalues decreased slowly, suggesting that noise is distributed across a broader set of principal components. For denoised outputs, eigenvalues dropped off more rapidly, indicating a concentration of variance in a smaller number of dominant components—consistent with successful noise suppression and dimensionality reduction. The sharp decay in the denoised spectrum implies that the network effectively emphasizes informative structures while discarding dispersed, unstructured noise.

[0274] This analysis quantitatively demonstrates the CONFIG-UNet denoising model’s capacity to compress signal content into a low-dimensional representation. Furthermore, the faster eigenvalue decay in clean features indicates a more structured and low-rank representation, which is a desirable trait for downstream tasks such as classification or reconstruction. The retained energy metric also serves as a potential proxy for denoising aggressiveness—a balance point between noise removal and detail preservation.

[0275] 3.3.4.3 Training dynamics and representation complexity

[0276] We observed that as training progresses, the model exhibits a reduction in weight entropy and compressed representation size while maintaining or improving denoising fidelity (SSIM, PSNR, CASSIM & TWIQ). This suggests a transition toward more structured and efficient (information-rich) representations without compromising expressiveness and overfitting (FIG. 20C (I-IV)).

[0277] 3.3.4.4 Layer-Wise Feature Distribution Analysis via Wasserstein Distance and Activation Sparsity

[0278] The majority of convolutional layers exhibit near-zero Wasserstein distances and sparsity levels, indicating stable, smooth transformation across depth. A large spike at the first skip connection, suggesting significant feature transformation when early-layer spatial details are re- injected. Smaller but detectable spikes at subsequent skip connections, implying further but more or less similar dramatic changes due to additional concatenation or merging operations (FIG. 20D (I-VI)). These spikes likely reflect the fusion of earlier, high-resolution spatial features with deeper, abstract representations, disrupting the otherwise smooth feature evolution. The skip connections act as informational resets, injecting spatial detail and causing noticeable changes in activation distributions. This pattern indicates that skip connections introduce non-trivial changes to the feature distribution and activation sparsity, disrupting the otherwise gradual transformation flow through the network. These results highlight the structural role of skip connections in theCONFIG-UNet architecture: they inject rich, high-resolution features back into deeper layers, resulting in abrupt changes in feature map distribution and sparsity. The magnitude of these spikes corresponds to the amount of information reintroduced and the semantic gap being bridged.

[0279] 3.3.4.5 Wavelet-Based Feature Decomposition Analysis Across Convolutional Layers

[0280] Across most convolutional layers in the network, the average wavelet coefficients (cA, cH, cV, cD) remained close to zero, indicating a consistent suppression or smoothing of spatial- frequency components during standard feature propagation (FIG.20E (I-VIII)). Exceptionally large spikes in all four coefficient types were consistently observed at the first skip connection layer. This indicates a substantial reintroduction of spatial details (especially low-frequency and edge-related structures) into the network through skip connections. Subsequent skip connection layers exhibited smaller yet distinct spikes, suggesting that while spatial-frequency information is re-injected multiple times, its magnitude and novelty diminish progressively. These findings are consistent across all the test samples, emphasizing the robustness of the observation. A notable exception was found in the horizontal detail component (cH), which exhibited additional smaller spikes in non-skip layers, i.e., between skip connections. These intermediate spikes were not clearly aligned with structural skip pathways, indicating the presence of horizontal edge structures being preserved or introduced by local dense blocks or attention-enhanced convolutions. This could also reflect anisotropic propagation of spatial information in the network—where horizontal detail is retained or emphasized differently than vertical or diagonal components.

[0281] This wavelet-domain analysis confirms that skip connections play a pivotal role in injecting spatial detail back into the network pipeline. The first skip connection appears to carry the most significant spatial-frequency information, while later ones provide supplementary adjustments. These results demonstrate that while skip connections are the primary mechanism for spatial detail recovery, certain internal convolutional operations may also selectively enhance specific directional features—particularly horizontal patterns. This may reflect architectural nuances such as dense connectivity or attention modules that can preserve or amplify specific spatial cues even in deeper layers.

[0282] 3.3.4.6 Saliency-Based Analysis for Signal Localization

[0283] Across all samples, the resulting saliency maps consistently concentrated over signal-rich regions while de-emphasizing noise-dominated areas, validating the model’s ability to correctly localize and distinguish true signal content from background noise (FIG.20F (I-IV)). These maps serve as a visual explanation of the model's denoising behavior and support the hypothesis that the model relies on physiologically relevant features for restoration.

[0284] 3.3.4.7 Visualization of Loss Independence in Weight Space

[0285] By generating and comparing loss landscapes after different training stages (e.g., after 20 and 40 epochs), a clear difference in the smoothness and geometry of the optimization landscape was observed (FIG.20G (I-IV)). The loss valley exhibited higher curvature, localized sharp regions, and asymmetries, indicating that the model was still in a relatively unstable region of the loss surface, potentially vulnerable to overfitting or poor generalization. The loss landscape became noticeably smoother and flatter in the neighborhood of the solution. This suggests that the model had transitioned to a more stable and robust minimum, potentially associated with better generalization performance and reduced sensitivity to perturbations in the weights

[0286] 3.3.4.8 Lie Algebra-Based Stability of Feature Representations in CONFIG-UNet

[0287] Early Layers exhibited higher Frobeni^^^ s norms of Lie brackets compared to deeperlayers (FIG.20H (I-VI)). This suggests that early representations are more sensitive to input perturbations and carry richer spatial diversity, possibly due to lower-level textures and noise sensitivity. The bottleneck layer showed intermediate Frobeniüs norms, reflecting compressed but still moderately sensitive feature representations, likely containing the most abstracted semantic information. Late layers demonstrated lowest Frobeniüs norms, indicating more stable and noise-invariant features, consistent with the network’s objective of reconstruction and denoising. Interestingly, when noisy and clean inputs were analyzed, the corresponding Liebrackets were well-aligned (i.e., had similar Frobeni^^^ s norms) across layers. This suggests thatthe network preserves a consistent internal representation structure, even when the input quality varies. This Lie algebra-based analysis provides a quantitative lens into the stability and geometric evolution of features in convolutional networks. It reinforces the understanding that early layers are more susceptible to noise, while deeper layers tend to abstract and stabilize the representation. Moreover, the alignment of Lie brackets between noisy and clean inputs emphasizes the robustness of the model's learned feature space.

[0288] 3.3.4.9 UMAP-Based Feature Space Analysis Across Network Depth

[0289] At early layers (shallow depths), the UMAP embeddings showed large separations between noisy, denoised, and clean feature clusters (FIG.20I). Both the noisy-to-denoised and denoised-to-clean distances were high, indicating minimal semantic alignment at this stage. As the features passed through the network's bottleneck, we observed a significant reduction in the distance between denoised and clean clusters. However, the noisy-to-denoised distance remained relatively high, suggesting that the network had not yet fully corrected for noise but was moving closer to a clean representation. At the final convolutional stage, the UMAP embeddings of noisy, denoised, and clean images collapsed into a nearly overlapping distribution. The distances between all three groups—particularly denoised-to-clean—approached zero, indicating successful signal reconstruction and structural alignment. To further visualize the progression, directional arrows were added in both 2D and 3D UMAP plots showing the trajectory of individual samples from noisy → denoised → clean. These arrows consistently demonstrated a converging pattern across all samples. The same arrangement was also plotted in 2D for neighborhood sizes 5 and 10 and 3D UMAP trajectory for neighborhood size 10.

[0290] This multi-layer UMAP analysis confirms that the network not only suppresses noise but also learns a progressive mapping toward clean representations, with convergence emerging gradually through depth. The results were qualitatively consistent across both neighborhood sizes and dimensionalities (2D and 3D), demonstrating the robustness of the learned transformation trajectory.

[0291] 3.3.4.10 Evolution of Feature Representations Across Network Depth and Training Epochs

[0292] The analysis revealed that in the early layers, feature maps derived from noisy images exhibited substantial dissimilarity from those of clean images, reflecting the strong influence of noise on initial representations (FIG.20J). In the bottleneck layer, the difference between noisy and clean features was notably reduced, indicating that the network had learned to suppress noise-related artifacts while preserving core signal information. By the final layers, the feature differences between noisy and clean inputs were minimal, suggesting that the network successfully reconstructed clean representations from noisy inputs by this stage. This progression highlights how the network incrementally learns to bridge the feature disparity between noisy and clean data as information flows from early to deep layers, achieving effective denoising through hierarchical representation learning.

[0293] 3.4 Discussion

[0294] LED-based PA imaging for StO2 measurement presents a compelling alternative to traditional laser-based illumination systems, offering significant advantages in terms of safety, affordability, and portability. These attributes align with the growing need for accessible, real- time medical imaging technologies. However, the inherent drawback of low energy delivery from LEDs gives rise to significant noise in StO2maps, compromising their interpretability if we conduct real-time imaging. Consequently, adopting a high number of frame averaging emerges as a viable strategy to mitigate noise and calculate accurate StO2 maps, albeit at the expense of real-time processing capabilities. This limitation has been a significant barrier to the widespread adoption of LED-based PA imaging despite its potential benefits.

[0295] This study directly addresses these challenges by introducing CONFIG-UNet, a novel approach that integrates a CNN-based DL technique with an adaptive intensity-aware and frequency-adjusted loss function to enhance SNR especially restoring very intensity signals (almost at the same levels as the noise floors) before performing linear unmixing. The development of CONFIG-UNet builds upon the recent advancements in applying DL to PA imaging. Previous studies have demonstrated the potential of DL in enhancing PA image SNR, reconstruction, and analyses. This approach is unique in restoring very low intensity signals of LED-based PA systems and its focus on real-time StO2 mapping. The efficacy of the DL-based denoising model was comprehensively evaluated in recovering meaningful PA signals under realistic, low SNR conditions, and its subsequent impact on the accuracy of derived parameters such as StO2. Through controlled phantom experiments with known concentrations of oxygenated and deoxygenated hemoglobin, we demonstrated that our model significantly outperformed conventional and state-of-the-art DL-based denoising methods, both in terms of signal restoration fidelity and in preserving quantitative oxygen saturation integrity. The denoising performance evaluation, centered around dual-wavelength PA inputs (750 nm and 850 nm), revealed important insights into model robustness under varying noise profiles. Notably, while all models performed reasonably well at 750 nm—where the input signal was inherently less stochastic—substantial performance disparities emerged at 850 nm, where input signals were severely degraded by noise. This wavelength, known to produce lower SNR due to higher absorption and scattering, highlighted the inability of several DL models to distinguish low- intensity signals from background noise, resulting in blurred reconstructions and loss ofstructural integrity. In contrast, the CONFIG-UNet model consistently preserved weak signal regions across both wavelengths, as visualized through zoomed-in regions and intensity line profiles. The fidelity of the recovered signals in our model, especially in low-intensity areas indicated by white arrows, was not only visually superior but also corroborated quantitatively through higher PSNR and SSIM scores. These results suggest that the model does not merely overfit to global noise trends but is capable of learning contextually relevant priors to preserve spatially sparse, yet diagnostically meaningful signals—an essential feature for real-world clinical deployment in low-fluence regimes.

[0296] Building on this foundational capability, it was further investigated how these denoising improvements translate into outcomes by assessing the accuracy of StO2estimation. Since StO2quantification in PAI is inherently sensitive to spectral fidelity and signal consistency across wavelengths, any degradation in signal quality—particularly in low-intensity regions—can directly impair quantitative interpretation. The experiments showed that most DL models yielded plausible StO2 maps in high-signal areas (e.g., top of the tubes), yet significantly failed to preserve saturation gradients in the bottom parts of the tubes, where light fluence and hence PA signal was minimal due to absorption and geometric attenuation. This failure was prominently seen as over-smoothed or misclassified StO2values, which can lead to clinically misleading interpretations in scenarios such as tumor hypoxia mapping or ischemia detection. In stark contrast, our model reliably preserved the spatial heterogeneity of oxygen saturation within the same tube, suggesting it effectively mitigates fluence-related degradation and retains the underlying chromophore distribution. This robustness was further quantified using a multiclass- stratified confusion matrix analysis, where we discretized StO2into relevant bins to probe model performance at various saturation levels. The results clearly established that while high-SNR regions were accurately predicted by most models, low-SNR zones were only faithfully reconstructed by our model.

[0297] Together, these findings underscore that denoising strategies in PAI must not only focus on reducing perceptual noise but also on retaining subtle, low-intensity signal patterns that encode vital physiological information. Evaluating denoising performance purely via image similarity metrics is insufficient; physiological task-specific validation (e.g., StO2estimation) is essential for real-world applicability. Moreover, the results highlight the importance of designing noise-robust architectures that can generalize across both spectrally and spatially heterogeneousimaging conditions. The superior performance of the CONFIG-UNet model in both denoising and downstream quantitative analysis reflects its potential as a practical solution for real-time, point-of-care PAI, especially in resource-constrained or intraoperative environments where acquisition constraints often preclude long averaging durations. The model not only restores signal fidelity under extreme noise conditions but also preserves physiological accuracy, setting a new benchmark for intelligent PAI enhancement pipelines.

[0298] The CONFIG-UNet model also demonstrates robust performance across multiple experimental contexts, including dynamic oxygen challenges in murine muscle, subcutaneous tumor imaging with histological validation, and real-time imaging of human peripheral vasculature. These results represent a major advancement in intelligent PAI enhancement, especially under the challenging constraints of low-fluence, LED-based PA imaging systems. In vivo validation under controlled oxygen challenge in murine thigh muscle represents a critical step in assessing physiological robustness. By alternating inspired oxygen between 21% and 100%, we introduced cyclical perturbations to vascular oxygenation. The DL method accurately tracked these fluctuations in StO2, achieving near-parity with HA-derived gold standard maps while operating in real-time. Notably, LA methods consistently overestimated oxygenation, a trend confirmed by both phantom and in vivo muscle data. This overestimation is likely due to the elevated noise floor, which distorts the PA amplitude ratios critical for spectral unmixing. By contrast, the DL method maintained both the absolute and relative oxygenation trends across multiple mice and oxygen cycles, demonstrating superior repeatability, and resilience to hemodynamic variability. The scatter plot analysis further quantifies these trends, showing significantly reduced variability and standard deviation in the DL-derived oxygenation estimates. This suggests not only improved signal fidelity but also enhanced physiological specificity—an essential attribute for translational application in both functional and molecular imaging domains.

[0299] Extending the method to subcutaneous tumor models introduced additional challenges due to intrinsic spatial heterogeneity, chaotic vasculature, and localized hypoxia typical of solid tumors. These conditions test the spatial resolving power and physiological accuracy of any imaging modality. DL-enhanced StO2maps closely replicated HA GT, capturing both rim- associated hyperoxic zones and core-localized hypoxia. Crucially, these observations were notmerely visual correlations: histological validation using CD31 and pimonidazole staining revealed strong spatial correspondence with StO2 distributions.

[0300] Translation of this methodology to human subjects further establishes its potential for clinical deployment. The superficial palmar vasculature, with its known anatomical layout and minimal overlying tissue, provides an ideal testbed. Here, DL-enhanced StO2 maps accurately localized the common palmar digital arteries and produced physiologically reasonable oxygenation values consistent with expectations from healthy peripheral perfusion. LA maps, by contrast, demonstrated diffuse, artifact-prone signal inflation across tissue boundaries, again underscoring their susceptibility to SNR-induced quantification errors. The anatomical validity of DL-derived StO2maps—corroborated by co-registered ultrasound—demonstrates reliable vessel-specific oxygenation tracking. This has significant implications for real-time vascular diagnostics, peripheral arterial disease screening, and microvascular function assessment. Furthermore, the LED-based acquisition system used in these studies facilitates point-of-care applicability, opening doors for widespread deployment in resource-limited settings where traditional laser-based PA systems may not be feasible. However, as evident from HA and DL supplementary 3D videos, some residual noise might appear as streaks when we denoise each frame in the 3D setup. The segmentation and denoising can be inconsistent across frames due to temporal coherence mismatch, where some frames are denoised better than others, and noise levels in those frames become prominent when stacked into a 3D volume. The model introduces subtle artifacts during denoising (e.g., streaks or blotches that aren't immediately noticeable in 2D frames). These artifacts can accumulate and become more visible when the frames are stacked into a 3D volume. Any slight differences in noise handling or artifact introduction across frames led to a non-uniform distribution of noise artifacts in the final 3D reconstruction, which manifested as scattered streaks. Another possibility could be structural misalignment between frames due to unavoidable biological motion and variability in the underlying structures between frames (e.g., due to physiological changes). Moreover, most noise levels depend on the PA signal intensity (e.g., higher noise in low-signal areas), and noise can be more challenging to distinguish from the actual signal in areas with low contrast or subtle gradients.

[0301] By enabling real-time, high-quality StO2mapping, our approach opens new possibilities across a range of medical applications. These include intraoperative tissue oxygenation monitoring, which could significantly improve surgical outcomes by providing immediatefeedback on tissue perfusion. Additionally, the method could enhance the rapid assessment of vascular diseases, potentially improving the diagnosis and management of conditions like peripheral artery disease and diabetic foot ulcers. In the realm of oncology, the real-time evaluation of tumor oxygenation could provide invaluable insights for cancer research and treatment planning, potentially guiding adaptive radiotherapy strategies and assessing the efficacy of anti-angiogenic therapies. Moreover, the combination of affordability and portability offered by LED-based systems with the advanced capabilities of our DL approach could significantly expand access to sophisticated tissue oxygenation imaging in resource-limited settings. This aligns with global health initiatives aimed at reducing disparities in medical technology access. The potential for widespread deployment of this technology could have far- reaching implications for early disease detection and management in underserved populations.

[0302] The success of the workflow presented herein is fundamentally dependent on the effectiveness of intelligent segmentation and denoising processes, particularly when considering the limitations posed by linear unmixing with only two wavelengths. Actual deployment of DL algorithms as post-processing tools in PA imaging for real-time output needs to address several challenges primarily related to computational efficiency, latency, and system integration. Optimized lightweight models (similar to advanced MobileNet or SqueezeNet) can significantly reduce the computational burden to operate on limited hardware settings. Pruning and Quantization might further optimize the models without substantial loss of accuracy, leading to faster inferencing. The DL algorithm should also be tightly coupled with the PA imaging system’s software and hardware involving custom software development and firmware optimization to minimize any overhead. One of the most important steps in the deployment process will be feedback mechanism in real-time by implementing adaptive algorithms that will adjust parameters and address bottlenecks immediately, ensuring consistent and accurate outcomes during live imaging sessions. Additionally, expanding the number of wavelengths used could significantly enhance the accuracy and optimization of measurements. The transition from preclinical studies to clinical applications often reveals new challenges.

[0303] Different tissue types and anatomical structures might also ensure that the model generalizes well across a broader range of clinical scenarios, addressing any potential biases introduced by specific tissue characteristics. Integrating CONFIG-UNet with other imaging modalities, such as ultrasound or MRI, could enhance the accuracy and depth of oxygensaturation mapping. Exploring these multi-modal approaches could lead to more comprehensive diagnostic tools, especially in complex vascular studies. The DL model can be deployed in many healthcare settings, including resource-limited environments, as it does not require costly medical hardware.

[0304] As used in this specification and the claims, the singular forms “a,” “an,” and “the” include plural forms unless the context clearly dictates otherwise.

[0305] As used herein, “about”, “approximately,” “substantially,” and “significantly” will be understood by persons of ordinary skill in the art and will vary to some extent on the context in which they are used. If there are uses of the term which are not clear to persons of ordinary skill in the art given the context in which it is used, “about” and “approximately” will mean up to plus or minus 10% of the particular term and “substantially” and “significantly” will mean more than plus or minus 10% of the particular term.

[0306] As used herein, the terms “include” and “including” have the same meaning as the terms “comprise” and “comprising.” The terms “comprise” and “comprising” should be interpreted as being “open” transitional terms that permit the inclusion of additional components further to those components recited in the claims. The terms “consist” and “consisting of” should be interpreted as being “closed” transitional terms that do not permit the inclusion of additional components other than the components recited in the claims. The term “consisting essentially of” should be interpreted to be partially closed and allowing the inclusion only of additional components that do not fundamentally alter the nature of the claimed subject matter.

[0307] The phrase “such as” should be interpreted as “for example, including.” Moreover, the use of any and all exemplary language, including but not limited to “such as”, is intended merely to better illuminate the invention and does not pose a limitation on the scope of the invention unless otherwise claimed.

[0308] Furthermore, in those instances where a convention analogous to “at least one of A, B and C, etc.” is used, in general such a construction is intended in the sense of one having ordinary skill in the art would understand the convention (e.g., “a system having at least one of A, B and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description or figures, should beunderstood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B.”

[0309] All language such as “up to,” “at least,” “greater than,” “less than,” and the like, include the number recited and refer to ranges which can subsequently be broken down into ranges and subranges. All ranges include each individual member. Thus, for example, a group having 1-3 members refers to groups having 1, 2, or 3 members. Similarly, a group having 6 members refers to groups having 1, 2, 3, 4, or 6 members, and so forth. A range of from 1 to 100 includes 1 and 100 and every value therebetween.

[0310] The modal verb “may” refers to the preferred use or selection of one or more options or choices among the several described embodiments or features contained within the same. Where no options or choices are disclosed regarding a particular embodiment or feature contained in the same, the modal verb “may” refers to an affirmative act regarding how to make or use an aspect of a described embodiment or feature contained in the same, or a definitive decision to use a specific skill regarding a described embodiment or feature contained in the same. In this latter context, the modal verb “may” has the same meaning and connotation as the auxiliary verb “can.”

[0311] The present disclosure has described one or more preferred embodiments, and it should be appreciated that many equivalents, alternatives, variations, and modifications, aside from those expressly stated, are possible and within the scope of the invention.

Claims

Claims 1. A system for improving signal-to-noise ratio (SNR) in real-time in photoacoustic imaging (PAI) comprising: an excitation probe configured to deliver light to a tissue; an acoustic detection probe configured to receive photoacoustic signals at a first frame acquisition rate from the tissue in response to excitation of the tissue by the light; a processor configured to: receive the photoacoustic signals; reconstruct the photoacoustic signals by performing a low number frame averaging (LA) on the photoacoustic signals to produce LA image data; process the LA image data in real-time using a deep learning (DL) model, wherein the SNR of the processed LA image data is statistically identical to a SNR of a high number frame averaging (HA) of the photoacoustic signals acquired at a second frame acquisition rate; generate one or more output images from the processed LA image data; and a display configured to display the one or more output images.

2. The system of claim 1, wherein the first frame acquisition rate is 200 times greater than the second acquisition frame rate and the LA is 200 times smaller than the HA.

3. The system of claim 1, wherein statistically identical is a probability-value (p-value) less than or equal to 0.

05.

4. The system of claim 1, wherein the SNR of the output images is at least three times greater than the SNR of the LA image data.

5. The system of claim 1, wherein the DL model includes a U-Net generator comprising a plurality of blocks in each of an encoding pathway and a decoding pathway.

6. The system of claim 5, wherein the plurality of blocks in the encoding pathway include: a convolution layer and a rectified linear unit (ReLU) activation in a first block of the plurality of blocks; a convolution layer, a batch normalization, and a ReLU activation in a second block and third block of the plurality of blocks; and a convolution layer, a batch normalization, a ReLU activation, and a dropout layer in a remainder of the plurality of blocks.

7. The system of claim 6, wherein a plurality of blocks in the decoding pathway each include a transposed convolution layer, a batch normalization, and a ReLU activation, and each block except for a final three blocks include a dropout layer.

8. The system of claim 7, wherein the U-Net generator further includes attention-dense residual blocks in the encoding pathway and decoding pathway.

9. The system of claim 5, wherein the DL model further includes a discriminator.

10. The system of claim 9, wherein the discriminator includes a plurality of blocks.

11. The system of claim 10, wherein the plurality of blocks of the discriminator include: a convolution layer and a rectified linear unit (ReLU) activation in a first block of the plurality of blocks;a convolution layer, a batch normalization, and a ReLU activation in each block between the first block and a final block; and a convolution layer, a batch normalization, and a leakyReLU activation in the final block.

12. A system for monitoring oxygen saturation (StO2) in real-time using photoacoustic imaging (PAI), the system comprising: an excitation probe configured to deliver light to a tissue, wherein the light includes two or more wavelengths; an acoustic detection probe configured to receive photoacoustic signals at a first frame acquisition rate from the tissue in response to excitation of the tissue by each of the two or more wavelengths; a processor configured to: receive the photoacoustic signals from each of the two or more wavelengths; reconstruct the photoacoustic signals from each of the two or more wavelengths by performing a low number frame averaging (LA) on the photoacoustic signals to produce LA image data; process the LA image data from each of the two or more wavelengths in real-time using a deep learning (DL) model, wherein a SNR of the reconstructed image data is statistically identical to a SNR of a high number frame averaging (HA) of the photoacoustic signals for each of the two or more wavelengths acquired at a second frame acquisition rate; generate one or more output images from the processed LA image data for each of the two or more wavelengths; perform a linear unmixing on the one or more output images of each of the two or more wavelengths; generate a StO2map from the linear unmixing of the output images; anda display configured to display the StO2 map.

13. The system of claim 12, wherein the first frame acquisition rate is 200 times greater than the second acquisition frame rate and the LA is 200 times smaller than the HA.

14. The system of claim 12, wherein statistically identical is a probability-value (p-value) less than or equal to 0.

05.

15. The system of claim 12, wherein the SNR of the output images is at least three times greater than the SNR of the LA photoacoustic signals.

16. The system of claim 12, wherein the DL model includes an encoder and a decoder and is configured to receive an input from each of the two or more wavelengths.

17. The system of claim 16, wherein DL model further includes residual connections, dense blocks, and attention mechanisms.

18. The system of claim 17, wherein the DL model comprises: for each LA photoacoustic signal input for each of two or more wavelengths, a first encoder layer including a residual convolution block, a batch normalization, a ReLU activation, a dense block a skip connection, and downward max pooling, a second encoder layer including a residual convolution block and a dense block, a skip connection, and downward max pooling,a third and fourth encoder layer, each including a dense block, a skip connection, and downward max pooling, a bottleneck layer including a residual convolution block, a skip connection, and downward max pooling; a first, second, and third decoder layer, each including an attention gate and upward deconvolution; and a fourth decoder layer including an attention gate and a convolution layer.

19. The system of claim 18, wherein each decoder layer includes a convolution layer, a dense block, batch normalization, and ReLU activation.

20. The system of claim 19, wherein each residual convolution block and each dense block includes a plurality of convolution sublayers, batch normalizations, and ReLU activations.

21. The system of claim 12, wherein the DL model is trained using a loss function including an intensity aware loss (IAL) component, a frequency loss (FL) component, and a mean squarederror (MSE) component, wherein the loss function is defined asℒ௧^௧^^ = ^^^^௧^^^^௧௬ ∙ ℒ^^௧^^^^௧௬ + ^^^^^^ ∙ ℒ^^^^where ℒ is the lossthe IAL, and ^^^^^^∈[0,1] is a weighting factor for the FL.

22. The system of claim 21, wherein the IAL is defined as1− ∙ ^^ ^^ ଶ= ^ ௧^௨^ ^ௗ^^ + ∙ ^^ − ^^^^ ൯where ^^ ∈ [0,1] is a proportion of unweighted MSE, ^^௧^௨^ ^^^ௗ^^ and ^^^^ are a true and a predictedimage intensity at pixel location (^^, ^^), respectively, and ^^23. The system of claim 21, wherein the FL is defined as 1ℒ = ∙ ^^ห^^^௧^௨^ห − ห^^^^ௗ^^^^ ^^ ௨௩ ^^௨௩ ห^where|∙|denotes aTransform (FFT) at spatialfrequency (^^, ^^), ^^ is a number of frequency bins, ห^^^௧^௨^௨௩ ห is ^^^^^^2(^^௧^௨^), and ห^^^^^^ௗ௨௩ ห is^^^^^^2(^^^^^ௗ).

24. The system of claim 12, wherein the DL model is evaluated using evaluation metrics including a Chaos-Aware Structural Similarity (CASSIM) metric, a Turbulence-Weighted Image Quality (TWIQ) metric, a Peak Signal to Noise Ratio (PSNR), and a Structural Similarity Index (SSIM).

25. The system of claim 24, wherein the CASSIM metric is defined as^^^^^^^^^^^^(^^௧^௨^ , ^^^^^ௗ) = ^^ିఒ ∙ ^^^^^^^^(^^௧^௨^ , ^^^^^ௗ)where ^^ is a1|]26. The system of claim 24, wherein the TWIQ metric is defined as^^^^^^^^(^^௧^௨^ , ^^^^^ௗ) = ห^^^^^ೠ^ − ^^^^^^^หwhere average spectral energy ^^ = ^ ∑ ^^(^^, ^^) , power spectrum ^ ( ) | { }( )|ଶ^ ே ௨,௩ ^ ^^ ^^, ^^ = ℱ ^^ ^^, ^^ ,and ℱ{^^} is a Fourier transform27. The system of claim 24, wherein the PSNR is defined as max(Target)^^^^^^^^ =^^^^^^^^^௨^ௗwhere max(Target) is a maximum intensity data of a region of interest (ROI) in the two or more output images and ^^^^^^^^^௨^ௗis a standard deviation of a selected background region of the two or more output images.

28. The system of claim 24, wherein the SSIM is defined as (2^^^^,^^ = ^^^^ + ^^^)(2^^^ + ^^ )SSIM( ) ^ ଶ^^ଶ ^^ ଶ ଶ^where ^^^is a sample(ROI) of the two or more output images, A, ^^^ଶis a variance of A, ^^^^is a covariance of A and a second image array B of the ROI, ^^^is defined as [^^^^^]ଶand ^^ଶis defined as [^^ଶ^^]ଶ, where ^^^and ^^ଶare 0.01 and 0.03, respectively, and L is a dynamic range of the ROI.

29. The system of claim 12, wherein the two or more wavelengths include a first wavelength at 750 nm and a second wavelength at 850 nm.

30. The system of claim 29, wherein the linear unmixing is defined as ^^^^^^ = ^^^^1 ^^2 ^^2 ^^1^^^^ ∗^^^^ −^^^^^^ ∗^^^^^^ ^ ^^where ^^ఒ = ^^ఒ − ^^ఒ ఒு^ைమ , ^^ு^coefficients forhemoglobin (Hb) and oxygenated hemoglobin (HbO2), respectively, at the two or morewavelengths ^^^^, and ^^^ఒdenotes an absorption coefficient at the two or more wavelengths ^^^^, and n is a number of a of the two or more wavelengths.

31. A method for improving signal-to-noise ratio (SNR) in real-time in photoacoustic imaging (PAI) comprising: delivering light to a tissue using an excitation probe; detecting photoacoustic signals using a detection probe at a first frame acquisition rate from the tissue in response to excitation of the tissue by the light; receiving the photoacoustic signals via a processor; reconstructing the photoacoustic signals by performing a low number frame averaging (LA) on the photoacoustic signals to produce LA image data; processing the LA image data in real-time using a deep learning (DL) model, wherein the SNR of the processed LA image data is statistically identical to a SNR of a high number frame averaging (HA) of the photoacoustic signals acquired at a second frame acquisition rate; generating one or more output images from the processed LA image data; and displaying, via a display, the one or more output images.

32. The method of claim 31, wherein the light includes two or more wavelengths.

33. The method of claim 32, wherein the processor reconstructs the photoacoustic signals, processes the LA image data, and generate one or more output images for each of the two or more wavelengths.

34. The method of claim 33, further comprising performing linear unmixing, via the processor, of the one or more output images from each of the two or more wavelengths, generating an oxygen saturation (StO2) map from the linear unmixing of the output images via the processor; and displaying the StO2map on the display.

Citation Information

Patent Citations

  • Pulmonary nodule data enhancement method based on RAU-GAN

    CN116091885A

  • Infrared weak and small target detection method based on space-time attention fusion coding

    CN117809151A

  • Low-noise photoacoustic imaging method and system based on dual-channel data acquisition and noise cancellation

    CN118236032A

  • Photoacoustic imaging of inflamed tissue

    EP3375353A1