Quantitative photoacoustic image method and system based on deep learning

By introducing the TransUNet model in photoacoustic imaging technology, combined with the advantages of Transformer and U-Net, the shortcomings of quantitative oxygen saturation imaging in the prior art are solved, and higher accuracy and robust oxygen saturation reconstruction are achieved.

CN120070412AActive Publication Date: 2025-05-30CHONGQING MEDICAL UNIVERSITY

Patent Information

Application Number
CN202510283971.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-05-30
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing photoacoustic imaging technology has shortcomings in quantitative blood oxygen saturation imaging. The traditional method assumes that the optical characteristics of the tissue are uniform, it is difficult to accurately reflect the complexity of biological tissues, and is highly sensitive to noise and data sampling density. Deep learning models are mostly limited to local feature extraction, making it difficult to fully utilize the global correlation of multi-wavelength photoacoustic data, and exhibit limited generalization capabilities when crossing data domains.

Method used

Using a quantitative photoacoustic image method based on deep learning, the TransUNet model is constructed. This model adds a Transformer module to the U-Net network structure to capture the long-distance dependence of the data in the dataset, and uses the multi-head attention mechanism to process multi-wavelength data to improve the modeling ability of global information.

Benefits of technology

The quantitative accuracy and robustness of blood oxygen saturation are improved, the problem of traditional methods being sensitive to noise and data density is overcome, the model's generalization ability across data domains is enhanced, and the higher precision reconstruction of blood oxygen saturation is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070412A_ABST
    Figure CN120070412A_ABST
Patent Text Reader

Abstract

The invention discloses a quantitative photoacoustic image method and system based on deep learning, a Transform module is constructed, the model combines the advantages of Transform and U-Net, and the modeling ability of global information is improved through a deep learning method. The problem that the quantitative precision of the blood oxygen saturation is limited due to the fact that an existing deep learning technology cannot fully capture the relevance between different wavelengths during multi-wavelength data processing is solved, meanwhile, the problem that the generalization ability of a deep learning model is insufficient during cross-data-domain (such as simulation and experimental data) is solved, and the accuracy of the blood oxygen saturation is improved. And the adaptability of the model to actual clinical application is improved. By comprehensively optimizing the photoacoustic imaging quantification process, the reconstruction with higher precision, higher robustness and wider applicability for the oxyhemoglobin saturation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of photoacoustic image processing and analysis, and particularly to a quantitative photoacoustic image method and system based on deep learning. Background Art

[0002] The existing photoacoustic imaging technology has the following deficiencies in quantitative blood oxygen saturation imaging. Traditional physics-based modeling methods usually assume uniform tissue optical properties and are difficult to accurately reflect the complexity of biological tissues. These methods are highly sensitive to noise and data sampling density, and the reconstruction accuracy decreases significantly under sparse sampling and limited view conditions. At the same time, the computational complexity is relatively high, making it difficult to meet the real-time requirement.

[0003] In recent years, deep learning methods have been introduced to solve the above problems. However, existing deep learning models are mostly limited to local feature extraction and are difficult to fully utilize the global correlation of multi-wavelength photoacoustic data.

[0004] In addition, the performance of deep learning models depends on high-quality labeled data and computing resources, and shows limited generalization ability when crossing data domains (such as simulation and experimental data). These limitations still pose challenges to the accurate quantification of blood oxygen saturation and its wide clinical applications in the existing technology.

[0005] Therefore, a quantitative photoacoustic image method is needed to overcome the limitations of traditional physics-based modeling methods, such as the assumption of uniform tissue optical properties. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide a quantitative photoacoustic image method based on deep learning, which solves the problem of difficult accurate inversion of optical absorption coefficient and blood oxygen saturation in complex biological tissues.

[0007] To achieve the above purpose, the present invention provides the following technical solutions: The quantitative photoacoustic image method based on deep learning provided by the present invention includes the following steps: S1. Photoacoustic data preparation: Obtain simulation data and experimental data, and the photoacoustic data includes photoacoustic image data of different wavelengths; S2. Division of model training set and test set: Divide the data set into a training data set, a validation data set, and a test data set; S3. Construct a neural network model: The neural network model uses the TransUNet model, and a Transformer module is added to the U-Net network structure to capture the long-range dependencies of the data in the data set through the Transformer module; S4. Model training: Use the training set to train the TransUNet model; optimize the model hyperparameters; S5. Model validation and evaluation. Use the test set to validate the model and calculate the model performance metrics. S6. Complete image processing through the trained TransUNet model. Input the photoacoustic image for absorption coefficient estimation and output the processed image or absorption coefficient distribution map.

[0008] Furthermore, the TransUNet model includes an input module, an encoder module, a Transformer module, a decoder module, and an output module. The multi-channel input module is used to input photoacoustic image data of different wavelengths in the form of different channels. The encoder module includes multiple cascaded convolutional blocks for extracting local features of the input data. Each convolutional block includes multiple convolutional layers, batch normalization layers, activation functions, and max pooling layers. Feature extraction is performed by gradually downsampling through each convolutional block. The Transformer module is used to capture global dependencies and correlations between different wavelengths. The decoder module is used to gradually restore the resolution of the feature map and fuse low-level and high-level features to obtain a two-dimensional distribution map representing the absorption coefficient and blood oxygen saturation. The output module is used to output the reconstructed quantitative photoacoustic image, including the distribution maps of the absorption coefficient and blood oxygen saturation.

[0009] Furthermore, the Transformer module includes several Transformer layers. Each Transformer layer includes a multi-head attention mechanism module and a feed-forward neural network. The multi-head attention mechanism module is used to calculate the feature dependencies and global correlations between wavelengths. The feed-forward neural network is used to further process and transform the features.

[0010] Furthermore, the decoder module includes several cascaded upsampling blocks for performing multiple upsampling operations. A skip connection is connected after each upsampling block to fuse features with the corresponding layer of the encoder to obtain a two-dimensional distribution map representing the absorption coefficient and blood oxygen saturation.

[0011] Furthermore, the photoacoustic data in step S1 is preprocessed in the following manner: S11: Data form: Loaded as multi-channel data with the shape of (H, W, C). H and W are the spatial dimensions of the image, and C represents the number of channels of different wavelengths. S12: Normalization: The data of each wavelength is respectively normalized to the interval [0, 1].

[0012] Furthermore, the specific steps of the Transformer module are as follows: Step 1: Calculate the query, key, and value, and generate three vectors for each input feature: a query vector, a key vector, and a value vector; Step 2: Calculate the attention weights: Obtain the attention weights by calculating the similarity between the query and the key; Step 3: Cross-wavelength correlation: For multi-wavelength data, each wavelength corresponds to an input channel, and calculate the correlation between different wavelengths through the self-attention mechanism.

[0013] Furthermore, the multi-head attention mechanism module performs parallel calculations using the multi-head attention mechanism, specifically according to the following formula: where, is the th attention head, is the output weight matrix.

[0014] Furthermore, the decoder module outputs the absorption coefficient and according to the following formula: where, represents the output, that is, the output absorption coefficient and the two-dimensional distribution of blood oxygen saturation .

[0015] Furthermore, the loss function in the TransUNet model is calculated according to the following formula: where L is the joint loss function, Lmse is the mean square error, used to evaluate the difference between the predicted value and the true value, α and β are weight coefficients; absorption represents the absorption coefficient; represents the blood oxygen saturation.

[0016] The present invention also provides a quantitative photoacoustic imaging system based on deep learning, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above method is implemented.

[0017] The beneficial effects of the present invention are as follows: A quantitative photoacoustic imaging method based on deep learning provided by the present invention constructs a Transformer module. This model combines the advantages of Transformer and U-Net, and improves the ability to model global information through deep learning methods. It overcomes the problem that existing deep learning technologies fail to fully capture the correlation between different wavelengths during multi-wavelength data processing, resulting in limited quantitative accuracy of blood oxygen saturation. At the same time, this method also solves the problem of insufficient generalization ability of deep learning models across data domains (such as simulation and experimental data), and improves the adaptability of the model to actual clinical applications. By comprehensively optimizing the quantitative process of photoacoustic imaging, the reconstruction of blood oxygen saturation with higher accuracy, higher robustness and wider applicability is achieved. This method has the following characteristics: 1. More accurate quantitative effect: Many traditional algorithms assume that the initial photoacoustic pressure distribution is linearly related to the optical absorption coefficient. However, in reality, due to the scattering and attenuation of light in tissues, this assumption is inaccurate, especially in deep tissues where the error will be significantly amplified, which may result in inaccurate estimation of the optical absorption coefficient and blood oxygen saturation. The multi-channel input of the TransUNet model can effectively model the mutual relationship between different wavelengths, thereby improving the reconstruction accuracy of photoacoustic images and the estimation accuracy of blood oxygen saturation.

[0018] 2. Stronger global understanding ability: Combining the advantages of Transformer and U-Net: The traditional U-Net model is widely used in medical image processing, especially in image segmentation tasks in medical imaging. It relies on convolutional operations for local feature extraction. Although the U-Net model performs well in image reconstruction and quantitative tasks, it usually has difficulty capturing long-distance global dependencies in images. By introducing Transformer into the U-Net model, TransUNet introduces a self-attention mechanism on the basis of local convolutional operations, enabling the network to capture long-distance dependencies. The Transformer module is particularly good at processing global context information, thereby effectively improving the model's understanding of global features. Especially in multi-wavelength data tasks, it can better capture the feature correlations between different wavelengths.

[0019] 3. Higher computing performance and efficiency: In terms of computing efficiency, although traditional convolutional neural networks (such as U-Net) perform well in image reconstruction tasks, their computing costs are relatively high. Especially when dealing with high-resolution images or large-scale data, the training speed and inference efficiency are usually slow. The parallel computing ability of the Transformer module enables TransUNet to improve computing efficiency while maintaining high accuracy. The multi-head self-attention mechanism can process multiple data channels simultaneously, enhancing the training and inference speed. Especially when facing large-scale datasets, it can maintain high efficiency.

[0020] Other advantages, objectives, and features of the present invention will, to some extent, be elaborated in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. Brief Description of the Drawings

[0021] To make the objectives, technical solutions, and beneficial effects of the present invention clearer, the present invention provides the following drawings for illustration.

[0022] Figure 1 It is a system flow chart.

[0023] Figure 2 It is the main structure diagram of the TransUNet neural network model.

[0024] Figure 3 It is the structure diagram of the Transformer module.

[0025] Figure 4 It is the comparison 1 of the experimental results of different methods for test samples.

[0026] Figure 5 It is the comparison 2 of the experimental results of different methods for test samples.

[0027] Figure 6 It is the comparison 3 of the experimental results of different methods for test samples. Detailed Embodiments

[0028] The following further illustrates the present invention in conjunction with the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the examples given are not intended to limit the present invention.

[0029] Embodiment 1 As Figure 1 shown, the deep learning-based quantitative photoacoustic image method provided in this embodiment includes the following steps: S1. Photoacoustic data preparation, obtaining simulation data and experimental data to ensure that the photoacoustic data includes photoacoustic image data of different wavelengths; S2. Division of the model training set and test set, divided according to the dataset ratio (such as 80% training, 10% validation, 10% test); S3. Construct a neural network model, the neural network model adopts the TransUNet model, and a Transformer module is added to the U-Net network structure, and the Transformer module captures long-range dependencies; S4. Model training, using the training set to train the TransUNet model; optimizing model hyperparameters (such as learning rate, batch size, etc.); S5. Model validation and evaluation, using the test set to validate the model, and calculating performance metrics (such as error, correlation, CNR, etc.); S6. Complete image processing through the trained TransUNet model, that is, the model application and the actual image processing process, input the photoacoustic image for absorption coefficient estimation, and output the processed high-quality image or absorption coefficient distribution map; The TransUNet model in this embodiment can effectively process multi-wavelength photoacoustic image data and achieve high-precision blood oxygen saturation and absorption coefficient reconstruction in complex biological tissues.

[0030] As Figure 2 shown, the structure of the TransUNet model provided in this embodiment includes an input module, an encoder module, a Transformer module, a decoder module, and an output module; the specific structures of each part are as follows: The input module (Input) is used to input photoacoustic image data, usually multi-channel data, and each channel corresponds to a different wavelength; it is used to input photoacoustic image data of different wavelengths in the form of different channels; In this embodiment, different channels correspond to different wavelengths, and the model can make full use of the unique information of each wavelength in the tissue, improving the quantitative analysis accuracy of key indicators such as absorption coefficient and blood oxygen saturation. Specifically: 1. Capturing different optical properties: The absorption properties of light of different wavelengths in biological tissues are different, and the optical absorption coefficient (μa) shows obvious differences at different wavelengths. Using multi-channel input of multiple wavelengths can more comprehensively reflect the optical properties of tissues, helping to improve the quantitative estimation accuracy of blood oxygen saturation and absorption coefficient. In addition, each wavelength can provide different tissue contrast and detail information. By combining the information of different wavelengths through multiple channels, the model can more accurately extract multi-level features of tissues, thereby enhancing the imaging quality and accuracy.

[0031] 2. Improve the estimation accuracy of blood oxygen saturation: Lights of different wavelengths have different absorption characteristics on oxyhemoglobin (HbO 2 ), and deoxyhemoglobin (HbR), which enables data of different wavelengths to provide more information to distinguish between the two. Through the input of multi-wavelength photoacoustic images, the model can more accurately distinguish the oxygenation state in the blood, thereby improving the estimation accuracy of blood oxygen saturation.

[0032] 3. Improve the recovery of depth and tissue structure: Since photoacoustic imaging involves the processes of light propagation and absorption, the imaging accuracy of deep tissues is usually affected. Lights of different wavelengths have different penetration abilities for deep tissues. By using multi-wavelength data, the reflections of different wavelengths on deep structures can be integrated to improve the reconstruction ability of deep tissues.

[0033] The encoder module includes multiple cascaded convolutional blocks (Conv Block) for extracting local features of the input data; The convolutional block includes multiple convolutional layers, batch normalization layers, activation functions (such as ReLU), and max pooling layers; each convolutional block gradually downsamples and extracts higher-level feature representations; for example, the input is mapped from 21 channels to 64 channels through the first convolutional block, and the number of channels is successively doubled (128, 256) through subsequent convolutional blocks, and the final output feature map has a shape of (H / 8, W / 8, 256); The Transformer module is used to capture global dependencies and the correlation between different wavelengths; In this embodiment, the Transformer module can effectively capture global dependencies and the correlation between different wavelengths, mainly due to its two major features: self-attention mechanism and multi-head attention mechanism. Specifically: 1. The core idea of the self-attention mechanism is that the network can autonomously determine the relevance with the features of other positions based on the features of each position in the input. This mechanism is not limited to local regions but can perform global modeling on the entire input data to capture the dependencies between different features. In image processing tasks, each pixel (or each spatial position) may have important relationships with other regions far away in the image. For example, in blood oxygen saturation estimation, the morphology and structure of blood vessels may span a large area. Traditional convolutional neural networks (CNNs) usually rely on local convolutional kernels for processing and are difficult to effectively capture long-distance context information. The self-attention mechanism can capture global dependencies by calculating the correlation between each position and all other positions, fully understanding the global structure of the image; in multi-wavelength photoacoustic imaging, the information of each wavelength is different, so the mutual correlation between wavelengths is crucial. The self-attention mechanism allows the model to calculate the similarity and correlation between different wavelengths, enabling the network to understand globally how different wavelengths cooperate on the optical properties of tissues (such as absorption coefficient and blood oxygen saturation). This helps the model capture the connections between wavelengths more accurately rather than just processing each wavelength independently.

[0034] 2. The multi-head attention mechanism makes the model able to independently learn different relationships between input features in multiple subspaces by parallelizing the attention mechanism, thus enhancing the model's expressive power. In the processing of multi-wavelength data, using multi-head attention allows the model to learn the dependencies between different wavelengths from multiple perspectives. For example, certain wavelengths may play a dominant role in absorption coefficient estimation, while other wavelengths provide supplementary information. The multi-head attention mechanism allows the model to parallelly learn these different associations between wavelengths through multiple different heads (subspaces), enabling the model to capture the contributions of different wavelengths to the final result from more dimensions. Multi-head attention enables the model to perform feature fusion on the outputs of different heads, thus enhancing the model's expressiveness. The features learned by each head can be weighted and combined, and this flexible feature fusion can effectively capture the complex dependencies between different wavelengths, improving the accuracy of the final quantitative photoacoustic imaging.

[0035] As Figure 3 shown, the Transformer module includes Transformer*12 (Multi-head + FFN); that is, it includes 12 Transformer layers, and each Transformer layer contains a multi-head attention mechanism module (Multi-headAttention) and a feed-forward neural network (Feed-Forward Network, FFN); The multi-head attention mechanism module is used to calculate the feature dependencies and global associations between wavelengths; In this embodiment, the process of calculating the characteristic dependencies between wavelengths mainly relies on the application of the self-attention mechanism in the Transformer module. This process can help the model capture the mutual influence and global dependencies between data of different wavelengths, thereby improving the accuracy in quantitative photoacoustic imaging. The specific steps are as follows: Step 1: Calculate the query, key, and value. For each input feature (such as each pixel in a multi-wavelength image or the data of each wavelength), three vectors are first generated: the query vector (Query), the key vector (Key), and the value vector (Value). These vectors are obtained by multiplying with learnable weight matrices. The following is the calculation of the query, key, and value: Given the input feature , we obtain the query, key, and value through linear transformation: is the learnable weight matrix, , and represent the query, key, and value vectors.

[0036] Step 2: Calculate the attention weights: The attention weights are obtained by calculating the similarity between the query and the key. This similarity is usually measured by the dot product: where, is the dot product of the query vector and the key vector, which measures the correlation between the query and the features at other positions; is the square root of the dimension of the key vector, which is used to prevent the dot product result from being too large and causing the gradient to vanish.

[0037] The softmax operation ensures that the weights are normalized over all positions, thereby obtaining the relative importance of each position.

[0038] Finally, the weights are applied to the value vector to obtain the weighted output.

[0039] Step 3: Cross-wavelength correlation: For multi-wavelength data, each wavelength corresponds to an input channel, and the self-attention mechanism can calculate the correlation between different wavelengths. For example, assuming the input data contains photoacoustic images of wavelengths, the attention mechanism can not only calculate the correlation within each wavelength but also calculate the dependencies between different wavelengths. This enables the model to globally understand how different wavelengths cooperate to affect the final estimation of the absorption coefficient and blood oxygen saturation.

[0040] The feedforward neural network is used to further process and transform features as follows: 1. Feature processing and transformation of input data: Photoacoustic imaging usually uses lasers of multiple wavelengths to obtain multi-channel data. In this case, the photoacoustic images of each wavelength represent different absorption characteristics of the photoacoustic signals. To enable the TransUNet model to effectively process this multi-wavelength data, the images of each wavelength need to be first converted into different channels. Assume the input data contains data of wavelengths, and the shape of the input photoacoustic image data is , where and are the height and width of the image respectively, and is the number of wavelengths (e.g., 21 wavelengths with a step of 10 nm between 700 nm and 900 nm). Each wavelength corresponds to a channel. Therefore, the entire dataset will be represented as a three-dimensional tensor with a shape of

[0041] . The meaning of the channels: Each channel corresponds to photoacoustic data of different wavelengths. For example, the image of wavelength 700 nm is in the first channel, the image of wavelength 710 nm is in the second channel, and so on. Normalization formula: where is the pixel value at position in the th wavelength channel, and and are the minimum and maximum values of the th wavelength channel respectively. Through such normalization processing, it is ensured that the data of all wavelengths are within the range of , which is convenient for model training.

[0042] Converted into a format suitable for the input of TransUNet. The input data processed by the TransUNet model is usually a two-dimensional sequence, while the image data is three-dimensional. To adapt to the structure of the TransUNet model, the multi-wavelength image data needs to be converted into a suitable format for processing, that is, feature flattening.

[0043] Spatial flattening: First, flatten the spatial dimensions of the image into one dimension. When processing multi-wavelength images, the input data will be converted from to , that is, the multiple wavelength information at each position of the image is used as a feature vector. In this way, the feature of each pixel point becomes a feature vector composed of A vector composed of wavelengths.

[0044] Output shape: The shape of the transformed data is , where , representing the number of pixels in the image, represents the feature dimension of each pixel (i.e., the number of wavelengths). This flattened input can be directly fed into TransUNet.

[0045] Feature processing in the Transformer module: Since the Transformer itself is a sequence-based model, it cannot directly capture the spatial location information in the input data. Therefore, when converting the image data into a one-dimensional sequence, positional encoding needs to be added to each position (i.e., each pixel point) to retain the spatial information. Positional encoding embeds each position to ensure that the model can recognize the spatial location. For each pixel in the image, the positional encoding can be generated using sine and cosine functions: , where 𝑖 is the position index, 𝑘 is the dimension index, and d is the embedding dimension. In this way, the positional encoding can provide unique spatial information for each pixel.

[0046] Multi-Head Self-Attention In the Transformer module, the multi-head self-attention mechanism is the core, which can help the model capture the global dependencies between the input features, including the dependencies between wavelengths. For each input pixel position (or each wavelength), the Transformer calculates the similarity between this position and other positions (by calculating the similarity between the query and key vectors through dot product), and then determines how each pixel focuses on the information of other pixels. In particular, the model can capture the dependencies between different wavelengths. For example, some wavelengths may contribute more to the estimation of the absorption coefficient, while other wavelengths provide supplementary information. The self-attention mechanism enables the network to calculate these dependencies globally. The formula is as follows: where are the query, key, and value vectors respectively, is the dimension of the key vector, ensures that the relevance of each position is normalized and calculates the final weighted value.

[0047] In photoacoustic imaging, the dependence between different wavelengths is crucial, especially in the estimation of absorption coefficient and blood oxygen saturation. Through the self-attention mechanism, the model can calculate the similarity between different wavelengths and automatically learn which wavelengths have higher correlation with the final quantitative results, thus enhancing the accuracy of quantitative imaging. Output feature processing: After being processed by the Transformer module, the features of the image are updated and passed to the decoder part. The decoder gradually restores the spatial resolution of the image and finally generates the distribution maps of absorption coefficient and blood oxygen saturation. The decoder uses skip connections corresponding to the encoder to combine low-level features with high-level features and restore more details.

[0048] Finally, the convolutional layer is used to map the feature map to the target output dimension (absorption coefficient) to complete the quantitative reconstruction.

[0049] Decoder module, used to gradually restore the resolution of the feature map and fuse low-level and high-level features; including several concatenated up-sampling blocks (Up-sample Block); used to implement multiple up-sampling operations (such as deconvolution layer or transposed convolution layer), and each up-sampling block is followed by a skip connection (Skip Connection) to perform feature fusion with the corresponding layer of the encoder, and finally obtain a two-dimensional distribution with dimensions (H, W, 2), representing the absorption coefficient and blood oxygen saturation (sO2).

[0050] Output module (Output), used to output the reconstructed quantitative photoacoustic image, including the distribution of absorption coefficient and blood oxygen saturation.

[0051] The TransUNet model provided in this embodiment combines the advantages of U-Net and Transformer to achieve efficient processing of multi-wavelength data. The encoder part extracts local features through convolutional blocks, the Transformer part captures global information through the multi-head attention mechanism, and the decoder part restores the original image size and fuses features through up-sampling blocks, and finally outputs a high-precision quantitative photoacoustic image.

[0052] The TransUNet model structure provided in this embodiment combines the advantages of U-Net and Transformer, and fuses local features and global dependencies: the convolutional operation in the encoder part effectively extracts the local features of the input data, and the self-attention mechanism is introduced in the Transformer part to capture the feature dependencies and global context information between different wavelengths. This combination enables the model to not only process local details but also understand the global information of the entire image, thus improving the accuracy of photoacoustic image reconstruction.

[0053] Meanwhile, considering the characteristics of multi-wavelength photoacoustic data, the TransUNet model utilizes its multi-channel input ability and the multi-head attention mechanism of Transformer to more precisely model the interrelationships between different wavelengths. This helps improve the estimation accuracy of physiological parameters such as blood oxygen saturation, especially in deep tissues.

[0054] The input data preprocessing in this embodiment includes the following methods: Data format: Loaded as multi-channel data with the shape of (H, W, C), where H and W are the spatial dimensions of the image, and C represents the number of channels for different wavelengths; Normalization: The data of each wavelength is normalized to the interval [0, 1] respectively to ensure the consistency of the dynamic range of all wavelength data.

[0055] Data classification: 80% is used for training, and 20% is used for testing (10% of the training set is randomly divided as the validation set).

[0056] The encoder module of this embodiment: Feature extraction convolution: It contains 3 convolution blocks, and each convolution block consists of a convolutional layer, batch normalization, ReLU activation function, and max pooling layer. The first convolution block maps the input from 21 channels to 64 channels, and the number of channels in subsequent convolution blocks doubles successively (128, 256). Finally, the output feature map has the shape of (H / 8, W / 8, 256). Feature vectorization: Flatten the two-dimensional feature map into a one-dimensional vector with the shape of (N, P, D), where N is the batch size, P is the number of feature maps, and D is the feature dimension.

[0057] Transformer module: Encoding position: Add sine-cosine position encoding to retain spatial information. Multi-head attention mechanism: Set parameters including the number of attention heads = 8, the number of Transformer layers = 12, and the hidden layer size = 256. Calculate the feature dependencies and global associations between wavelengths through this mechanism. Feed-forward neural network (FFN): Each Transformer layer is followed by two fully connected networks. The first layer expands the feature dimension from D to 4D, and the second layer compresses it back to D.

[0058] Decoder module: Upsampling and fusion: Gradually perform upsampling operations on the corresponding features of the encoder until the original image size is restored. Each upsampling stage includes a transposed convolutional layer, skip connection (concatenate with the feature map of the corresponding layer in the encoder), and convolution block (with the same structure as the encoder).

[0059] Output layer: Use a separate convolutional layer to map the high-dimensional features to the target output dimensions (absorption coefficient and sO2), and the final output has the shape of (H, W, 2).

[0060] Loss function design in this embodiment: Multi-task loss: Define the combined loss function L, which combines the mean squared error and other weight coefficients to evaluate the difference between the predicted value and the true value. Wavelength consistency: Introduce a wavelength regularization term to ensure the global consistency of multi-wavelength data.

[0061] Training and optimization strategy: Use the Adam optimizer with an initial learning rate of 10 -4 , and adopt a learning rate decay strategy; set the batch size to 16. Data augmentation techniques include random cropping, rotation, horizontal flipping, and adding random noise to the wavelength data to improve the robustness of the model.

[0062] The Transformer module provided in this embodiment improves the accuracy and efficiency of photoacoustic image reconstruction by combining the advantages of U-Net and Transformer, and performs particularly well in processing deep tissues and large-scale data.

[0063] Embodiment 2 The photoacoustic image quantitative analysis method based on the TransUNet model provided in this embodiment has high quantitative accuracy and image reconstruction quality, and can be widely applied in the field of medical imaging to provide effective support for clinical diagnosis and treatment; first, use the photoacoustic image training dataset to train and obtain the neural network of the TransUNet model; then use the trained TransUNet model to reconstruct the quantitative photoacoustic image, which is specifically carried out according to the following steps: Step 1: Input data preprocessing: Preparation of data form: Load the obtained simulation data and experimental data as multi-channel data of 21 wavelengths, with the shape of (H, W, 21), where H and W are the spatial dimensions of the image, and 21 represents the number of wavelength channels; Experimental data: Experimental data usually comes from actual photoacoustic imaging experiments, using real photoacoustic imaging equipment to image biological tissues. The advantage of experimental data is that it reflects the characteristics of real-world biological tissues, including tissue inhomogeneity, noise, and measurement errors. Its advantage lies in directly reflecting the optical absorption and scattering characteristics of real tissues, and can provide a true reflection of actual pathological conditions (such as tumors, blood vessels, etc.); moreover, training and evaluation using experimental data can help the model better apply in a clinical environment and improve the practicality of the model. However, experimental data often contains noise and measurement errors, and in some cases, data collection may be incomplete or have limitations (such as limited view or sparse sampling); in addition, experimental data is usually expensive and difficult to obtain, especially labeled data, which limits the training and verification of the model.

[0064] Simulation data: Simulation data is generated through physical models or simulation platforms, simulating processes such as light wave propagation, tissue absorption, and scattering in photoacoustic imaging. Its advantages are that simulation data can generate a large number of diverse samples through models, which is suitable for the training of deep learning models. Especially when experimental data is scarce, simulation data can make up for this deficiency; simulation data allows researchers to precisely control factors such as the optical properties of tissues, scattering effects, and noise, so that training and testing can be carried out under different experimental conditions; in addition, the labels of simulation data (such as absorption coefficient and blood oxygen saturation) can be accurately generated through physical models, avoiding the problem of human annotation errors during the experimental process. However, there is often a domain gap between simulation data and actual experimental data, that is, the simulation model may not be able to fully capture the complexity and variability of actual biological tissues. This gap may lead to the model's performance on real data being inferior to that on simulation data. By inputting the above simulation data and experimental data in different proportions, the advantages of both can be brought into play and their respective deficiencies can be made up for.

[0065] Enhance the diversity of training data: Simulation data can provide large-scale, accurately labeled, and diverse training samples, while experimental data reflects the complexity of the real world. Combining the two can enable the model to learn more diverse features and variation situations, improving the generalization ability.

[0066] Reduce the influence of experimental noise: By combining simulation data, the model can be trained on cleaner and clearly labeled data, thereby reducing the interference of noise in experimental data on training and improving the robustness of the model.

[0067] Improve the generalization ability of the model: Using simulation data for pre-training and experimental data for fine-tuning can enable the model to learn the general laws obtained from simulation data while retaining the specific characteristics of experimental data, thereby enhancing the model's performance and robustness in real scenarios.

[0068] Optimize the accuracy of quantitative imaging: Combining these two types of data can effectively improve the accuracy of quantitative photoacoustic imaging tasks, especially the estimation of blood oxygen saturation and absorption coefficient. Simulation data can help the model identify the laws under different tissues and conditions, while experimental data enables the model to better adapt to the complex situations in actual applications.

[0069] Normalization processing: Use the following formula to normalize the data of each wavelength to [0, 1] respectively to eliminate the influence of the data scale of different wavelengths and ensure the consistency of the dynamic range of all wavelength data: Among them, in this formula, the meanings of each letter are as follows: : Represents the value of the original data, i.e., a specific value in the input data; : Represents the minimum value in the original dataset xxx; : Represents the maximum value in the original dataset xxx; Represents the value of the normalized data, i.e., the result of normalizing the original data through this formula.

[0070] Data classification: All data is divided into an 80% training set and a 20% test set (randomly divide 10% of the training set as the validation set) Set model parameters: Input data processing: The input data includes multi-wavelength photoacoustic images in the form of , where: and are the spatial dimensions of the image; is the number of wavelengths, which are 21 wavelengths from 700 nm to 900 nm with a step of 10 nm. Features are extracted from the data channels of each wavelength through convolution operations, and the image is flattened and fed into the TransUNet module for further processing. The input data needs to be normalized to have the same scale range.

[0071] Step 2. Encoder The encoder uses multiple convolutional blocks to gradually extract the features of the input image. Each convolutional block includes the following steps: Convolutional layer: Use convolutional kernels for feature extraction, and the number of convolutional kernels increases layer by layer (e.g., 64 -> 128 -> 256).

[0072] Among them, is the input feature, is the convolutional kernel, is the bias.

[0073] Batch Normalization: Normalize the convolutional output to accelerate training and reduce overfitting.

[0074] Activation function (ReLU): The ReLU activation function ensures the non-linear expression of the network: MaxPooling: Use pooling layer for downsampling, reducing the spatial dimension and retaining the most significant features.

[0075] Feature flattening: After the convolutional layer, the feature map is flattened into a one-dimensional vector to fit the format of the Transformer input: , where is the number of pixels in the image, is the feature dimension of each pixel.

[0076] Step 3: Transformer module The self-attention mechanism is the core of the Transformer module. It determines the degree of attention of each feature to other features by calculating the similarity between the query, key, and value of the input features. The specific steps are as follows: Calculate the query, key, and value vectors: Among them, is the input feature matrix, is the learnable weight matrix, , , are the query, key, and value vectors respectively.

[0077] Calculate the attention weights: The similarity between the query and the key is calculated by the dot product and passed through the softmax operation to obtain the attention weights: Among them, is the dimension of the key vector, is the dot product of the query and the key, measuring the similarity between features. Through the operation, the weights are normalized to ensure that the sum of the attention weights at all positions is 1.

[0078] Capture of wavelength-dependent relationships: The Transformer module can capture the interaction between different wavelengths by calculating the similarity (dot product calculation) of different wavelengths, thereby modeling the wavelength-dependent relationships. The self-attention mechanism can handle multi-wavelength image data, enabling the model to integrate information from each wavelength and improve the quantitative estimation ability of the absorption coefficient and blood oxygen saturation.

[0079] The multi-head attention mechanism is used for parallel computing to enhance the model's ability to model the relationships between different wavelength features. Each attention head learns different relationships and patterns, enhancing the model's understanding ability of the data.

[0080] Among them, is the th attention head, is the output weight matrix.

[0081] Each layer in the Transformer module is followed by a feed-forward neural network for further processing of features. The feed-forward neural network usually consists of two fully-connected layers: where and are learnable weight matrices, and are bias terms, and the activation function is used to introduce non-linearity.

[0082] Step 3: Decoder part The goal of the decoder in this embodiment is to gradually restore the spatial resolution of the image. At each layer, the decoder first performs transposed convolution, i.e., upsampling through a convolutional kernel and convolution with a stride of 2. Then, through skip connections, the features in the encoder are concatenated with the feature maps in the decoder to combine low-level and high-level features.

[0083] Through gradual upsampling, the original image size is finally restored.

[0084] Step 4: Output layer The last step of the decoder is to map the restored feature map through a convolutional layer to the final target output dimension (such as the accurate estimation of the absorption coefficient and .

[0085] where represents the output, i.e., the output absorption coefficient and the two-dimensional distribution of blood oxygen saturation , and the number of output channels is 2.

[0086] In this embodiment, the mean squared error (MSE) is used as the loss function for model training to optimize the estimation accuracy of the absorption coefficient and .

[0087] where is the predicted value, is the true value, and N is the number of samples.

[0088] The encoder part of this embodiment: ① Feature extraction convolution: Use 3 convolutional blocks to sequentially extract the features of multi-wavelength data. Each convolutional block includes: a convolutional layer (kernel = 3x3, stride = 1, padding = 1), batch normalization, an activation function (ReLU), and a max pooling layer (MaxPooling, kernel = 2x2, stride = 2). The first convolutional block maps the 21-channel input to 64 channels, and the number of channels doubles successively thereafter (128, 256). The output feature shape is (H / 8, W / 8, 256). ② Feature vectorization: Flatten the two-dimensional feature map into a one-dimensional vector with a shape of (N, P, D), where: N = batchsize; p = H / 8 * W / 8; the number of patches of the image; D = 256: the feature dimension; The Transformer part of this embodiment: ① Encoding position: Add position encoding to retain the spatial information of the image, using the sine-cosine position encoding method; ② Multi-head attention mechanism: Set the core parameters of the Transformer: the number of attention heads (num_heads) = 8; the number of Transformer layers (num_layers) = 12; the hidden layer size (hidden_size) = 256; calculate the feature dependencies and global associations between wavelengths through the multi-head attention mechanism; The feed-forward neural network (FFN) of this embodiment: A feed-forward neural network is connected after each Transformer layer, using a 2-layer fully connected network: where the dimensions of W1 and W2 are D * 4D and 4D * D in sequence; and represents the bias term; represents the input; The decoder part of this embodiment: ① Upsampling and fusion: The decoder structure gradually upsamples the encoded features until the size of the original image is restored; ② Each upsampling stage includes: a transposed convolution layer (Transposed Convolution, kernel = 2x2, stride = 2), a skip connection: concatenating the feature maps of the corresponding layer of the encoder; a convolutional block (the same structure as the encoder); the final output dimension (H, W, 2), representing the two-dimensional distributions of the absorption coefficient and sO2; ③ Output layer: Use a separate convolutional layer to map the number of channels from the high-dimensional feature to the target output dimension (absorption coefficient and sO2); Loss function design of this embodiment (1). Multi-task loss: Define the joint loss function L: where Lmse is the mean squared error, used to evaluate the difference between the predicted value and the true value, α and β are weight coefficients with an initial value of 1; absorption represents the absorption coefficient; represents the blood oxygen saturation; (2) Wavelength consistency: Introduce a wavelength regularization term to constrain the global consistency of multi-arm length data; Training and optimization of this embodiment: (1). Parameter settings: ① Optimizer: Adam, learning rate is 10 -4 , and use the learning rate decay strategy; batchsize = 16; number of training epochs; ② Data augmentation: Add the following augmentations during training: random cropping, rotation, horizontal flipping, and spectral interference (add random noise to the wavelength data to improve robustness); Verification and testing of the model of this embodiment: (1). Testing: Compare the model reconstruction with the true situation distribution in the test set; (2). Verification: Use an independent validation set to evaluate the performance of the model in reconstructing the sO2 and absorption coefficient distributions; Embodiment 3 This embodiment demonstrates three representative cases for illustration, which are illustrated by the line profile of the inclusion at a wavelength of 800 nm, as follows: First, check the performance of the estimated image of µa (absorption coefficient) in all test samples; As Figure 4 shown, in this embodiment, the GT-φ (Monte Carlo optical flux correction-based) method may overestimate µa in the deep region of the test object, while the DL-Exp method can reconstruct a smoother and more uniform background; Figure 4 The meanings are described as follows: Figure 4 A in : The image of the optical absorption coefficient (µa) estimated by the linear calibration method (Cal.), and the sample line is shown in yellow. The estimated µa of this method in the background material and inclusions is generally lower than the reference value; Figure 4B in: The μa image estimated by the Monte Carlo fluence correction method (GT-φ), with the sample line shown in green. This method has an overestimation bias in the μa estimation of the deep region and distortion appears at the edges of the inclusions; Figure 4 C in: The μa image estimated by the deep learning model trained on simulated data (DL-Sim), with the sample line shown in blue. This model overestimates μa in the top region of the sample and generates artifacts in the coupling medium region (outside the sample); Figure 4 D in: The μa image estimated by the deep learning model trained on experimental data (DL-Exp), with the sample line shown in red. This model generates clear demarcations at the boundaries of the sample and the edges of the inclusions, providing a smoother and more accurate μa estimation spatially; Figure 4 E in: Shows the comparison of the sample line profiles of the above four methods (Cal., GT-φ, DL-Sim, DL-Exp), with the colors consistent with Figures A to D, for showing the estimated values of different methods at different depths.

[0089] As Figure 5 shown, in this embodiment, a challenging test sample is adopted, where there are halo artifacts around a certain inclusion. In this test, DL-Exp can reconstruct a relatively smooth and uniform background; Figure 5 The meanings are described as follows: Figure 5 A in: The μa image estimated by the linear calibration (Cal.) method of the second test sample, with the sample line shown in yellow. This sample is more complex than the sample in Figure A and contains more halo artifacts; Figure 5 B in: The μa image estimated by the Monte Carlo fluence correction (GT-φ) of the second test sample, with the sample line shown in green. This method produces overestimated μa estimations at the edges of the inclusions and in the background region, resulting in distortion; Figure 5 C in: The μa image estimated by the deep learning model trained on simulated data (DL-Sim) of the second test sample, with the sample line shown in blue. This method overestimates μa in the background region and there are artifacts, especially in the coupling medium region; Figure 5 D in: The μa image estimated by the deep learning model trained on experimental data (DL-Exp) of the second test sample, with the sample line shown in red. This method can provide a smooth background estimation and the estimations for most inclusions are close to the reference values; Figure 5E in: The line profile of the second test sample, showing the µa estimations in the inclusion region by different methods. The red (DL-Exp) is close to the reference value and performs the best, while the blue (DL-Sim) and green (GT-φ) have larger deviations.

[0090] As Figure 6 shown, this embodiment uses a test sample outside the training data distribution range, and the geometry of its inclusions has not appeared in the training set. In this case, the µa estimation results of the DL-Sim and DL-Exp methods show a highly uneven distribution inside the inclusions, but DL-Exp is slightly better than DL-Sim.

[0091] Figure 6 The meanings are as described below: Figure 6 A in: The µa image estimated by the linear calibration (Cal.) of the third test sample, and the sample line is shown in yellow. The geometry of this sample is different from the training data, and the distribution of inclusions is more complex; Figure 6 B in: The µa image estimated by the Monte Carlo photon flux correction (GT-φ) of the third test sample, and the sample line is shown in green. This method overestimates µa in the deep region, and the boundaries of the inclusions are distorted; Figure 6 C in: The µa image estimated by the deep learning model trained on simulated data (DL-Sim) of the third test sample, and the sample line is shown in blue. Due to the geometry not being in the training data, this method produces uneven µa estimations inside the inclusions and generates artifacts in some regions; Figure 6 D in: The µa image estimated by the deep learning model trained on experimental data (DL-Exp) of the third test sample, and the sample line is shown in red. This method provides the smoothest and most accurate estimation, being closer to the reference value; Figure 6 E in: The line profile of the third test sample, showing the µa estimations in the inclusion region by different methods. The red (DL-Exp) performs the best among all methods, while the µa estimated by the blue (DL-Sim) and green (GT-φ) is more uneven with larger errors.

[0092] Overall, the control methods are as follows: The Cal method (linear calibration) systematically underestimates the µa values in both the inclusion and background regions; The GT-φ method (Monte Carlo photon flux correction) overestimates µa in the large inclusions and the central region of the background material, resulting in distorted inclusion edges; The DL-Sim method (a deep learning model trained based on simulated data) systematically overestimates µa in the top region of the sample; The DL-Exp method (a deep learning model trained based on experimental data) can generate a smoother µa estimate and provide a clearer demarcation at the sample edge and inclusion boundary.

[0093] From the quantitative analysis results, the deep learning method has significantly better µa estimation accuracy in the inclusion region than other methods, and these results are consistent in all test samples.

[0094] In addition, it shows that the DL-Exp method is also superior to DL-Sim in terms of spatial accuracy and quantitative estimation, highlighting the advantages of deep learning models trained on experimental data.

[0095] The above-described embodiments are only preferred embodiments given to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are within the protection scope of the present invention. The protection scope of the present invention is subject to the claims.

Claims

1. A quantitative photoacoustic imaging method based on deep learning, characterized by: The following steps are involved: S1. Photoacoustic data preparation, obtaining simulation data and experimental data, wherein the photoacoustic data includes photoacoustic image data of different wavelengths; S2. Model training set and test set division, divide the data set into training data set, verification data set, and test data set; S3. Construct a neural network model, wherein the neural network model adopts a TransUNet model, adds a Transformer module to the U-Net network structure, and captures the long-distance dependency relationship of data in the data set through the Transformer module; S4. Model training: use the training set to train the TransUNet model; optimize the model hyperparameters; S5. Model validation and evaluation: Use the test set to validate the model and calculate the model performance indicators; S6. Complete image processing through the trained TransUNet model, input the photoacoustic image to estimate the absorption coefficient, and output the processed image or absorption coefficient distribution map.

2. The quantitative photoacoustic imaging method based on deep learning according to claim 1, characterized in that: The TransUNet model includes an input module, an encoder module, a Transformer module, a decoder module and an output module; The multi-channel input module is used to input photoacoustic image data of different wavelengths in the form of different channels; The encoder module includes a plurality of convolution blocks connected in series, which are used to extract local features of input data; the convolution blocks include a plurality of convolution layers, a batch normalization layer, an activation function and a maximum pooling layer; each convolution block is gradually downsampled and features are extracted; The Transformer module is used to capture the global dependencies and the correlations between different wavelengths; The decoder module is used to gradually restore the resolution of the feature map and fuse low-level and high-level features; A two-dimensional distribution diagram representing the absorption coefficient and blood oxygen saturation is obtained; The output module is used to output the reconstructed quantitative photoacoustic image, including the distribution diagram of the absorption coefficient and the blood oxygen saturation.

3. The quantitative photoacoustic imaging method based on deep learning according to claim 2, characterized in that: The Transformer module includes several Transformer layers, each of which includes a multi-head attention mechanism module and a feedforward neural network. The multi-head attention mechanism module is used to calculate the feature dependencies and global associations between wavelengths; the feedforward neural network is used to further process and transform features.

4. The method for quantitative photoacoustic imaging based on deep learning according to claim 2, characterized in that: The decoder module includes several serially connected upsampling blocks, which are used to implement multiple upsampling operations. Each upsampling block is followed by a jump connection, and features are fused with the corresponding layer of the encoder to obtain a two-dimensional distribution map representing the absorption coefficient and blood oxygen saturation.

5. The quantitative photoacoustic imaging method based on deep learning according to claim 1, characterized in that: The photoacoustic data in step S1 is preprocessed in the following manner: S11: Data format: loaded as multi-channel data of shape (H, W, C), where H and W are the spatial dimensions of the image, and C represents the number of channels at different wavelengths; S12: Normalization: The data of each wavelength is normalized to the interval [0, 1].

6. The quantitative photoacoustic imaging method based on deep learning according to claim 2, characterized in that: The specific steps of the Transformer module are as follows: Step 1: Calculate the query, key, and value. For each input feature, generate three vectors: query vector, key vector, and value vector. Step 2: Calculate attention weight: Get the attention weight by calculating the similarity between the query and the key; Step 3: Cross-wavelength correlation: For multi-wavelength data, each wavelength corresponds to an input channel, and the correlation between different wavelengths is calculated through the self-attention mechanism.

7. The method for quantitative photoacoustic imaging based on deep learning according to claim 3, characterized in that: The multi-head attention mechanism module adopts the multi-head attention mechanism for parallel calculation, which is specifically performed according to the following formula: in, It is A head of attention, is the output weight matrix.

8. The method for quantitative photoacoustic imaging based on deep learning according to claim 2, characterized in that: The decoder module outputs the absorption coefficient and : in, Represents the output, that is, the output absorption coefficient and blood oxygen saturation Two-dimensional distribution of .

9. The quantitative photoacoustic imaging method based on deep learning according to claim 1, characterized in that: The loss function in the TransUNet model is calculated according to the following formula: Where L is the joint loss function, Lmse is the mean square error, which is used to evaluate the difference between the predicted value and the true value, α and β are weight coefficients; absorption represents the absorption coefficient; Indicates blood oxygen saturation.

10. A quantitative photoacoustic imaging system based on deep learning, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Quantitative photoacoustic imaging method on basis of deep neural networks

    CN108309251A

  • Laser energy correction method and prompting method in photoacoustic imaging system, and photoacoustic imaging system

    CN116269203A

  • Photoacoustic image restoration method and system based on Transform neural network

    CN118115374A

  • Machine Learning-Based Quantitative Photoacoustic Tomography (PAT)

    US20190192008A1

Cited By

  • Photoacoustic chromatography absorption coefficient reconstruction method and device based on implicit neural representation

    CN120782906A

  • Arrhythmia real-time detection method, system and device based on morphological fidelity consistency constraint

    CN121101514A

  • CTD section data interpolation reconstruction method based on deep learning

    CN122176095A