Photoacoustic image enhancement method and device combining Mama and CNN (Convolutional Neural Network)

By combining the Mamba U-network and the CNN hybrid two-branch model, the problems of fuzzy organ boundary segmentation and high computational overhead in photoacoustic imaging technology are solved, achieving high-precision organ segmentation with low computational overhead and reducing labeling costs.

CN120707580AActive Publication Date: 2025-09-26THE FIRST MEDICAL CENT CHINESE PLA GENERAL HOSPITAL

Patent Information

Application Number
CN202511197074.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-09-26
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

Existing photoacoustic imaging technology has problems in organ segmentation, such as fuzzy organ boundary segmentation, low segmentation accuracy, high computational overhead and low computational efficiency, which are particularly evident in high-noise photoacoustic images. In addition, the labeling cost is high and the weak supervision strategy lacks adaptability.

Method used

A hybrid two-branch model combining Mamba U-type network and CNN is adopted. Through pre-training and fine-tuning, CNN is used to obtain local features and Mamba U-type network to obtain global context features. The fusion module weightedly fuses the organ category probability maps output by the two to achieve high-precision organ segmentation.

Benefits of technology

It achieves high-precision organ segmentation, reduces computational overhead, reduces annotation costs, and improves the segmentation accuracy of organ edge areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707580A_ABST
    Figure CN120707580A_ABST
Patent Text Reader

Abstract

The invention provides a photoacoustic image enhancement method and device combining Mama and CNN, and relates to the technical field of image processing. The opto-acoustic image enhancement method combining the Mama and the CNN is realized through a hybrid double-support model, and the hybrid double-support model comprises the CNN, the Mama U-shaped network and a fusion module; the method can comprise the following steps: acquiring a first pixel-level organ category probability graph of a photoacoustic image by using a CNN (Convolutional Neural Network); acquiring a second pixel-level organ category probability graph of the photoacoustic image by using a Mamba U-shaped network; the first pixel-level organ category probability graph and the second pixel-level organ category probability graph are subjected to weighted fusion through a fusion module to obtain a third pixel-level organ category probability graph, and the third pixel-level organ category probability graph comprises the probability value of each pixel point belonging to various organs in the photoacoustic image. The method and the device have a high-precision organ segmentation function and low calculation overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to a photoacoustic image enhancement method and device combining a structured spatial model (Mamba) with a convolutional neural network (CNN). Background Art

[0002] Photoacoustic imaging (PAI) is a hybrid medical imaging modality based on the photoacoustic effect. It uses pulsed lasers to excite biological tissues, generating ultrasonic signals and reconstructing images. This technology combines the high contrast of optical imaging with the deep tissue penetration of ultrasound imaging, making it valuable in areas such as brain functional imaging and early tumor detection.

[0003] Currently, organ segmentation in photoacoustic imaging is primarily achieved using models such as CNNs and Vision Transformers (ViTs). However, due to the limitations of CNNs' local receptive fields, they struggle to model long-range spatial dependencies, resulting in blurred organ boundary segmentation. Segmentation accuracy is particularly degraded in noisy photoacoustic images. While ViTs can capture global context, their self-attention mechanism has a computational complexity that grows quadratically with the input image size (O(*n*²)), resulting in high computational overhead and low computational efficiency. This limits their applicability and makes them difficult to apply to high-resolution photoacoustic images. Therefore, a photoacoustic image enhancement solution that combines high-precision segmentation performance with low computational overhead is urgently needed. Summary of the Invention

[0004] In view of this, the present disclosure provides a photoacoustic image enhancement method and device combining Mamba and CNN.

[0005] According to a first aspect of the present disclosure, a method for photoacoustic image enhancement combining Mamba and CNN is provided. The method is implemented by a pre-trained hybrid two-branch model, wherein the hybrid two-branch model includes a CNN, a Mamba U-shaped network, and a fusion module. The method includes: Use CNN to obtain the first pixel-level organ category probability map of the photoacoustic image; Obtaining a second pixel-level organ category probability map of the photoacoustic image using a Mamba U-network; The first pixel-level organ category probability map and the second pixel-level organ category probability map are weightedly fused by the fusion module to obtain a third pixel-level organ category probability map, wherein the third pixel-level organ category probability map includes a probability value of each pixel point in the photoacoustic image belonging to each type of organ.

[0006] In some embodiments of the first aspect of the present disclosure, the Mamba U-type network adopts a symmetric encoder-decoder structure; and obtaining the second pixel-level organ category probability map of the photoacoustic image using the Mamba U-type network includes: performing encoder processing on the photoacoustic image to obtain contextual compression features, the encoder processing comprising: segmenting the photoacoustic image to obtain a sequence of photoacoustic image blocks, linearly transforming the sequence of photoacoustic image blocks to obtain embedded features, and sequentially processing the embedded features through multiple levels of encoding units to obtain contextual compression features, wherein each level of encoding unit processing comprises at least two levels of visual state space module processing in series; Based on the context compression features and the output features obtained by the encoding units at each level, decoder processing is performed to obtain the second pixel-level organ category probability map. The decoder processing includes: multi-level decoding unit processing, upsampling the sharpened reconstruction features obtained by the multi-level decoding unit processing and performing linear projection processing. The decoding unit processing at each level includes: upsampling the context compression features or the sharpened reconstruction features output by the previous level decoding unit to obtain upsampled features, fusing the upsampled features with the output features processed by the symmetrical hierarchical encoding units to obtain fused features, and performing two-level visual state space module processing in series on the fused features to obtain the sharpened reconstruction features of the current level decoding unit.

[0007] In some embodiments of the first aspect of the present disclosure, the visual state space module processing at each level in the two-level visual state space module processing in series includes: main path processing, the main path processing includes depth convolution operations, SS2D operator operations and SiLU activation function processing performed in sequence; gated branch path processing, the gated branch path processing includes fully connected layer processing for generating a gated signal; gated multiplication and linear projection of the feature map obtained by the main path processing and the feature map obtained by the gated branch path processing; and, output feature map and residual addition operation of the gated multiplication and linear projection.

[0008] In some embodiments of the first aspect of the present disclosure, the weighted fusion of the first pixel-level organ class probability map and the second pixel-level organ class probability map by the fusion module to obtain a third pixel-level organ class probability map includes: The fusion module performs weighted summation on the first pixel-level organ category probability map and the second pixel-level organ category probability map based on the following formula to obtain a third pixel-level organ category probability map;

[0009] Among them, P represents the third pixel-level organ category probability map, represents the first pixel-level organ category probability map, Represents the second pixel-level organ category probability map, and the value range of weight α is [0.7, 0.8].

[0010] In some embodiments of the first aspect of the present disclosure, the hybrid two-branch model is trained based on a synthetic dataset, and the synthetic dataset is constructed in the following manner: obtaining a pre-selected abdominal MRI image in a CHAOS dataset, wherein the abdominal MRI image is annotated with an organ segmentation mask; converting the abdominal MRI image into a simulated photoacoustic tomography (PAT) image, performing a skeletonization process of corrosion on the organ segmentation mask of the abdominal MRI image to generate graffiti labels along the central axis of the organ, and annotating the simulated PAT image with the graffiti labels to form a synthetic dataset.

[0011] In some embodiments of the first aspect of the present disclosure, the hybrid two-branch model is trained by: pre-training based on the synthetic data set to determine the parameters of the hybrid two-branch model; and fine-tuning the parameters of the hybrid two-branch model based on real data, wherein the real data includes mouse abdominal PAT data obtained by real photoacoustic scanning and annotated with an organ segmentation mask.

[0012] In some embodiments of the first aspect of the present disclosure, the hybrid two-branch model is trained based on the following loss function:

[0013]

[0014]

[0015]

[0016] in, represents the total loss, represents the Dice loss, represents the cross entropy loss of the CNN graffiti supervision part, represents the cross entropy loss of the graffiti supervision part of the Mamba U-network, represents the set of labeled pixels, represents the probability value of pixel i in category c, represents the pixel-level organ category probability value predicted by model k, express A labeled pixel in , c represents the category index, k represents the model branch identifier, represents the pixel-level organ category probability map predicted by CNN, Represents the pixel-level organ category probability map predicted by the Mamba U-network, represents the pseudo label of the hybrid two-branch model, Dice() represents the soft Dice coefficient between the predicted probability map and the pseudo label, Represents the fusion weight of CNN.

[0017] According to a second aspect of the present disclosure, a photoacoustic image enhancement device combining Mamba and CNN is provided, the photoacoustic image enhancement device combining Mamba and CNN comprising: a first segmentation unit, configured to obtain a first pixel-level organ category probability map of the photoacoustic image using a CNN; a second segmentation unit, configured to obtain a second pixel-level organ category probability map of the photoacoustic image using a Mamba U-type network; A weighted fusion unit is configured to weightedly fuse the first pixel-level organ category probability map and the second pixel-level organ category probability map to obtain a third pixel-level organ category probability map, wherein the third pixel-level organ category probability map includes a probability value of each pixel point in the photoacoustic image belonging to each type of organ.

[0018] According to a third aspect of the present disclosure, an electronic device is provided, comprising: one or more processors and a memory storing a program, wherein the program comprises instructions, and when the instructions are executed by the processor, the processor executes the above method.

[0019] According to a fourth aspect of the present disclosure, a computer-readable storage medium storing a program is provided, wherein the program includes instructions, which, when executed by one or more processors of a computing device, cause the computing device to perform the above-mentioned method.

[0020] It can be seen from the above technical solution that the embodiment of the present disclosure adopts a hybrid two-branch model including CNN and Mamba U-type network to achieve high-precision organ segmentation of photoacoustic images under weak supervision, which has both high-precision organ segmentation function and low computational overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0022] Figure 1 A flowchart of a photoacoustic image enhancement method combining Mamba and CNN provided in an embodiment of the present disclosure; Figure 2 Schematic diagram of the structure of the hybrid two-branch model involved in the embodiment of the present disclosure; Figure 3A schematic diagram of an abdominal MRI image and its corresponding simulated PAT image, organ mask, and graffiti label involved in an embodiment of the present disclosure; Figure 4 A schematic diagram of the structure of the Mamba U-type network involved in the embodiments of the present disclosure; Figure 5 Schematic diagram of the structure of the VSS Block involved in the embodiment of the present disclosure; Figure 6 Schematic diagram of parameter sensitivity analysis of weight α involved in the embodiment of the present disclosure; Figure 7 Schematic diagram for comparing the segmentation effects of synthetic data; Figure 8 Schematic diagram of the comparison of segmentation results of mouse PAT images; Figure 9 A schematic diagram of the structure of a photoacoustic image enhancement device combining Mamba and CNN provided in an embodiment of the present disclosure; Figure 10 A schematic structural block diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present disclosure in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0024] The terms used in the embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit the present disclosure. The singular forms "a," "an," "the," and "the" used in the embodiments of the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0025] As used herein, the words "if," "if," and the like may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.

[0026] As mentioned above, related technologies mainly use models such as CNN and visual transformers to achieve organ segmentation in photoacoustic images, which have problems such as fuzzy organ boundary segmentation, low segmentation accuracy, high computational overhead, low computational efficiency and limited usage.

[0027] In addition, the related technology still has the following problems: 1) High labeling cost: Pixel-level labeling of photoacoustic imaging usually requires manual drawing by professional physicians, which is extremely costly.

[0028] 2) Insufficient Adaptability of Weakly Supervised Strategies: Among various weakly supervised (WSL) strategies, traditional WSL methods, such as Conditional Random Fields (CRFs) and pseudo-labeling, struggle to effectively integrate multi-scale features. Their ability to collaboratively model global context and local details is insufficient, making them unsuitable for the aforementioned models used for organ segmentation in photoacoustic images. While scribble-based supervision is effective, the sparse nature of scribble annotations limits their ability to capture object boundaries and complex shapes, often leading to inaccurate segmentation near edges and in fuzzy regions, resulting in a sharp drop in segmentation accuracy in edge regions. Furthermore, the incompleteness of scribble labels poses a challenge to supervised loss functions.

[0029] In light of this, the present disclosure provides a method and apparatus for photoacoustic image enhancement that combines Mamba with a CNN. This method employs a hybrid two-branch model comprising a CNN and a Mamba U-shaped network to achieve high-precision organ segmentation in weakly supervised photoacoustic images, combining high-precision organ segmentation with low computational overhead. Furthermore, this hybrid two-branch model can be trained using a scribbling-based weak supervision method, achieving both high segmentation accuracy for organ edge regions and low annotation costs.

[0030] Figure 1 Schematic diagram showing the flow of the photoacoustic image enhancement method combining Mamba and CNN provided by the embodiment of the present disclosure. Figure 1 , the method of the embodiment of the present disclosure may include the following steps: Step 101, using CNN to obtain a first pixel-level organ category probability map of the photoacoustic image; Step 102, using a Mamba U-network to obtain a second pixel-level organ category probability map of the photoacoustic image; Step 103: weightedly fuse the first pixel-level organ category probability map and the second pixel-level organ category probability map through a fusion module to obtain a third pixel-level organ category probability map, where the third pixel-level organ category probability map includes the probability value of each pixel point in the photoacoustic image belonging to each type of organ.

[0031] The pixel-level organ category probability map involved in the embodiments of the present disclosure, namely the first pixel-level organ category probability map, the second pixel-level organ category probability map, and the third pixel-level organ category probability map, can be a single-channel probability map or a multi-channel probability map. In the multi-channel probability map, each channel corresponds to a semantic category (such as background, liver, etc.). For example, if C organ categories need to be segmented, the output is a C-channel probability map, and each channel stores the probability value of an organ category. Each channel is a matrix of the same size as the photoacoustic image, and each element value in the matrix ∈ [0,1] represents the probability that the pixel belongs to the organ category corresponding to the current channel. For example, in multi-organ segmentation, channel 1 can represent the probability that each pixel belongs to "liver" (the higher the value, the more likely it is liver), channel 2 can represent the probability that each pixel belongs to "background", and channel 3 can represent the probability that each pixel belongs to "kidney".

[0032] The photoacoustic image enhancement method combining Mamba and CNN provided in the embodiments of the present disclosure is implemented through a hybrid two-branch model, which includes the aforementioned CNN, Mamba U-type network and a fusion module for weighted fusion of the first pixel-level organ category probability map and the second pixel-level organ category probability map.

[0033] Figure 2 The structure of the hybrid two-branch model provided by the embodiment of the present disclosure and a schematic diagram of its processing process are shown. Figure 2 In the example, “(a)Visual Mamba-CNN” represents the custom name of the hybrid two-branch model, F cnn (X;θ) represents the processing process of CNN, F mamba (X;θ) represents the processing of the Mamba U-type network, X represents the input photoacoustic image (Input), θ represents the parameters, and both CNN and Mamba U-type networks can adopt the encoder-decoder structure, see Figure 2 The hybrid two-branch model includes two branches, one is CNN and the other is Mamba U-type network. The fusion module can fuse the organ segmentation result Pred1 predicted by CNN with the organ segmentation result Pred2 predicted by Mamba U-type network to obtain the final organ segmentation result.

[0034] See also Figure 2 , the hybrid two-branch model can be trained by Scribble labels. The pseudo label S can be used when training the hybrid two-branch model. pseudo , pseudo label S pseudo The fusion model can be used to fuse the organ segmentation result Pred1 predicted by CNN and the organ segmentation result Pred2 predicted by Mamba U-type network.

[0035] In some embodiments, a hybrid two-branch model can be trained based on a synthetic dataset, which can be constructed by obtaining pre-selected abdominal magnetic resonance imaging (MRI) images in a CHAOS dataset, where the abdominal MRI images are annotated with organ segmentation masks, converting the abdominal MRI images into simulated photoacoustic tomography (PAT) images, performing eroded skeletonization on the organ segmentation masks of the abdominal MRI images to generate graffiti labels along the central axes of the organs, and annotating the simulated PAT images with the graffiti labels to form a synthetic dataset.

[0036] Figure 3 An abdominal MRI image and its corresponding simulated PAT image, organ masks, and corresponding graffiti labels are shown. Figure 3 In the figure, from left to right are: abdominal MRI image, simulated PAT image converted from abdominal MRI image, organ mask of abdominal MRI image, and graffiti label.

[0037] In some examples, the K-Wave toolbox can be used to convert abdominal MRI images into simulated photoacoustic tomography (PAT) images. Specifically, a computational grid is established, i.e., a uniform mesh is constructed. A 20-point perfectly matched layer is added to suppress boundary reflections. Acoustic properties such as velocity and attenuation coefficient are assigned based on tissue type. 512 acoustic sensors are distributed in a circular array (4.5 cm radius). The K-Wave toolbox is used to simulate acoustic wave propagation, and the PAT image is reconstructed using the delay-sum (DAS) algorithm.

[0038] The hybrid two-branch model can be obtained through a two-stage training process. Specifically, the training process of the hybrid two-branch model includes: pre-training based on a synthetic dataset to determine the parameters of the hybrid two-branch model; and fine-tuning the parameters of the hybrid two-branch model based on real data, such as mouse abdominal PAT data obtained through real photoacoustic scanning and annotated with organ segmentation masks. Therefore, the strategy of pre-training with synthetic data and fine-tuning with real data can address the scarcity of real data and enable the hybrid two-branch model to better meet the needs of specific segmentation tasks.

[0039] In specific applications, a synthetic dataset can be constructed using 16 abdominal MRI images with a resolution of 224×224 from the CHAOS dataset. Mouse abdominal PAT data can be acquired using a 1064nm laser and a 512-channel array. For training, the hardware configuration can be an NVIDIA RTX 3090 GPU, with a batch size of 24, an SGD optimizer, a momentum of 0.9, and a learning rate of 0.01.

[0040] The hybrid two-branch model can be trained based on the loss functions shown in the following equations (1) to (3). Equation (4) shows how pseudo labels are generated.

[0041] (1) (2) (3) (4) in, represents the total loss, represents the Dice loss, represents the cross entropy loss of the CNN graffiti supervision part, represents the cross entropy loss of the graffiti supervision part of the Mamba U-network, represents the set of labeled pixels, represents the probability value of pixel i in category c, represents the pixel-level organ category probability value predicted by model k, express A labeled pixel in , c represents the category index, k represents the model branch identifier, represents the pixel-level organ category probability map predicted by CNN, Represents the pixel-level organ category probability map predicted by the Mamba U-network, represents the pseudo label of the hybrid two-branch model, Dice() represents the soft Dice coefficient between the predicted probability map and the pseudo label, Represents the fusion weight of CNN.

[0042] From the above, we can combine the Dice loss with the cross-entropy loss of the scribble supervision part of each branch to train the hybrid two-branch model, further improving the organ segmentation accuracy of the hybrid two-branch model.

[0043] The following describes in detail the specific implementation of each step of the photoacoustic image enhancement method combining Mamba and CNN provided in the embodiment of the present disclosure.

[0044] In a hybrid two-branch model, a CNN can include a decoder and an encoder. The encoder is used to gradually extract local features and compress spatial dimensions through convolution and downsampling, while the decoder is used to restore spatial details and fuse multi-level features through upsampling and skip connections. In some examples, the CNN decoder may include a 3×3 convolutional layer, a ReLU activation function layer, and a max pooling layer, while the CNN decoder may include a bilinear upsampling layer and skip connections. A CNN can preserve key local information, such as organ edges, during downsampling, restoring organ structure during reconstruction.

[0045] In step 101, the specific implementation process of using CNN to obtain the first pixel-level organ category probability map of the photoacoustic image may include the following steps a1 to a5: Step a1: Use a 3×3 convolutional layer to process the photoacoustic image to obtain local neighborhood features. The local neighborhood features include edge and texture features of the photoacoustic image. Step a2, using the RelU activation function layer to enhance the effective features in the local neighborhood features to obtain local enhanced features; The ReLU activation function layer can retain effective features, suppress noise, introduce nonlinearity, suppress negative responses, highlight significant features, enhance the activation values ​​of low-contrast areas, and prevent details such as organ edges from being submerged by noise.

[0046] Step a3: The local enhanced features are processed by a maximum pooling layer to further enhance the effective features in the local enhanced features to obtain compressed local features. The spatial size of the compressed local features is smaller than that of the local enhanced features, but the number of channels remains unchanged. Through the max pooling layer (i.e., downsampling), the strongest edges are preserved, background noise is compressed, and the feature map size is reduced while retaining the most significant features in the local area, thereby improving computational efficiency. For example, the max pooling layer can use a 2×2 window to avoid excessive compression that may cause the loss of details such as organ edges.

[0047] Step a4, using a bilinear upsampling layer to perform bilinear upsampling on the compressed local features to obtain semantic features, where the semantic features contain semantic context information of the organ in the photoacoustic image; Bilinear upsampling can amplify low-resolution compressed local features to their original size and restore their coarse spatial structure through interpolation. Bilinear upsampling is not only computationally efficient but also avoids checkerboard artifacts, making it suitable for smooth transition regions in photoacoustic images.

[0048] In step a5, the local enhanced features are fused with the semantic features through skip connections to obtain a first pixel-level organ category probability map. The first pixel-level organ category probability map includes the pixel-level position features of the organ edge in the photoacoustic image and the semantic context information of the organ.

[0049] Here, the fusion method may be, but is not limited to, channel concatenation, element-by-element addition, and the like.

[0050] The local enhancement features retain the edge, texture and other detailed information of the organs in the photoacoustic image, and can provide spatial details and local structures. The semantic features contain high-level semantic information such as target category and global context, which can provide semantic guidance and context perception. The first pixel-level organ category probability map is obtained by fusing the local enhancement features and semantic features through jump connections, and the local features can be restored with high precision.

[0051] The Mamba U-shaped network adopts a symmetrical encoder-decoder structure. The encoder and decoder are connected through a Visual State Space Block (VSS Block). The VSS Block has stronger long-range dependency modeling capabilities. By replacing traditional CNN or Transformer with the VSS Block, it can overcome the limitations of local receptive field, reduce computational complexity, and enhance the ability to capture semantic associations in photoacoustic images, while maintaining hardware-friendly linear efficiency. The encoder of the Mamba U-network is responsible for feature extraction and downsampling, while the decoder of the Mamba U-network is responsible for feature reconstruction and upsampling. The encoder and decoder in the Mamba U-network can each include multiple symmetrical layers, with skip connections connecting the corresponding layers of the encoder and decoder. For example, the encoder of the Mamba U-network can include a preprocessing layer that performs segmentation, linear transformation, and other processes, and multiple layers of encoding units. Each encoding unit includes a block merging layer and a VSS layer, and each VSS layer includes two visual state space modules connected in series. Similarly, the decoder of the Mamba U-network can include multiple layers of decoding units and post-processing layers that perform upsampling and linear projection. Each decoding unit includes a block expansion layer and a VSS layer, and each VSS layer includes two visual state space modules connected in series.

[0052] Specifically, step 102 may include the following steps b1 and b2: Step b1, performing encoder processing on the photoacoustic image to obtain context compression features; The encoder processing includes: segmenting the photoacoustic image to obtain a photoacoustic image block sequence, linearly transforming the photoacoustic image block sequence to obtain embedded features, and sequentially processing the embedded features through multiple levels of coding units to obtain contextual compression features. Each level of coding unit processing includes at least two levels of visual state space module processing in series. Step b2: performing decoder processing based on the variable context compression features and the output features processed by the encoding units at each level in the encoder to obtain a second pixel-level organ category probability map.

[0053] Among them, the decoder processing includes multi-level decoding unit processing, upsampling the sharpened reconstructed features obtained by multi-level decoding unit processing and performing linear projection processing. The processing of each level of decoding unit includes: upsampling the context compression features or the sharpened reconstructed features output by the previous level decoding unit to obtain upsampled features, fusing the upsampled features with the output features processed by the encoding unit at the same level to obtain fused features, and performing two-level visual state space module processing in series on the fused features to obtain the sharpened reconstructed features of the current level decoding unit.

[0054] Figure 4 A schematic diagram of an exemplary structure of a Mamba U-Net is shown. Figure 4 In the example, "(b) Mamba U-Net" is a custom name for the Mamba U-Net, which uses a 4-level symmetric encoder-decoder structure. Figure 4 The encoder and decoder each contain four layers. Skip connections connect the output features of the VSS Block × 2 modules in each encoder layer directly to the input of the VSSB Block × 2 modules in the corresponding decoder layer. This allows the encoder's rich features to be directly transmitted to the decoder through skip connections, helping the decoder recover these details during feature reconstruction and compensating for some of the spatial details lost by the encoder during layer-by-layer downsampling, thereby accurately segmenting organ boundaries in photoacoustic image segmentation.

[0055] See also Figure 4 , the corresponding layers of the encoder and decoder are respectively provided with VSS layers, and the VSS layer adopts two VSS Blocks connected in series, that is, a VSS Block×2 structure. In the VSS layer, the first VSS Block can be used to model local spatial dependencies to obtain a basic feature map containing organ contours and primary textures. The second VSS Block can be used to dynamically adjust parameters based on the basic feature map output by the first VSSBlock through the selective scanning mechanism of SS2D to focus on key areas to obtain a refined feature map containing cross-organ associations and global semantic enhancements. In addition, the VSS layer adopts a VSS Block×2 structure, and the output of the first layer of VSS Block is used as the input of the second layer of VSS Block to form a residual learning path, which can alleviate the vanishing gradient of the deep network and effectively improve the convergence stability of the model.

[0056] See also Figure 4 The specific implementation process of using the Mamba U-type network to obtain the second pixel-level organ category probability map of the photoacoustic image in step 102 may include the following steps c1 to c2: Step c1: The photoacoustic image is processed by the encoder of the Mamba U-network to obtain contextual compression features; Specifically, Figure 4 For example, the encoder processing flow of the Mamba U-type network may include the following steps c11 to c14: Step c11: input the photoacoustic image, divide the photoacoustic image into non-overlapping image blocks of fixed size through the patch partition layer to obtain a photoacoustic image block sequence, and project each image block in the photoacoustic image block sequence into a high-dimensional feature space through the linear transformation of the linear embedding layer to obtain embedded features. Figure 4 The first VSS Block×2 from top to bottom in the encoder part performs feature transformation on the embedded features to obtain context-aware features.

[0057] Converting photoacoustic images into a sequence of blocks preserves the original spatial topology and prevents smoothing of organ microstructures during subsequent convolution. Linear transformation converts each block in the photoacoustic image sequence (i.e., the photoacoustic signal of each small region in the image) into a feature vector containing abstract information. This not only improves feature representation and enhances the signal-to-noise ratio of weakly absorbing regions of an organ, but also avoids detail loss caused by early CNN pooling, while also reducing the dimensionality of subsequent processing.

[0058] The VSS Block process is performed twice on the feature vectors of each image block in the photoacoustic image block sequence corresponding to the embedded features to extract a photoacoustic feature block sequence. Each feature block in the photoacoustic feature block sequence corresponds to an image block in the photoacoustic image block sequence. Each feature block contains the edge texture features and semantic features of the photoacoustic signal in the corresponding area in the photoacoustic image. The photoacoustic feature block sequence is the aforementioned context-aware feature.

[0059] The VSS layer can effectively capture long-range dependencies and local features. By stacking two VSS Block processes, multi-scale features can be extracted from the high-dimensional feature representation of the photoacoustic image block sequence. Specifically, the first VSS Block process can extract edge texture features, while the second VSS Block process can extract semantic features such as vascular structure and tissue boundaries. This can capture the continuity of organ tissues (for example, vascular networks) in the photoacoustic image and the differences between different tissues.

[0060] The VSS Block utilizes a state-space model (SSM) to model the long-range dependencies of a sequence of photoacoustic image blocks. Its computational complexity is linear, resulting in high efficiency. The specific implementation of a single VSS Block is described below.

[0061] Step c12, through the first patch merging layer (ie, Figure 4 The first Patch Merging from top to bottom in the encoder part) downsamples the context-aware features to obtain the first downsampled features, which are then passed through the second VSS layer (i.e., Figure 4 The second VSS Block × 2 from top to bottom in the middle encoder part performs two superimposed VSS Block processes on the first down-sampled feature to obtain the first semantic enhancement feature; Here, the downsampling operation of patch merging can include performing 2×2 adjacent block merging and channel compression on context-aware features. This downsampling operation can reduce spatial resolution and increase feature dimensions. While compressing spatial dimensions, it also increases feature richness, integrating information over a larger receptive field, thereby capturing more macroscopic structural information such as organ contours. This also improves computational efficiency.

[0062] The twice-superimposed VSS Block processing in this step can model the spatial constraints of organs and capture the global context information in the photoacoustic image, such as the global position and structure of different organs.

[0063] Step c13, merge the layers through the second block (i.e., Figure 4 The second PatchMerging from top to bottom in the encoder part) downsamples the first semantic enhancement feature to obtain the second downsampled feature, which is then passed through the third VSS layer (i.e., Figure 4 The third VSS Block × 2 from top to bottom in the middle encoder part performs two superimposed VSS Block processes on the second down-sampled features to obtain the second semantic enhancement features; Step c14, merge the layers through the third block (i.e., Figure 4 The third PatchMerging from top to bottom in the encoder part) performs the final downsampling on the second semantic enhancement feature to obtain the third downsampling feature, which is then passed through the fourth VSS layer (i.e., Figure 4 The VSS Block × 2 connected between the encoder part and the decoder performs a VSS Block process of stacking twice on the third down-sampled feature to obtain the context compression feature of the global representation.

[0064] The ultimate downsampling through Patch Merging can condense the core semantics of organs. The superposition of two VSS Block processes on the ultimate downsampling features can capture the contextual information of the entire photoacoustic image, such as the spatial relationship between different organs, and establish cross-regional organ associations to obtain contextual compression features.

[0065] Step c2: Based on the context compression features obtained by the encoder processing of the Mamba U-type network and the output features of each VSS layer of the encoder, the decoder processing of the Mamba U-type network is performed to obtain a second pixel-level organ category probability map.

[0066] Specifically, Figure 4 For example, the decoder processing flow of the Mamba U-type network may include the following steps c21 to c25: Step c21, through the first block expansion layer (Patch Expanding) (ie, Figure 4 The first Patch Expanding from bottom to top in the decoder part) performs an upsampling operation on the context compressed feature to obtain the first upsampled feature, and the first upsampled feature is combined with the symmetrical VSS layer in the encoder (i.e., Figure 4 The output features of the third VSSBlock×2) from top to bottom in the encoder part (i.e., the second semantic enhancement feature) are fused to obtain the first fusion feature, which is passed through the first VSS layer of the decoder (i.e., Figure 4 The first VSS Block × 2 in the decoder from bottom to top performs two superimposed VSS Block processes on the first fusion feature to obtain the first sharpened reconstructed feature; The upsampling operation in the block expansion layer gradually restores the spatial dimensions of the image, bringing the resolution of the feature map close to that of the photoacoustic image. In photoacoustic image segmentation tasks, this process can help reconstruct the fine shape of organs. Fusion of the upsampled features with the output features of the symmetric layers in the encoder integrates multi-scale information through methods such as channel concatenation. Performing a VSS Block process by stacking the first fused feature twice can repair broken structures and enhance boundary contrast.

[0067] Step c22, through the second block extension layer (ie, Figure 4 The second PatchExpanding from bottom to top in the decoder part) performs an upsampling operation on the first sharpened reconstructed feature to obtain a second upsampled feature, and the second upsampled feature is combined with the VSS layer at a symmetrical position in the encoder (i.e., Figure 4 The output features of the second VSS Block×2 from top to bottom in the encoder part (i.e., the first semantic enhancement feature in the previous text) are fused to obtain the second fused feature, which is then passed through the second VSS layer of the decoder (i.e., Figure 4 The second VSS Block × 2 from bottom to top in the middle decoder part performs two superimposed VSS Block processes on the second fusion feature to obtain the second sharpened reconstructed feature; Step c23, through the third block extension layer (ie, Figure 4The decoder part of the third PatchExpanding from bottom to top performs an upsampling operation on the second sharpened reconstructed feature to obtain a third upsampled feature, and the third upsampled feature is compared with the VSS layer of the encoder at a symmetrical position (ie, Figure 4 The output features of the first VSS Block×2 from top to bottom in the encoder part (i.e., the context-aware features in the previous text) are fused to obtain the third fused features, which are then passed through the third VSS layer of the decoder (i.e., Figure 4 The third VSS Block × 2 from bottom to top in the middle decoder part performs a VSSBlock process of superimposing twice on the third fusion feature to obtain the third sharpened reconstructed feature; By fusing the third upsampled feature with the context-aware features obtained at the middle level of the encoder, we can integrate the highest resolution details into the upsampled features, achieving detail enhancement. Performing VSSBlock processing on the third fused feature by stacking it twice achieves final edge sharpening and noise suppression.

[0068] Step c24, through the fourth block extension layer (ie, Figure 4 The decoder part upsamples the third sharpened reconstructed feature from the fourth PatchExpanding from bottom to top to obtain the fourth upsampled feature, and performs linear projection processing of the linear projection layer (Linear Projection) on the fourth upsampled feature to map the features output by the decoder back to the pixel space of the photoacoustic image to obtain the second pixel-level organ category probability map.

[0069] In the disclosed embodiments, during the VSS Block ×2 processing of each VSS layer in the decoder, skip connections are used to provide detailed information from the VSS layer output at a symmetrical position in the encoder and upsampled features obtained by expansion of the previous block, gradually reconstructing spatial details and refining features. The fusion of upsampled features in the decoder with the output features of the same layer in the encoder can be, but is not limited to, channel concatenation.

[0070] In some examples, the patch merging operation can include: merging adjacent 2×2 blocks → number of channels × 4 → linear layer compression to 2x (e.g., D → 2D) to achieve spatial dimensionality reduction (H / 2, W / 2), similar to convolutional downsampling but preserving the block structure. Patch expanding is the opposite of patch merging, expanding (upsampling) the feature map blocks, reducing the number of channels and increasing spatial resolution.

[0071] The Visual State Space Block (VSS Block) combines the state space model with the characteristics of visual data, dynamically capturing the global contextual relationships in the image through a selective scanning mechanism while maintaining linear computational complexity.

[0072] Furthermore, each level of visual state space module processing in the two-level visual state space module processing in series includes: main path processing, the main path processing includes depth convolution operations, SS2D operator operations and SiLU activation function processing performed in sequence; gated branch path processing, the gated branch path processing includes fully connected layer processing for generating gating signals; gated multiplication and linear projection of the feature map obtained by the main path processing and the feature map obtained by the gated branch path processing; and, output feature map of the gated multiplication and linear projection and residual addition operation.

[0073] Figure 5 A schematic diagram of the structure of a VSS Block is shown. Figure 5 , VSS Block includes depth convolution (DWCNN), SS2D operator and SiLU activation component (not shown in the figure). The depth convolution is located in the left branch path, the SS2D operator is located in the center of the main path, and the SiLU activation component is not directly shown in the figure, but the multiplication operation (×) of the gated branch needs to be dynamically weighted through the SiLU activation component.

[0074] See also Figure 5 The processing of a single VSS Block may include the following steps d1 to d10: Step d1, input features Figure X , the shape is [H, W, C], H represents the height, W is the width, and C is the number of channels; Step d2: Layer Normalization (LayerNorm, LN) uses the second feature map to keep the shape unchanged; Specifically, the input X is layer-normalized to make its mean 0 and variance 1, eliminating the attenuation difference of the photoacoustic signal in deep tissue and reducing the internal covariate shift.

[0075] Step d3, linear projection (Linear Projection) to output the third feature map, the shape becomes [H, W, 2C]; Specifically, the number of channels can be mapped from C to a higher-dimensional space (e.g., C → 2C) through linear transformation to enhance the expressive power of features.

[0076] Step d4, main branch processing, that is, performing depthwise convolution (DW CNN) and SS2D operator operations on the third feature map to obtain the fifth feature map. The shape of the fifth feature map is still [H, W, 2C], but the features at each position are integrated with global context information; The third feature map is processed by a depthwise convolution (DW CNN) to output a fourth feature map. The fourth feature map is then processed by an SS2D operator to obtain a fifth feature map. The DW CNN process involves processing the third feature map using a depthwise separable convolution, which performs spatial convolution independently on each input channel. This is lightweight and can extract local features. The depthwise convolution does not change the number of channels, and the shape of the fourth feature map remains [H, W, 2C]. The SS2D operator performs a 2D selective scan on the fourth feature map, which is an operation in the state-space model (SSM). This expands the 2D feature map into a 1D sequence and applies the following equation (5) to model long-range dependencies. This operation captures the global contextual relationships in the image.

[0077] Step d5, gating branch processing, dynamic gating generation: The third feature map is processed through the fully connected layer to generate the gating signal, and the sixth feature map is output, and the shape is still [H, W, 2C].

[0078] Here, the gated branch is directly connected to the subsequent gated multiplication operation (×) without going through depthwise convolution and SS2D.

[0079] In step d6, layer normalization (LN) and linear projection (Linear) are performed on the fifth feature map to obtain the seventh feature map, which still has the shape of [H, W, 2C].

[0080] Here, layer normalization normalizes the feature map output by SS2D, and linear projection applies a fully connected layer again to perform feature transformation.

[0081] Step d7: Apply the SiLU activation function to the seventh feature map to output the eighth feature map. The shape of the eighth feature map remains unchanged, still [H, W, 2C]. Among them, the SiLU activation function combines linear and nonlinear characteristics and can enhance key features.

[0082] Step d8, gated multiplication (×), outputting the ninth feature map with a shape of [H, W, 2C]; Specifically, the eighth feature map obtained through SiLU activation is element-wise multiplied with the sixth feature map output by the gated branch in step d5. Through the gating mechanism, the weights generated by the gated branch dynamically adjust the features of the main branch, emphasizing important features and suppressing noisy or irrelevant features.

[0083] Step d9, linear projection (Linear), outputs the tenth feature map, and the shape becomes [H, W, C].

[0084] Here, the result of the gated multiplication, the ninth feature map, is linearly transformed, and the number of channels is adjusted (for example, from 2C to C) to output the tenth feature map. In step d10, the residuals are added (+) and the eleventh feature map is output.

[0085] Specifically, the tenth feature map output by step d9 is compared with the original input feature map Figure X The resulting eleventh feature map, with a shape of [H, W, C], is the output feature of the VSS block. Residual connections preserve the original features, preventing gradient vanishing or degradation in deep networks while allowing the learned features to be optimized as residuals.

[0086] The VSS Block uses deep convolution to capture local features, SS2D to model global dependencies, a gating mechanism to dynamically adjust feature importance, and residual connections to ensure gradient stability. This effectively models local-global feature relationships while maintaining low computational complexity, making it particularly suitable for processing multi-scale targets (such as organs) in photoacoustic images.

[0087] The VSS Block's main path includes an SS2D operator and short-path residual connections. SS2D models long-range dependencies with linear complexity, avoiding the high complexity of the self-attention mechanism in traditional Transformers. The gating mechanism enhances feature expressiveness and noise immunity, while residual connections ensure network training stability. This allows for a balanced approach to optical image segmentation, significantly improving segmentation accuracy (e.g., the Dice coefficient).

[0088] The SS2D operator in the VSS Block implements recursive global modeling through discrete state recursive equations. Combined with selective mechanisms and discretization transformations, it can dynamically adapt to the physical characteristics of photoacoustic images, surpassing the segmentation performance of traditional CNN and Transformer while maintaining linear complexity.

[0089] The discrete state recursion equation of the SS2D operator can be expressed as the following equations (5) to (6).

[0090] (5) , (6) in, Represents the discrete state vector at time t, which is a hidden state used to encode the input from the history ( arrive ). In photoacoustic images, It can represent the organ structure information along the scanning path up to the current position. Represents the discrete state transfer matrix, which is used to control the previous state How to transfer to the current state. In the photoacoustic image, Determines the degree of attenuation and retention of historical information. is a discrete input matrix that can be used to control the current input How to affect the state update. In the photoacoustic image, The influence weight of the current pixel on the state can be adjusted. Represents the input data at time t. In SS2D, the input data is the pixel value (or feature value) along the scan path (row or column). For example, in a horizontal scan, It may represent the pixel value of the i-th row and j-th column of the image. Represents the output projection matrix, which is used to transform the state vector Mapping to output . Represents the output at time t.

[0091] Equation (5) simulates a dynamic system, where the current state is determined by both the previous state and the current input. In photoacoustic image processing, recursively updating the state along the scanning path (e.g., from left to right, from top to bottom) can capture long-range dependencies. For example: 、 、 is the input conditional mapping, which is a learnable projection function used to map the input features to the parameter space. is the activation function, which is used to ensure Is positive.

[0092] represents the continuous input matrix at time t, represents the continuous output matrix at time t, The discretization step size of time t, a scalar or vector. Determines the granularity of the discretization of a continuous system and can be used to control the frequency or speed of state updates. 、 、 and step length By input Dynamic generation enables the SS2D operator to adapt its behavior based on the input content.

[0093] In photoacoustic imaging, The state update speed can be adjusted according to the importance of the input pixels. For example, at the boundary of an organ (where the feature changes greatly), a larger To quickly update the state; in a uniform tissue, a smaller To maintain stability.

[0094] above 、 、 The defined selectivity mechanism enables the model to adapt to the local characteristics of the photoacoustic image, thereby improving the segmentation accuracy.

[0095] In order to discretize the underlying continuous-time dynamics, the state transition matrix can be calculated using the following equations (7)~(8).

[0096] (7) (8) Among them, exp is the matrix exponential operation, I is the identity matrix, To invert the matrix, Represents the continuous state transition matrix, which is used to capture long-range dependencies.

[0097] The transformation of Equations (7) to (8) transforms the continuous-time state space model into a discrete-time model (i.e., recursive equation), making it applicable to image pixel sequences. The discretization process is affected by Impact, greater will make Decay faster, and smaller make Decays more slowly.

[0098] In the SS2D operator, the aforementioned formulas (5) to (8) work together to achieve the following processing: 1) Selective generation of parameters, that is, dynamically generating according to the input features 、 and step length ; 2) Discretization transformation: use Δ and fixed A to calculate discrete parameters 、 3) Discrete State Recursion: Recursively compute states and outputs along the scan path. In the task of organ segmentation in photoacoustic images, this mechanism enables the SS2D operator to adapt to image content (e.g., organ vs. background), efficiently model long-range dependencies (e.g., organs throughout the image), maintain linear computational complexity while modeling global context, and encode the biophysical properties of photoacoustic images as differentiable mathematical processes through discrete state equations, achieving efficient and accurate global modeling.

[0099] In step 103, a weighted summation of the first pixel-level organ category probability map and the second pixel-level organ category probability map may be performed based on the following formula (9) to obtain a third pixel-level organ category probability map.

[0100] (9) Among them, P represents the third pixel-level organ category probability map, represents the first pixel-level organ category probability map, Represents the second pixel-level organ category probability map, and the value range of weight α is [0.7, 0.8].

[0101] Figure 6 Figure 2 shows the results of parameter sensitivity analysis of weight α. Figure 6 , the Dice coefficient is optimal when α∈[0.7,0.8], with a peak value of 0.652.

[0102] Figure 7 A schematic diagram showing the comparison of synthetic data segmentation effects is shown. Figure 7 In , Image represents the photoacoustic image, Ours represents the segmentation result of the embodiment of the present disclosure, Ground Label represents the true label of the photoacoustic image, and pCE+UNet, USTM+Unet, Mumford+Unet, pCE+SwinUNet+UNet, USTM+SwinUNet, Mumford+SwinUNet, and Gated CRF+SwinUNet represent the segmentation results of the corresponding models respectively. Figure 7 It can be seen that the method (Ours) of the embodiment of the present disclosure is more accurate in segmenting the boundaries of the liver (yellow) and kidney (blue) in the photoacoustic image.

[0103] Table 1 below shows a quantitative performance comparison of the method of the embodiment of the present disclosure and the organ segmentation method based on traditional SwinUNet and Gated CRF.

[0104] Table 1

[0105] As can be seen from the above, the method of the embodiment of the present disclosure can achieve at least the following technical effects: 1) Segmentation accuracy is effectively improved: The disclosed embodiment improves the Dice coefficient by 8.5% and reduces HD95 by 36.7% through global dependency modeling (Mamba) + local detail preservation (CNN), fully demonstrating that the organ boundary segmentation results obtained by the method of the disclosed embodiment are more accurate.

[0106] 2) High computational efficiency: Mamba's linear complexity (O(n)) is smaller than Transformer's complexity (O(n²), and training takes only 4 hours per model.

[0107] 3) Annotation costs are effectively reduced: Graffiti annotation takes only 10% of the time of pixel-level annotation. Pre-training with synthetic data can effectively reduce reliance on real annotations.

[0108] 4) Enhanced generalization: A training strategy that combines pre-training with synthetic data and fine-tuning with real data can address the problem of scarce training data and reduce training costs.

[0109] Figure 8 Schematic diagram showing the segmentation results of mouse PAT data. Figure 8 In the figure, Image represents the mouse PAT image, Ours represents the segmentation result of the mouse PAT image according to the embodiment of the present disclosure, Ground Label represents the true label of the mouse PAT image, Scribble Label represents the graffiti label of the mouse PAT image, and pCE+UNet, USTM+Unet, Mumford+Unet, pCE+SwinUNet+UNet, USTM+SwinUNet, and Gated CRF+SwinUNet represent the segmentation results of the corresponding models respectively. Figure 8 The method of the embodiment of the present disclosure can accurately segment the contours of the liver, kidney, and spleen (red frame area) in the mouse PAT image.

[0110] Figure 9 FIG1 shows a schematic diagram of the structure of a photoacoustic image enhancement device combining Mamba and CNN provided by an embodiment of the present disclosure, and the photoacoustic image enhancement device combining Mamba and CNN is applied to an electronic device 1000. Figure 9 , a photoacoustic image enhancement device 900 combining Mamba and CNN may include: A first segmentation unit 901 is configured to obtain a first pixel-level organ category probability map of the photoacoustic image using CNN; A second segmentation unit 902 is configured to obtain a second pixel-level organ category probability map of the photoacoustic image using a Mamba U-type network; The weighted fusion unit 903 is configured to weightedly fuse the first pixel-level organ category probability map and the second pixel-level organ category probability map to obtain a third pixel-level organ category probability map, wherein the third pixel-level organ category probability map includes a probability value of each pixel point in the photoacoustic image belonging to each type of organ.

[0111] Furthermore, the second segmentation unit 902 can be specifically used to: perform encoder processing on the photoacoustic image to obtain contextual compression features, the encoder processing includes: segmenting the photoacoustic image to obtain a photoacoustic image block sequence, linearly transforming the photoacoustic image block sequence to obtain embedded features, and the embedded features are sequentially processed by multi-level encoding units to obtain contextual compression features, and each level of encoding unit processing includes at least two levels of visual state space module processing in series; and based on the contextual compression features and the output features obtained by the encoding units at each level, perform decoder processing to obtain the second pixel-level organ category probability map, the decoder processing includes: multi-level decoding unit processing, upsampling the sharpened reconstruction features obtained by the multi-level decoding unit processing and performing linear projection processing, and each level of decoding unit processing includes: upsampling the contextual compression features or the sharpened reconstruction features output by the previous level decoding unit to obtain upsampled features, fusing the upsampled features with the output features processed by the symmetric level encoding unit to obtain fused features, and performing two levels of visual state space module processing in series on the fused features to obtain the sharpened reconstruction features of the current level decoding unit.

[0112] Furthermore, the weighted fusion unit 903 may be specifically configured to perform weighted summation on the first pixel-level organ category probability map and the second pixel-level organ category probability map based on the aforementioned formula (9) to obtain a third pixel-level organ category probability map.

[0113] For further technical details regarding the photoacoustic image enhancement device 900 combining Mamba and CNN, please refer to the previous section on the photoacoustic image enhancement method combining Mamba and CNN, and will not be repeated here. In specific applications, the photoacoustic image enhancement device 900 combining Mamba and CNN can be implemented using the electronic device 1000 described below, or as software within the electronic device 1000.

[0114] In addition, an embodiment of the present disclosure further provides a computer-readable storage medium on which a computer program is stored. The program includes instructions that, when executed by one or more processors of a computing device, perform the steps of the aforementioned photoacoustic image enhancement method combining Mamba and CNN.

[0115] Figure 10 Schematic diagram of the structure of the electronic device provided by the embodiment of the present disclosure is shown. Figure 10 The electronic device 1000 may include: one or more processors 1001, and a memory 1002 storing one or more programs, which are executed by the one or more processors 1001 to implement the method flow shown in the above embodiments of the present disclosure and / or the program units corresponding to each unit in the device.

[0116] The various components are interconnected using various buses and may be mounted on a common motherboard or in other configurations as needed. Processor 1001 may process instructions for execution within the electronic device, including instructions stored in or on memory for displaying graphical information for a user interface on an external input / output device (such as a display device coupled to the interface). In other embodiments, multiple processors and / or multiple buses may be used with multiple memories and multiple storage devices, if desired.

[0117] The processor 1001 may include one or more single-core processors or multi-core processors. The processor 1001 may include any combination of general-purpose processors or dedicated processors (such as an image processor, an application processor, a baseband processor, etc.).

[0118] The memory 1002 is a computer-readable storage medium provided by the present disclosure, which can be used to store non-transient software programs, non-transient computer executable programs and units, such as the embodiment of the present disclosure. Figure 1 The processor 1001 executes the non-transient software programs, instructions and units stored in the memory 1002 to perform the above-mentioned method embodiment. Figure 1 The programs, instructions, and units corresponding to the photoacoustic image enhancement method combining Mamba and CNN are shown.

[0119] The electronic device 1000 may further include an input device 1003 and an output device 1004. The processor 1001, the memory 1002, the input device 1003 and the output device 1004 may be connected via a bus or other means. Figure 10 The bus connection is taken as an example.

[0120] The programs (also referred to as software, software applications, or code) described above include machine instructions for a programmable processor and may be implemented using an object-oriented programming language, assembly, or machine language.

[0121] Over time and with the advancement of technology, the meaning of "medium" has become increasingly broad. The dissemination of computer programs is no longer limited to tangible media and can also be directly downloaded from the Internet, for example. Any combination of one or more computer-readable storage media may be used. Computer-readable storage media may include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component.

[0122] The technical solutions provided by the present disclosure are described in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present disclosure. The descriptions of the above embodiments are intended only to help understand the methods and core concepts of the present disclosure. Furthermore, those skilled in the art will appreciate that variations in the specific implementation methods and scope of application may occur based on the concepts of the present disclosure. In summary, the contents of this specification should not be construed as limiting the present disclosure.

[0123] The above description is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent substitutions, etc. made within the spirit and principles of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. A photoacoustic image enhancement method combining Mamba and CNN, characterized in that: The method is implemented by a pre-trained hybrid two-branch model, which includes a CNN, a Mamba U-shaped network, and a fusion module; the method includes: Use CNN to obtain the first pixel-level organ category probability map of the photoacoustic image; Obtaining a second pixel-level organ category probability map of the photoacoustic image using a Mamba U-network; The first pixel-level organ category probability map and the second pixel-level organ category probability map are weightedly fused by the fusion module to obtain a third pixel-level organ category probability map, wherein the third pixel-level organ category probability map includes a probability value of each pixel point in the photoacoustic image belonging to each type of organ.

2. The method according to claim 1, characterized in that The Mamba U-type network adopts a symmetrical encoder-decoder structure; the method of obtaining a second pixel-level organ category probability map of the photoacoustic image using the Mamba U-type network includes: performing encoder processing on the photoacoustic image to obtain contextual compression features, the encoder processing comprising: segmenting the photoacoustic image to obtain a sequence of photoacoustic image blocks, linearly transforming the sequence of photoacoustic image blocks to obtain embedded features, and sequentially processing the embedded features through multiple levels of encoding units to obtain contextual compression features, wherein each level of encoding unit processing comprises at least two levels of visual state space module processing in series; Based on the context compression features and the output features obtained by the encoding units at each level, decoder processing is performed to obtain the second pixel-level organ category probability map. The decoder processing includes: multi-level decoding unit processing, upsampling the sharpened reconstruction features obtained by the multi-level decoding unit processing and performing linear projection processing. The decoding unit processing at each level includes: upsampling the context compression features or the sharpened reconstruction features output by the previous level decoding unit to obtain upsampled features, fusing the upsampled features with the output features processed by the symmetrical hierarchical encoding units to obtain fused features, and performing two-level visual state space module processing in series on the fused features to obtain the sharpened reconstruction features of the current level decoding unit.

3. The method according to claim 2, characterized in that Each level of visual state space module processing in the two-level series visual state space module processing includes: Main path processing, wherein the main path processing includes sequentially executing a depthwise convolution operation, an SS2D operator operation, and a SiLU activation function processing; Gated branch path processing, the gated branch path processing including fully connected layer processing for generating a gating signal; performing gated multiplication and linear projection of the feature map obtained by processing the main path and the feature map obtained by processing the gated branch path; and The output feature map of the gated multiplication and linear projection is added to the residual.

4. The method according to claim 1, wherein The weighted fusion of the first pixel-level organ category probability map and the second pixel-level organ category probability map by the fusion module to obtain a third pixel-level organ category probability map includes: The fusion module performs weighted summation on the first pixel-level organ category probability map and the second pixel-level organ category probability map based on the following formula to obtain a third pixel-level organ category probability map; Among them, P represents the third pixel-level organ category probability map, represents the first pixel-level organ category probability map, Represents the second pixel-level organ category probability map, and the value range of weight α is [0.7, 0.8].

5. The method according to claim 1, wherein The hybrid two-branch model is trained based on a synthetic dataset, which is constructed as follows: Obtain a pre-selected abdominal MRI image in the CHAOS dataset, wherein the abdominal MRI image is annotated with an organ segmentation mask; The abdominal MRI image is converted into a simulated photoacoustic tomography (PAT) image, the organ segmentation mask of the abdominal MRI image is subjected to erosion skeletonization processing to generate graffiti labels along the central axis of the organ, and the simulated PAT image is annotated with the graffiti labels to form a synthetic dataset.

6. The method according to claim 5, characterized in that The hybrid two-branch model is trained in the following way: Pre-training based on the synthetic dataset to determine parameters of the hybrid two-branch model; The parameters of the hybrid two-branch model were fine-tuned based on real data, including mouse abdominal PAT data obtained by real photoacoustic scanning and annotated with organ segmentation masks.

7. The method according to claim 1 or 5, characterized in that The hybrid two-branch model is trained based on the following loss function: in, represents the total loss, represents the Dice loss, represents the cross entropy loss of the CNN graffiti supervision part, represents the cross entropy loss of the graffiti supervision part of the Mamba U-network, represents the set of labeled pixels, represents the probability value of pixel i in category c, represents the pixel-level organ category probability value predicted by model k, express A labeled pixel in , c represents the category index, k represents the model branch identifier, represents the pixel-level organ category probability map predicted by CNN, Represents the pixel-level organ category probability map predicted by the Mamba U-network, represents the pseudo label of the hybrid two-branch model, Dice() represents the soft Dice coefficient between the predicted probability map and the pseudo label, Represents the fusion weight of CNN.

8. A photoacoustic image enhancement device combining Mamba and CNN, characterized in that: The photoacoustic image enhancement device combining Mamba and CNN includes: a first segmentation unit, configured to obtain a first pixel-level organ category probability map of the photoacoustic image using a CNN; a second segmentation unit, configured to obtain a second pixel-level organ category probability map of the photoacoustic image using a Mamba U-type network; A weighted fusion unit is configured to weightedly fuse the first pixel-level organ category probability map and the second pixel-level organ category probability map to obtain a third pixel-level organ category probability map, wherein the third pixel-level organ category probability map includes a probability value of each pixel point in the photoacoustic image belonging to each type of organ.

9. An electronic device, characterized in that: include: One or more processors and a memory storing a program, wherein the program includes instructions, and when the instructions are executed by the processor, the processor performs the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a program, the program comprising instructions, which, when executed by one or more processors of a computing device, cause the computing device to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Head and neck organ image segmentation method and device, electronic equipment and storage medium

    CN113487622A

  • AI assistance-based parathyroid gland identification method and device

    CN119722659A

  • Neural network infrared small target detection method and system fusing CNN and Mama

    CN119992064A

  • Unsupervised medical image registration method based on Mamba-CNN (Convolutional Neural Network) coding feature fusion

    CN120259385A

  • Super-pixel segmentation method and system based on Mamba state space model

    CN120279271A

Cited By

  • Mangrove forest dense time sequence detection method and device

    CN121259613A

  • Arc array photoacoustic tomography blood vessel recognition and mask generation method based on physical simulation

    CN121353306A

  • Deep sea polymetallic nodule image segmentation method based on multi-modal data fusion

    CN121937472A