Photoacoustic image reconstruction method and system
By employing a multi-velocity radio frequency signal fusion learning method, the problems of photoacoustic image artifacts and blurring caused by sound velocity mismatch in heterogeneous biological tissues using linear array transducers were solved, achieving high-quality photoacoustic image reconstruction and improving image accuracy and the reliability of clinical diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-19
- Publication Date
- 2026-03-27
AI Technical Summary
In photoacoustic imaging of heterogeneous biological tissues, linear array transducers suffer from image artifacts, structural blurring, and distortion due to sound velocity mismatch, which affects image quality and the reliability of clinical diagnosis.
A multi-velocity radio frequency signal fusion learning method is adopted. Through a time-of-flight inversion module, a beamforming enhancement module, and a multi-scale feature fusion network, photoacoustic image reconstruction under multi-velocity conditions is adaptively processed. This includes array element coordinate definition, spatial radio frequency signal matrix construction, beamforming, and tubular structure feature extraction and fusion processing.
It enables adaptive reconstruction of high-quality photoacoustic images in heterogeneous biological tissues, significantly reducing artifacts, improving vascular continuity and anatomical structure accuracy, and enhancing imaging quality and clinical diagnostic reliability.
Smart Images

Figure CN121730746A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical imaging technology and artificial intelligence, and in particular to a photoacoustic image reconstruction method and system based on multi-velocity radio frequency signal fusion learning. Background Technology
[0002] Photoacoustic imaging is an emerging biomedical imaging technology that reconstructs images of internal structures by detecting the ultrasound waves (i.e., photoacoustic signals) generated when tissue absorbs pulsed laser light. Linear array transducers are favored in preclinical research and clinical applications due to their ease of integration and operational flexibility.
[0003] Beamforming based on delay superposition is one of the most commonly used image reconstruction methods in linear array photoacoustic imaging. The general process of this method is as follows: assuming that the sound velocity in biological tissue is uniform, the time delay of the signal received by each sensor unit is calculated based on this fixed sound velocity, and then the delayed signals of all channels are superimposed to reconstruct the distribution image of the initial sound pressure.
[0004] However, this reconstruction process based on the assumption of uniform sound velocity has inherent and unavoidable technical flaws when dealing with real biological tissues, which have non-uniform sound velocity distributions. Biological tissues are heterogeneous bodies with varying sound velocities. If the reconstruction algorithm adopts a fixed sound velocity assumption, it will introduce signal time delay errors, resulting in geometric distortion, double artifacts, and blurred blood vessels in the reconstructed image. The one-dimensional arrangement of the linear array transducers limits its receiving angle, making it impossible to capture signals in all directions. This leads to severe stripe artifacts in the reconstructed image and disrupts or even eliminates the continuity of vascular structures perpendicular to the array. These two major problems severely restrict image quality and the reliability of clinical diagnosis. Summary of the Invention
[0005] In view of the above problems, the present invention provides a photoacoustic image reconstruction method and system for overcoming or at least partially solving the above problems. It solves the problems of image artifacts, structural blurring, and distortion caused by sound velocity mismatch in photoacoustic imaging of heterogeneous biological tissues using linear array transducers. Simultaneously, it solves the problem of severely degraded image quality in linear array photoacoustic computed tomography (PACT) in heterogeneous tissues due to sound velocity mismatch and limited viewing angle.
[0006] This invention provides the following solution: A photoacoustic image reconstruction method includes: Acquire the raw radio frequency signal collected by the array transducer and determine the probe parameters, imaging range, target resolution, and target sound velocity; The time-of-flight inversion module defines the array element coordinate vector in physical space based on the element spacing. Based on the imaging range and target resolution, an image reconstruction grid covering the region of interest is established. The distance between each imaging grid point and each array element is calculated, and the flight time from each grid point to each array element is obtained based on the target sound velocity. According to the sampling position corresponding to the flight time, the corresponding amplitude is obtained from the original radio frequency signal through interpolation, and the amplitude is filled into the corresponding position of the image reconstruction grid to obtain a spatial radio frequency signal matrix that contains spatial information for a single sound velocity while retaining the original radio frequency signal. Multiple sets of sound velocities are processed repeatedly and spliced along the sound velocity dimension or channel dimension to obtain a multi-sound velocity spatial radio frequency signal matrix. The multisonic spatial radio frequency signal matrix is subjected to enhanced beamforming processing using a beamforming enhancement module. The enhanced beamforming processing includes time alignment and weighted superposition of spatial radio frequency signals of different array elements in each sound speed channel to suppress sidelobes and defocus artifacts. In the imaging plane, a spatial weight distribution of the same size as the imaging grid is formed according to the response characteristics of each position and local neighborhood. The synthesis results at different spatial positions are selectively enhanced and suppressed to obtain multiple initial reconstructed images from the multisonic spatial radio frequency signal matrix. A multi-scale feature fusion and enhancement network is used to extract and fuse tubular features from multiple initial reconstructed images to obtain photoacoustic reconstructed images. The tubular feature extraction is used to highlight vascular structures that are distributed in tubular or strip-like patterns and to suppress point-like noise and block artifacts. The fusion processing is used to integrate complementary structural information from the initial results of different sound velocities and different beams, so that small blood vessels and large-diameter blood vessels are presented continuously and smoothly in the same reconstructed image.
[0007] Preferably, defining the array element coordinate vector in physical space based on the array element spacing includes: establishing a regular grid of the imaging area in physical space based on the physical size of the imaging area and the preset imaging resolution, and obtaining the physical coordinates of each grid point; determining the physical coordinates of each array element in the same coordinate system based on the array element spacing and array arrangement.
[0008] Preferably, the workflow of the beamforming enhancement module includes: The first-level convolutional unit receives the multisonic spatial radio frequency signal matrix, so that the first-level convolutional unit can perform local spatial convolution operations on the input feature map and compress the number of channels to extract more compact spatial radio frequency features. After the convolution operation, a spatial attention module is connected to adaptively adjust the weights of different positions according to the response differences of each spatial position and its local neighborhood, in order to compensate for micron-level time delay errors and reduce local artifacts caused by subtle time delay deviations. By receiving the outputs of the first-level convolutional unit and the spatial attention module from the second-level convolutional unit, the second-level convolutional unit effectively forms a large receptive field through two concatenated convolutional kernels, aggregating neighborhood information over a large spatial range and further compressing the number of channels to reduce redundant channels. After the second-level convolution, the spatial attention module is connected again. By weighting the response consistency between different array elements and different sound velocity channels at the same spatial position, the effect of cross-array phase is synchronized, thereby enhancing the main lobe response and suppressing defocus artifacts caused by phase inconsistency. The third-level convolutional unit receives the output of the second-level convolutional unit, which further compresses the number of channels while maintaining the local convolutional form. This allows the network to reorganize the spatial radio frequency characteristics within a large effective receptive field, correcting macroscopic wavefront distortions caused by factors such as sound velocity mismatch and probe placement. After the third-level convolution, the final spatial attention module is set to selectively suppress the responses in background and high-noise regions, thereby highlighting the coherent signals in the regions where target structures such as blood vessels are located.
[0009] Preferably, the stride of each convolutional stage is set to 1, and it is used in conjunction with a non-linear activation function to enhance the feature representation capability.
[0010] Preferably, the workflow of the spatial attention module includes: The input feature map is subjected to strip-shaped average pooling operations in the horizontal and vertical directions respectively. One-dimensional average pooling is performed in the horizontal direction to obtain the width direction feature vector representing the global response of each column; one-dimensional average pooling is performed in the vertical direction to obtain the height direction feature vector representing the global response of each row, thereby extracting global context information in the horizontal and vertical directions respectively. One-dimensional convolution operations are applied to the width-direction feature vector and the height-direction feature vector respectively to model the spatial dependency and response pattern difference between adjacent positions in their respective directions, thereby obtaining intermediate features that include directional correlation. The one-dimensional convolution outputs in both directions are group normalized to eliminate scale differences between different samples and different locations. The weight values in each direction are constrained to between 0 and 1 by the Sigmoid activation function to obtain the normalized horizontal and vertical weight vectors. The horizontal weight vector and the vertical weight vector are multiplied by an outer product to generate a bidirectional coupled two-dimensional spatial weight map. The two-dimensional spatial weight map is then multiplied element-wise with the original feature map to selectively enhance and suppress the response at each spatial location, thereby highlighting the target area with continuous structure and high contrast, and suppressing background noise and scattering artifacts.
[0011] Preferably, the multi-scale feature fusion and enhancement network includes a tubular structure feature extraction module and a tubular structure multi-branch convolutional unit; the tubular structure feature extraction module is used to adaptively fit the trajectory of slender structures along different directions using a serpentine convolutional kernel, and the tubular structure multi-branch convolutional unit is used to extract complementary features of small blood vessels and large blood vessels under different receptive fields.
[0012] Preferably, the tubular structure feature extraction module comprises an encoder module, a decoder module, and a skip connection module; The encoder module is used to downsample the input photoacoustic initial reconstructed image or feature map step by step and extract high-level semantic features; the decoder module is used to upsample the features output by the encoder step by step and restore spatial resolution; the skip connection module is used to concatenate the features of each layer of the encoder with the features of the corresponding layer of the decoder in the channel dimension to preserve high-resolution spatial detail information.
[0013] Preferably, the tubular structure multi-branch convolutional unit is embedded in each layer of convolutional operation of the encoder and the decoder to explicitly enhance the vascular tubular structure distributed along different directions and to fuse local and global features at multiple scales.
[0014] Preferably, the tubular multi-branch convolutional unit takes the same input feature map as input, and the tubular multi-branch convolutional unit includes parallel standard convolutional branches, horizontal serpentine convolutional branches, and vertical serpentine convolutional branches; The standard convolutional branch uses a standard convolutional kernel to perform local convolution operations on the input feature map, which is used to extract basic local features such as edges and textures, ensuring that the network has the ability to represent conventional image structures. The horizontal serpentine convolution branch is used to employ a serpentine convolution kernel that extends in the horizontal direction. By predicting a set of one-dimensional offsets and deforming the initial straight convolution kernel, the sampling position of the convolution kernel forms a smooth, curved trajectory in the lateral direction, which is used to produce a higher response to slender tubular structures in and around the horizontal direction. The vertical serpentine convolution branch is used to employ a serpentine convolution kernel that extends along the vertical direction. Through offset prediction and deformation operations, the sampling position of the convolution kernel forms a smooth, curved trajectory along the axial direction, which is used to generate a higher response to tubular structures that are oriented vertically and in the vicinity.
[0015] A photoacoustic image reconstruction system for performing the above-described photoacoustic image reconstruction method, the system comprising: The raw radio frequency signal acquisition unit is used to acquire the raw radio frequency signal collected by the array transducer and determine the probe parameters, imaging range, target resolution and target sound velocity. The time-of-flight interpolation and inversion module unit is used to define array element coordinate vectors in physical space based on the array element spacing using the time-of-flight inversion module; establish an image reconstruction grid covering the region of interest based on the imaging range and the target resolution; calculate the distance between each imaging grid point and each array element; and calculate the flight time from each grid point to each array element based on the target sound velocity; obtain the corresponding amplitude from the original radio frequency signal through interpolation based on the sampling position corresponding to the flight time; and fill the amplitude into the position corresponding to the image reconstruction grid to obtain a spatial radio frequency signal matrix that contains spatial information for a single sound velocity while retaining the original radio frequency signal; repeat the processing for multiple sets of sound velocities and stitch them together along the sound velocity dimension or channel dimension to obtain a multi-sound velocity spatial radio frequency signal matrix; A beamforming enhancement unit is used to perform enhanced beamforming processing on the multisonic spatial radio frequency signal matrix using a beamforming enhancement module. The enhanced beamforming processing includes time alignment and weighted superposition of spatial radio frequency signals of different array elements in each sound speed channel to suppress sidelobes and defocus artifacts. In the imaging plane, a spatial weight distribution of the same size as the imaging grid is formed according to the response characteristics of each position and local neighborhood. The synthesis results at different spatial positions are selectively enhanced and suppressed to obtain multiple initial reconstructed images from the multisonic spatial radio frequency signal matrix. The tubular structure feature extraction unit is used to extract and fuse tubular features from multiple initial reconstructed images using a multi-scale feature fusion and enhancement network to obtain a target photoacoustic reconstructed image. The tubular feature extraction is used to highlight vascular structures that are distributed in tubular or strip-like shapes and to suppress point-like noise and block artifacts. The fusion processing is used to integrate complementary structural information from the initial results of different sound velocities and different beams, so that small blood vessels and large-diameter blood vessels are presented continuously and smoothly in the same reconstructed image.
[0016] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects: This invention provides a photoacoustic image reconstruction method and system that, after acquiring the original radio frequency signal and probe parameters, can adapt to different sensor models without requiring a priori information on tissue sound velocity. Users only need to input the acquired radio frequency signal and can choose to provide five values covering the possible range of sound velocities in the tissue or use the system default settings. The system automatically fuses signal features from multiple sound velocities to intelligently reconstruct photoacoustic images with low artifacts, high vascular continuity, and accurate anatomical structures, significantly improving imaging quality and the reliability of clinical diagnosis while ensuring reconstruction efficiency. Simultaneously, this method effectively solves the problems of photoacoustic image artifacts and structural blurring caused by uneven sound velocity and limited viewing angle. It also possesses strong generalization ability; one model can adapt to multiple probe parameters, ensuring physical rationality while balancing reconstruction efficiency and clinical applicability.
[0017] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly described below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0019] Figure 1 This is a flowchart of a photoacoustic image reconstruction method provided in an embodiment of the present invention; Figure 2 This is a schematic flowchart of a method for reconstructing photoacoustic images from radio frequency signals according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the process for calculating a multi-velocity spatial radio frequency signal matrix from the original detection signal, provided in an embodiment of the present invention. Figure 4 This is a schematic flowchart of the beamforming enhancement module provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the spatial attention module provided in an embodiment of the present invention; Figure 6 This is the multi-scale feature fusion network architecture with an embedded tubular structure feature extraction module provided in the embodiments of the present invention; Figure 7 This is a schematic diagram of the tubular structure feature extraction module provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of a photoacoustic image reconstruction system provided in an embodiment of the present invention; Figure 9 This is a schematic diagram of a photoacoustic image reconstruction device provided in an embodiment of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention are within the scope of protection of the present invention.
[0021] See Figure 1 This invention provides a photoacoustic image reconstruction method, such as... Figure 1 As shown, the method may include: S101: Acquire the raw radio frequency signal collected by the array transducer and determine the probe parameters, imaging range, target resolution, and target sound velocity; S102: Using the time-of-flight inversion module, define the array element coordinate vector in physical space based on the array element spacing; based on the imaging range and the target resolution, establish an image reconstruction grid covering the region of interest, calculate the distance between each imaging grid point and each array element, and calculate the flight time from each grid point to each array element based on the target sound speed; based on the sampling position corresponding to the flight time, obtain the corresponding amplitude from the original radio frequency signal through interpolation, and fill the amplitude into the position corresponding to the image reconstruction grid to obtain a spatial radio frequency signal matrix that contains spatial information for a single sound speed while retaining the original radio frequency signal; repeat the processing for multiple sets of sound speeds and splice them along the sound speed dimension or channel dimension to obtain a multi-sound speed spatial radio frequency signal matrix; in specific implementation, the embodiment of the present invention can provide that defining the array element coordinate vector in physical space based on the array element spacing includes: establishing a regular grid of the imaging region in physical space based on the physical size of the imaging region and the preset imaging resolution to obtain the physical coordinates of each grid point; determining the physical coordinates of each array element in the same coordinate system based on the array element spacing and array arrangement.
[0022] S103: The beamforming enhancement module is used to perform enhanced beamforming processing on the multisonic spatial radio frequency signal matrix; the enhanced beamforming processing includes time alignment and weighted superposition of spatial radio frequency signals of different array elements in each sound channel to suppress sidelobes and defocus artifacts; in the imaging plane, a spatial weight distribution of the same size as the imaging grid is formed according to the response characteristics of each position and local neighborhood, and the synthesis results of different spatial positions are selectively enhanced and suppressed so as to obtain multiple initial reconstructed images from the multisonic spatial radio frequency signal matrix; In specific implementation, the workflow of the beamforming enhancement module can be provided by the embodiments of the present invention as follows: The first-level convolutional unit receives the multisonic spatial radio frequency signal matrix, so that the first-level convolutional unit can perform local spatial convolution operations on the input feature map and compress the number of channels to extract more compact spatial radio frequency features. After the convolution operation, a spatial attention module is connected to adaptively adjust the weights of different positions according to the response differences of each spatial position and its local neighborhood, in order to compensate for micron-level time delay errors and reduce local artifacts caused by subtle time delay deviations. By receiving the outputs of the first-level convolutional unit and the spatial attention module from the second-level convolutional unit, the second-level convolutional unit effectively forms a large receptive field through two concatenated convolutional kernels, aggregating neighborhood information over a large spatial range and further compressing the number of channels to reduce redundant channels. After the second-level convolution, the spatial attention module is connected again. By weighting the response consistency between different array elements and different sound velocity channels at the same spatial position, the effect of cross-array phase is synchronized, thereby enhancing the main lobe response and suppressing defocus artifacts caused by phase inconsistency. The third-level convolutional unit receives the output of the second-level convolutional unit, further compressing the number of channels while maintaining the local convolutional form. This allows the network to reorganize spatial radio frequency features within a large effective receptive field, correcting macroscopic wavefront distortions caused by factors such as sound velocity mismatch and probe placement. Following the third-level convolution, a final spatial attention module is implemented to selectively suppress responses in background and high-noise regions, thereby highlighting coherent signals in areas containing target structures such as blood vessels. Furthermore, the stride of each convolutional stage is set to 1, and a nonlinear activation function is used to enhance feature representation capabilities.
[0023] The workflow of the spatial attention module includes: The input feature map is subjected to strip-shaped average pooling operations in the horizontal and vertical directions respectively. One-dimensional average pooling is performed in the horizontal direction to obtain the width direction feature vector representing the global response of each column; one-dimensional average pooling is performed in the vertical direction to obtain the height direction feature vector representing the global response of each row, thereby extracting global context information in the horizontal and vertical directions respectively. One-dimensional convolution operations are applied to the width-direction feature vector and the height-direction feature vector respectively to model the spatial dependency and response pattern difference between adjacent positions in their respective directions, thereby obtaining intermediate features that include directional correlation. The one-dimensional convolution outputs in both directions are group normalized to eliminate scale differences between different samples and different locations. The weight values in each direction are constrained to between 0 and 1 by the Sigmoid activation function to obtain the normalized horizontal and vertical weight vectors. The horizontal weight vector and the vertical weight vector are multiplied by an outer product to generate a bidirectional coupled two-dimensional spatial weight map. The two-dimensional spatial weight map is then multiplied element-wise with the original feature map to selectively enhance and suppress the response at each spatial location, thereby highlighting the target area with continuous structure and high contrast, and suppressing background noise and scattering artifacts.
[0024] S104: A multi-scale feature fusion and enhancement network is used to extract and fuse tubular features from multiple initial reconstructed images to obtain photoacoustic reconstructed images. The tubular feature extraction is used to highlight vascular structures that are distributed in tubular or strip-like shapes and to suppress point-like noise and block artifacts. The fusion processing is used to integrate complementary structural information from the initial results of different sound velocities and different beams, so that small blood vessels and large-diameter blood vessels are presented continuously and smoothly in the same reconstructed image.
[0025] The multi-scale feature fusion and enhancement network includes a tubular structure feature extraction module and a tubular structure multi-branch convolutional unit. The tubular structure feature extraction module is used to adaptively fit the trajectory of slender structures along different directions using a serpentine convolutional kernel. The tubular structure multi-branch convolutional unit is used to extract complementary features of small blood vessels and large blood vessels under different receptive fields.
[0026] The tubular structure feature extraction module includes an encoder module, a decoder module, and a skip connection module. The encoder module is used to downsample the input photoacoustic initial reconstructed image or feature map step by step and extract high-level semantic features; the decoder module is used to upsample the features output by the encoder step by step and restore spatial resolution; the skip connection module is used to concatenate the features of each layer of the encoder with the features of the corresponding layer of the decoder in the channel dimension to preserve high-resolution spatial detail information.
[0027] The tubular multi-branch convolutional unit is embedded in each layer of convolutional operation of the encoder and the decoder to explicitly enhance the vascular tubular structure distributed in different directions and to fuse local and global features at multiple scales.
[0028] The tubular multi-branch convolutional unit takes the same input feature map as input, and the tubular multi-branch convolutional unit includes parallel standard convolutional branches, horizontal serpentine convolutional branches, and vertical serpentine convolutional branches. The standard convolutional branch uses a standard convolutional kernel to perform local convolution operations on the input feature map, which is used to extract basic local features such as edges and textures, ensuring that the network has the ability to represent conventional image structures. The horizontal serpentine convolution branch is used to employ a serpentine convolution kernel that extends in the horizontal direction. By predicting a set of one-dimensional offsets and deforming the initial straight convolution kernel, the sampling position of the convolution kernel forms a smooth, curved trajectory in the lateral direction, which is used to produce a higher response to slender tubular structures in and around the horizontal direction. The vertical serpentine convolution branch is used to employ a serpentine convolution kernel that extends along the vertical direction. Through offset prediction and deformation operations, the sampling position of the convolution kernel forms a smooth, curved trajectory along the axial direction, which is used to generate a higher response to tubular structures that are oriented vertically and in the vicinity.
[0029] The photoacoustic image reconstruction method provided in this invention can adapt to different sensor models after acquiring the original radio frequency signal and probe parameters, without requiring a preset tissue sound velocity prior. Users only need to input the acquired radio frequency signal and can choose to provide five values covering the possible range of sound velocities in the tissue or use the system default settings. The system will automatically fuse signal features from multiple sound velocities to intelligently reconstruct a photoacoustic image with low artifacts, high vascular continuity, and accurate anatomical structure, significantly improving imaging quality and the reliability of clinical diagnosis while ensuring reconstruction efficiency.
[0030] The photoacoustic image reconstruction method provided in the embodiments of the present invention will be described in detail below.
[0031] This invention provides an intelligent photoacoustic image reconstruction method based on radio frequency (RF) signals, primarily employing a time-of-flight interpolation inversion module, a beamforming enhancement module, and a tubular structure feature extraction module. The time-of-flight inversion module maps the acquired raw RF signal back to the imaging field-of-view grid, generating features that contain spatial information while retaining the original RF signal. The beamforming enhancement module adaptively assigns weights to different spatial regions of these features for preliminary beamforming. The tubular structure feature extraction module selectively captures subtle anatomical features such as blood vessels from the beamformed features, achieving clear imaging of vascular structures. Figure 2 As shown, the method includes the following steps: Step one: A multi-velocity spatial radio frequency signal matrix is calculated from the original detection signal. The photoacoustic signal acquired by the user is input into the system, and probe parameters, imaging range, and desired resolution are set. Users can choose to customize the sound velocity or keep the default sound velocity. The time-of-flight inversion module establishes an image reconstruction grid covering the region of interest based on the user-defined imaging range and resolution. Then, the distance to each sensor is calculated based on the user-provided probe parameters, and the flight time from the reconstruction grid to each sensor is calculated based on the set sound velocity. The sampling points of the photoacoustic signal are located based on the flight time. Since the sampling points obtained from the flight time are not integers, linear interpolation is used to obtain the photoacoustic signal at this flight time, which is then mapped back into the imaging grid to obtain features that contain spatial information while retaining the original radio frequency signal.
[0032] Specifically, in step one, the photoacoustic imaging system uses a linear / rectangular array transducer to acquire the raw radio frequency (RF) signal. First, array element coordinate vectors are defined in physical space according to the element pitch to describe the spatial position of each element. Then, a three-dimensional / two-dimensional imaging grid with physical dimensions is established based on the region of interest to be imaged. Subsequently, the distance between each imaging grid point and each element is calculated, and the flight time from each grid point to each element is obtained based on multiple preset sets of sound velocity parameters. Based on the sampling position corresponding to the flight time, the corresponding amplitude is obtained from the raw RF signal through interpolation and filled into the corresponding position of the imaging grid to obtain a spatial RF signal matrix for a single sound velocity. The above process is repeated for multiple sets of sound velocities, and the matrices are stitched along the sound velocity dimension or channel dimension to obtain a multi-velocity spatial RF signal matrix.
[0033] Step two involves enhancing beamforming the spatial radio frequency matrix to obtain multiple initial reconstruction results. Then, using a beamforming enhancement module, the original radio frequency signal features are input, and a first-level convolution is used to capture microscale sound velocity perturbations and compensate for time delay differences between adjacent array elements. A spatial attention module then enhances the region with accurate sound velocity estimation. A second-level convolution senses mesoscale acoustic propagation interference, achieving cross-channel phase synchronization compensation. A third-level convolution focuses on a clear acoustic inversion region. A third-level convolution integrates large-scale wavefront distortion information to correct deep tissue reconstruction biases. Finally, spatial attention suppresses background noise and improves the integrity of the target structure.
[0034] Specifically, the multisonic spatial radio frequency signal matrix obtained in step one is subjected to enhanced beamforming processing to acquire multiple initial reconstructed images. Enhanced beamforming processing includes: time-aligning and weighted superposition of spatial radio frequency signals from different array elements within each sound velocity channel to suppress sidelobes and defocus artifacts; and forming a spatial weight distribution of the same size as the imaging grid based on the response characteristics of each location and its local neighborhood within the imaging plane, selectively enhancing and suppressing the synthesized results at different spatial locations, thereby highlighting continuous vascular structures and weakening the influence of background areas. Through the above enhanced beamforming processing, multiple corresponding initial reconstructed images can be obtained from the multisonic spatial radio frequency signal matrix.
[0035] Step three involves extracting and fusing tubular features from multiple initial reconstruction results to obtain a photoacoustic reconstructed image. Finally, the feature map generated in the previous stage is input into a multi-scale feature fusion and enhancement network to generate the photoacoustic image. The core of this network lies in embedding a tubular structure feature extraction module to solve the problem of blurry reconstruction of small blood vessels. The tubular structure feature extraction module is used to complete tubular structure feature extraction and photoacoustic reconstruction. The sub-processes of the tubular structure feature extraction module are as follows: 1. Deformation perception: For the input feature map, the offset prediction network inside the module will automatically generate a set of deformation offsets based on the geometric shape of the local features.
[0036] 2. Adaptive kernel deformation: An initial linear kernel will be superimposed with the above offset, and the dynamic deformation will be a curved curve (snake-like) that is consistent with the current local blood vessel direction.
[0037] 3. Sampling along the structure: The deformed serpentine convolution kernel closely follows the direction of the blood vessel wall, and pixel values are sampled from the feature map along its trajectory using bilinear interpolation. This process can accurately capture the continuous direction and branching morphology of the blood vessel, avoiding irrelevant background noise introduced by traditional square convolution kernels.
[0038] 4) Feature-weighted output: The sampled feature sequence is weighted and summed with the trainable weights to output a new feature map that can significantly enhance the response of the tubular structure.
[0039] The decoder of this multi-scale feature fusion and enhancement network gradually restores spatial resolution through upsampling and skip connections, and focuses attention on the vascular contours and tissue boundaries enhanced in the previous stage, ultimately outputting a high-quality photoacoustic image with continuous vascular structure, clear edges, and significant artifact suppression.
[0040] Specifically, tubular feature extraction and fusion processing are performed on multiple initial reconstructed images obtained in step two to obtain the final photoacoustic reconstructed image. The tubular feature extraction is used to highlight vascular structures distributed in tubular or strip-like patterns and suppress point-like noise and blocky artifacts. The fusion processing is used to integrate complementary structural information from the initial results synthesized with different sound velocities and beams, enabling small and large-diameter vessels to be presented continuously and smoothly in the same reconstructed image. Through the above tubular feature extraction and fusion processing, the clarity of vessel boundaries can be further enhanced and the overall quality of the reconstructed image can be improved.
[0041] This invention employs a multi-velocity spatial rearrangement, enhanced beamforming, and tubular feature extraction and fusion processing on the original photoacoustic radio frequency signal to progressively suppress artifacts introduced by sound velocity mismatch and spatial noise, thereby enhancing the imaging performance of vascular structures. Specifically, the construction of the multi-velocity spatial radio frequency matrix explicitly characterizes the spatial information of the radio frequency signal in the imaging domain, providing a unified spatial representation for subsequent coherent synthesis. Enhanced beamforming adaptively adjusts the contribution of each channel and spatial position in the synthesis process based on the consistency of the responses of the multi-velocity channels and each array element on the imaging grid and their contrast with the local background, thus simultaneously suppressing sidelobes and defocus artifacts in both the channel and spatial dimensions, obtaining multiple initial reconstruction results with a clear overall structure. Furthermore, tubular feature extraction and fusion utilizes the prior knowledge that photoacoustic vessels are distributed in a slender, continuous tubular shape in the image to selectively enhance the tubular response in the initial reconstruction results and synergistically fuses the multi-velocity initial reconstruction results, thereby achieving high-contrast, continuous reconstruction of the photoacoustic vascular network, improving the visibility of small vessels and overall imaging quality.
[0042] Based on the above embodiments, as an optional embodiment, the process of calculating the radio frequency signal from the sensor coordinates and the field of view grid and interpolating it to the imaging domain in step one includes: determining the array element coordinate vector according to the spacing between the array elements of the photoacoustic imaging probe, and using the array element coordinate vector to characterize the position of each array element in physical space; establishing an imaging grid with a fixed spatial resolution in physical space according to a preset imaging region of interest, and obtaining the spatial coordinates of each grid point in the imaging grid; calculating the Euclidean distance between any grid point in the imaging grid and each array element for each set of preset sound velocity parameters, and calculating the flight time from the grid point to each array element based on the distance and the corresponding sound velocity; according to The relationship between flight time and sampling time interval is used to map flight time to the sampling point position in the original radio frequency signal. When the sampling point position is not an integer, the corresponding amplitude is obtained from the original radio frequency signal by interpolation. The interpolated amplitude is filled into the imaging grid position corresponding to the grid point and array element position to obtain the spatial radio frequency signal matrix under the assumption of sound speed. According to the processing method of the first group of sound speeds to the last group of sound speeds, the above flight time calculation and interpolation filling process is repeated for all sound speed parameters to obtain multiple spatial radio frequency signal matrices corresponding to sound speeds. The spatial radio frequency signal matrices are then spliced in the channel dimension to obtain a multi-sound speed spatial radio frequency signal matrix.
[0043] Specifically, Figure 3 This is a schematic diagram illustrating the process of calculating a multi-velocity spatial radio frequency signal matrix from the original detection signal, as provided in an embodiment of the present invention. Figure 3 As shown, (1) Establish the physical size grid of the imaging area and the physical coordinates of the sensor. Based on the physical size of the imaging area and the preset imaging resolution, establish a regular grid of the imaging area in the physical space to obtain the physical coordinates of each grid point; based on the sensor element spacing and array arrangement, determine the physical coordinates of each sensor element in the same coordinate system, thereby obtaining the "physical size grid of the imaging area" and the "physical coordinates of the sensor".
[0044] (2) Calculate the physical distance from the grid point to the sensor. Based on the geometric relationship between the grid point coordinates and the sensor coordinates, calculate the straight-line distance from each grid point to each sensor array element to obtain the "physical distance from the grid point to the sensor".
[0045] (3) Set multiple sets of sound speeds and calculate the flight time from multiple grid points to the sensor. Multiple sets of sound speed parameters are preset, and each set of sound speeds corresponds to one propagation assumption. For any physical distance d from any grid point to any sensor and the current sound speed c, the corresponding flight time is calculated as d / c. By performing the above calculation on all grid points, all sensors, and all sound speed sets, the "flight time from multiple grid points to the sensor" is obtained.
[0046] (4) Based on the sensor sampling frequency, the flight time is converted into the sampling time point of the radio frequency signal. Based on the sensor sampling frequency, each flight time is converted into the sampling time position in the radio frequency signal; the converted sampling time position is usually a non-integer, forming "multiple grid points to the sampling time point of the sensor's radio frequency signal".
[0047] (5) Obtain the corresponding amplitude from the photoacoustic signal based on the signal sampling time point. For each non-integer sampling time point, take the adjacent integer sampling points before and after it as reference, and use interpolation (e.g., linear interpolation) to obtain the corresponding amplitude from the original photoacoustic radio frequency signal, so as to achieve approximate sampling of continuous time positions and obtain the photoacoustic signal amplitude of each grid point under different sound speeds and different sensors.
[0048] (6) Construct a multi-velocity spatial radio frequency signal matrix. Finally, the interpolated amplitudes obtained in step (5) are organized and stored according to grid point positions, sensor numbers, and sound velocity group indices to construct a multi-velocity spatial radio frequency signal matrix in the imaging domain. This matrix corresponds one-to-one with the imaging grid in spatial coordinates and simultaneously characterizes the radio frequency response under multi-element and multi-velocity conditions in the channel dimension, providing a unified and accurate spatial input for subsequent enhanced beamforming and vascular structure reconstruction.
[0049] Based on the above embodiments, as an optional embodiment, the multi-sonic spatial radio frequency signal matrix is subjected to enhanced beamforming processing to obtain multiple initial reconstructed images. This includes: inputting the multi-sonic spatial radio frequency signal matrix into a multi-level convolutional beamforming enhancement module, in which the responses of different sound speed channels and different array elements on the imaging grid are adaptively coherently synthesized and suppressed through stepwise local convolution feature extraction and spatial weighting, thereby obtaining an initial reconstruction result with fewer artifacts and higher contrast.
[0050] Specifically, Figure 4 This is a schematic flowchart of the beamforming enhancement module provided in an embodiment of the present invention. Figure 4 As shown, it includes: (1) The multisonic spatial radio frequency signal matrix is input to the first-level convolutional unit. The first-level convolutional unit performs local spatial convolution operation on the input feature map and compresses the number of channels to extract more compact spatial radio frequency features. After the convolution operation, a spatial attention module is connected to adaptively adjust the weights of different positions according to the response differences of each spatial position and its local neighborhood, in order to compensate for micron-level time delay errors and reduce local artifacts caused by subtle time delay deviations.
[0051] (2) The outputs of the first-level convolutional unit and the spatial attention module are fed into the second-level convolutional unit. The second-level convolutional unit forms a larger receptive field by using two concatenated convolutional kernels, which aggregates neighborhood information in a larger spatial range and further compresses the number of channels to reduce redundant channels. After the second-level convolution, the spatial attention module is connected again. By weighting the response consistency of different array elements and different sound speed channels at the same spatial position, it plays a role in synchronizing the phase across array elements, thereby enhancing the main lobe response and suppressing defocus artifacts caused by phase inconsistency.
[0052] (3) The output of the second stage is fed into the third stage convolutional unit. While maintaining the local convolutional form, the third stage convolutional unit further compresses the number of channels, enabling the network to reorganize spatial radio frequency features within a larger effective receptive field to correct macroscopic wavefront distortions caused by factors such as sound velocity mismatch and probe placement. A final spatial attention module is set after the third stage convolution to selectively suppress responses in background and high-noise regions, thereby highlighting coherent signals in areas containing target structures such as blood vessels. The stride of each of the above convolutional stages is set to 1 and used in conjunction with a nonlinear activation function (such as ReLU) to enhance feature representation capabilities.
[0053] Specifically, Figure 5 This is a schematic diagram of the spatial attention module. (Example) Figure 5 As shown, it includes: (1) Perform strip-shaped average pooling operations on the input feature map in the horizontal and vertical directions respectively: perform one-dimensional average pooling along the horizontal direction to obtain the width direction feature vector representing the global response of each column; perform one-dimensional average pooling along the vertical direction to obtain the height direction feature vector representing the global response of each row, thereby extracting global context information in the horizontal and vertical directions respectively.
[0054] (2) Apply one-dimensional convolution operations to the width and height feature vectors respectively to model the spatial dependency and response pattern differences between adjacent positions in their respective directions, and obtain intermediate features containing directional correlation.
[0055] (3) The one-dimensional convolution outputs in the two directions are group normalized to eliminate the scale difference between different samples and different positions. The weight values in each direction are constrained to between 0 and 1 by the Sigmoid activation function to obtain the normalized horizontal weight vector and vertical weight vector.
[0056] (4) Perform an outer product operation on the horizontal and vertical weight vectors to generate a two-dimensional spatial weight map with bidirectional coupling; multiply the two-dimensional spatial weight map with the original feature map element by element to selectively enhance and suppress the response at each spatial location, thereby highlighting the target area with continuous structure and high contrast, and suppressing background noise and scattering artifacts.
[0057] This invention, through the introduction of a multi-level convolutional beamforming enhancement module onto a multi-velocity spatial radio frequency signal matrix, combined with a three-layer progressively compressed channel, a gradually expanding receptive field local convolutional structure, and a spatial attention module based on global context in the horizontal and vertical directions, achieves joint adaptive weighting in both the channel and spatial dimensions. This module effectively compensates for micro-delay errors, synchronizes phase across array elements, and suppresses background noise and defocus artifacts. Thus, even under conditions of viewpoint and velocity mismatch, it can still obtain an initial reconstructed image with high signal-to-noise ratio and good contrast, providing high-quality input for subsequent vascular tubular feature extraction and fusion.
[0058] This invention enhances the vascular structure in photoacoustic reconstructed images by embedding a tubular structure feature extraction module into the U-Net encoder-decoder network and combining it with a multi-scale convolutional feature extraction and fusion mechanism. Specifically, the tubular structure feature extraction module adaptively fits the trajectory of slender structures along different directions using serpentine convolutional kernels. The multi-scale convolution module extracts complementary features of small and large blood vessels through parallel convolutional branches in different receptive fields. The U-Net encoder-decoder network maintains the synergy of spatial resolution and semantic information during multi-level downsampling and upsampling, resulting in better continuity of vascular structures and clearer edges in the final reconstructed photoacoustic image, simultaneously balancing the saliency of blood vessels of different diameters and background suppression capabilities.
[0059] like Figure 6 As shown, Figure 6 This is a schematic diagram of a multi-scale feature fusion network architecture with an embedded tubular structure feature extraction module provided in an embodiment of the present invention. The network adopts a U-Net encoder-decoder structure and includes: an encoder module 61, a decoder module 62, a skip connection module 63, and a tubular structure multi-branch convolutional unit 64 uniformly used in each layer of the encoder and decoder.
[0060] The encoder module 61 downsamples the input photoacoustic initial reconstructed image or feature map step by step and extracts high-level semantic features; the decoder module 62 upsamples the features output by the encoder step by step and restores the spatial resolution; the skip connection module 63 concatenates the features of each layer of the encoder with the features of the corresponding layer of the decoder in the channel dimension to preserve high-resolution spatial detail information; the tubular structure multi-branch convolutional unit 64 is embedded in each layer of convolution operation of the encoder and decoder to explicitly enhance the vascular tubular structure distributed in different directions and fuse local and global features at multiple scales.
[0061] Specifically, in any layer of U-Net, the original single-path convolution is replaced by a tubular multi-branch convolutional unit 64. This convolutional unit takes the same input feature map as input and has three parallel convolutional branches, namely: 1. Standard Convolution Branch: The standard convolution kernel is used to perform local convolution operations on the input feature map to extract basic local features such as edges and textures, ensuring that the network has the ability to represent conventional image structures. 2. Horizontal serpentine convolution branch: A serpentine convolution kernel extending in the horizontal direction is used. By predicting a set of one-dimensional offsets and deforming the initial straight convolution kernel, the sampling position of the convolution kernel forms a smooth curved trajectory in the horizontal direction, which is used to produce a higher response to slender tubular structures in and around the horizontal direction. 3. Vertical serpentine convolution branch: Using a serpentine convolution kernel that extends along the vertical direction, through similar offset prediction and deformation operations, the sampling position of the convolution kernel forms a smooth, curved trajectory along the axial direction, which is used to produce a higher response to tubular structures in the vertical direction and its vicinity.
[0062] Three branches operate in parallel on the input feature map at the same level, each outputting a set of intermediate feature maps. Then, the three sets of feature maps are concatenated along the channel dimension to obtain a fused feature map with three times the original number of channels. This fused feature map is then input into a fusion convolutional layer with the same number of channels. This fusion convolutional layer adaptively learns the complementary relationship between the standard convolutional branch and the two serpentine convolutional branches, weightedly integrating multi-directional and multi-shaped tubular structure features to generate the final output feature of this layer, which serves as the convolutional result of this layer in U-Net.
[0063] like Figure 7 As shown, Figure 7 This is a schematic diagram of the internal structure of the tubular structure feature extraction module. The module first predicts the deformation offset of several sampling points using a convolutional layer and constrains the offset range using the Tanh function. Based on this, an initial linear convolutional kernel (e.g., a horizontal or vertical convolutional kernel) is superimposed on the offset, and the discrete offset sequence is smoothed into a continuous serpentine sampling curve through cumulative summation. Subsequently, bilinear interpolation sampling is performed on the input feature map along this serpentine curve, and then multiply-accumulate with trainable convolutional weights to obtain an enhanced feature map with high response to tubular structures in the specified direction.
[0064] In the encoding path, input features are processed sequentially from top to bottom through downsampling and the aforementioned tubular multi-branch convolutional units, achieving the construction of a feature pyramid with progressively increasing channel number and progressively decreasing spatial resolution. In the decoding path, features are sequentially upsampled from bottom to top and concatenated with features from the corresponding layer of the encoding path via skip connections along the channel dimension. This is then refined using the same multi-branch convolutional units, achieving layer-by-layer fusion of multi-scale information and gradual restoration of spatial resolution. Through this self-encoding-decoding multi-scale structure and the use of uniformly applied tubular multi-branch convolutional units in each layer, this embodiment of the invention can continuously enhance the tubular structures of blood vessels at different scales and directions throughout the entire network depth, ensuring that small blood vessels are not submerged during the downsampling process while maintaining smooth and continuous boundaries for large-diameter blood vessels.
[0065] In this embodiment of the invention, the horizontal serpentine convolution branch and the vertical serpentine convolution branch adopt serpentine sampling trajectories in the horizontal and vertical directions, respectively, so that the network can perform directional enhancement of vascular structures in different directions.
[0066] In summary, the photoacoustic image reconstruction method provided in this application effectively solves the problems of photoacoustic image artifacts and structural blurring caused by uneven sound velocity and limited viewing angle. It also possesses strong generalization ability; a single model can be adapted to various probe parameters, ensuring both physical rationality and reconstruction efficiency, as well as clinical applicability. See Figure 8 This application embodiment can also provide a photoacoustic image reconstruction system, such as... Figure 8 As shown, the system for performing the above-described photoacoustic image reconstruction method may include: The raw radio frequency signal acquisition unit 801 is used to acquire the raw radio frequency signal collected by the array transducer and determine the probe parameters, imaging range, target resolution and target sound velocity. The time-of-flight interpolation inversion module unit 802 is used to define the array element coordinate vector in physical space based on the array element spacing using the time-of-flight inversion module; establish an image reconstruction grid covering the region of interest based on the imaging range and the target resolution; calculate the distance between each imaging grid point and each array element; and calculate the flight time from each grid point to each array element based on the target sound speed; obtain the corresponding amplitude from the original radio frequency signal through interpolation based on the sampling position corresponding to the flight time; and fill the amplitude into the position corresponding to the image reconstruction grid to obtain a spatial radio frequency signal matrix that contains spatial information for a single sound speed while retaining the original radio frequency signal; repeat the processing for multiple sets of sound speeds and splice them along the sound speed dimension or channel dimension to obtain a multi-sound speed spatial radio frequency signal matrix; The beamforming enhancement unit 803 is used to perform enhanced beamforming processing on the multisonic spatial radio frequency signal matrix using the beamforming enhancement module. The enhanced beamforming processing includes time alignment and weighted superposition of spatial radio frequency signals of different array elements in each sound channel to suppress sidelobes and defocus artifacts; and forming a spatial weight distribution of the same size as the imaging grid in the imaging plane according to the response characteristics of each position and local neighborhood, selectively enhancing and suppressing the synthesis results at different spatial positions, so as to obtain multiple corresponding initial reconstructed images from the multisonic spatial radio frequency signal matrix. The tubular structure feature extraction unit 804 is used to extract and fuse tubular features from multiple initial reconstructed images using a multi-scale feature fusion and enhancement network to obtain a target photoacoustic reconstructed image. The tubular feature extraction is used to highlight vascular structures that are distributed in tubular or strip-like shapes and to suppress point-like noise and block artifacts. The fusion processing is used to integrate complementary structural information from the initial results of different sound velocities and different beams, so that small blood vessels and large-diameter blood vessels are presented continuously and smoothly in the same reconstructed image.
[0067] This application embodiment can also provide a photoacoustic image reconstruction device, the device including a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the steps of the photoacoustic image reconstruction method described above according to the instructions in the program code.
[0068] Figure 9 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention. Figure 9 As shown, the electronic device may include: a processor 10, a memory 11, a communications interface 12, and a bus 13, wherein the processor 10, the memory 11, and the communications interface 12 communicate with each other through the bus 13.
[0069] The communication interface 12 is used for data interaction between the electronic device and the photoacoustic imaging probe and data acquisition system. For example, it can receive information such as multi-channel photoacoustic radio frequency signals, sensor array element coordinates, sampling frequency, preset sound velocity parameters and imaging area grid parameters, and send the photoacoustic reconstructed image to the display terminal or host computer after reconstruction.
[0070] The memory 11 stores program code and data required for the operation of the electronic device, including but not limited to: program instructions for executing the method of the present invention, raw photoacoustic radio frequency data, a multi-velocity spatial radio frequency signal matrix, initial reconstruction results obtained by enhanced beamforming, intermediate feature maps, and the final photoacoustic reconstructed image. The processor 10 can call the logical instructions in the memory 11 to execute the following method steps: 1. Receive multi-channel photoacoustic radio frequency signals collected by the photoacoustic imaging probe from the communication interface 12, as well as the corresponding sensor array element coordinates, sampling frequency, preset sound velocity group and physical size grid parameters of the imaging area, and write the above data into the memory 11; 2. Based on the sensor array element coordinates and the grid division within the imaging field of view, calculate the geometric distance from each imaging grid point to each array element, obtain the corresponding flight time under multiple preset sound velocity parameters, and map the flight time to the sampling time position of the radio frequency signal according to the signal sampling frequency; for non-integer sampling time positions, use interpolation to obtain the corresponding amplitude from the original photoacoustic radio frequency signal stored in memory 11, organize the interpolation results into a multi-sound velocity spatial radio frequency signal matrix according to grid points, array elements and sound velocity groups, and temporarily store the matrix in memory 11; 3. Call the enhanced beamforming module to perform convolution processing and spatial weighting on the multisonic spatial radio frequency signal matrix in memory 11: The processor 10 inputs the multisonic spatial radio frequency matrix into the beamforming enhancement network composed of multi-level convolution and spatial attention. By suppressing sidelobes and background noise through multi-level local convolution and spatial weight modulation, one or more sets of initial reconstruction results are obtained, and these initial reconstruction images or feature maps are written into memory 11. 4. The multi-scale feature fusion network with embedded tubular structure feature extraction module is invoked to further reconstruct the initial reconstruction results. Specifically, the processor 10 inputs the initial reconstruction results into a neural network with a U-Net encoder-decoder structure. In the convolutional units of each layer of the encoder and decoder, tubular structure features are extracted through parallel convolution operations of standard convolution branches, horizontal serpentine convolution branches, and vertical serpentine convolution branches. After channel stitching, features are integrated through fusion convolution. At the same time, skip connections are combined during multi-scale downsampling and upsampling to preserve high-resolution detail information. Finally, a photoacoustic reconstructed image with spatial resolution restored to a preset size (e.g., 256×256) is output, and the photoacoustic reconstructed image is stored in the memory 11. 5. The photoacoustic reconstructed image stored in the memory 11 is sent to the display device or archiving system via the communication interface 12 for use by doctors or subsequent analysis software.
[0071] In this embodiment of the invention, the memory 11 may include high-speed random access memory (RAM) and / or non-volatile memory (such as read-only memory ROM, flash memory, hard disk, etc.) for storing program code, network parameters, and raw photoacoustic radio frequency data, multisonic spatial radio frequency signal matrix, enhanced beamforming results, and final reconstructed images, respectively. During operation, the processor 10 can cache frequently accessed intermediate results (such as multisonic matrices and intermediate network feature maps) in the high-speed storage area to improve overall processing efficiency.
[0072] Furthermore, the logical instructions in the aforementioned memory 11 can be implemented as software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a computer-readable storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, workstation, or dedicated photoacoustic imaging system host, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention.
[0073] The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, a USB flash drive, a portable hard drive, or other media capable of storing program code. The program code is used to execute the steps of the photoacoustic image reconstruction method described above.
[0074] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0075] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0076] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, for system or system embodiments, since they are fundamentally similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. Components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0077] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A photoacoustic image reconstruction method, characterized in that, The method includes: Acquire the raw radio frequency signal collected by the array transducer and determine the probe parameters, imaging range, target resolution, and target sound velocity; The time-of-flight inversion module defines the array element coordinate vector in physical space based on the element spacing. Based on the imaging range and target resolution, an image reconstruction grid covering the region of interest is established. The distance between each imaging grid point and each array element is calculated, and the flight time from each grid point to each array element is obtained based on the target sound velocity. According to the sampling position corresponding to the flight time, the corresponding amplitude is obtained from the original radio frequency signal through interpolation, and the amplitude is filled into the corresponding position of the image reconstruction grid to obtain a spatial radio frequency signal matrix that contains spatial information for a single sound velocity while retaining the original radio frequency signal. Multiple sets of sound velocities are processed repeatedly and spliced along the sound velocity dimension or channel dimension to obtain a multi-sound velocity spatial radio frequency signal matrix. The multisonic spatial radio frequency signal matrix is subjected to enhanced beamforming processing using a beamforming enhancement module. The enhanced beamforming processing includes time alignment and weighted superposition of spatial radio frequency signals of different array elements in each sound speed channel to suppress sidelobes and defocus artifacts. In the imaging plane, a spatial weight distribution of the same size as the imaging grid is formed according to the response characteristics of each position and local neighborhood. The synthesis results at different spatial positions are selectively enhanced and suppressed to obtain multiple initial reconstructed images from the multisonic spatial radio frequency signal matrix. A multi-scale feature fusion and enhancement network is used to extract and fuse tubular features from multiple initial reconstructed images to obtain photoacoustic reconstructed images. The tubular feature extraction is used to highlight vascular structures that are distributed in tubular or strip-like patterns and to suppress point-like noise and block artifacts. The fusion processing is used to integrate complementary structural information from the initial results of different sound velocities and different beams, so that small blood vessels and large-diameter blood vessels are presented continuously and smoothly in the same reconstructed image.
2. The photoacoustic image reconstruction method according to claim 1, characterized in that, The step of defining the array element coordinate vector in physical space based on the array element spacing includes: establishing a regular grid of the imaging area in physical space based on the physical size of the imaging area and the preset imaging resolution, and obtaining the physical coordinates of each grid point; and determining the physical coordinates of each array element in the same coordinate system based on the array element spacing and array arrangement.
3. The photoacoustic image reconstruction method according to claim 1, characterized in that, The workflow of the beamforming enhancement module includes: The first-level convolutional unit receives the multisonic spatial radio frequency signal matrix, so that the first-level convolutional unit can perform local spatial convolution operations on the input feature map and compress the number of channels to extract more compact spatial radio frequency features. After the convolution operation, a spatial attention module is connected to adaptively adjust the weights of different positions according to the response differences of each spatial position and its local neighborhood, in order to compensate for micron-level time delay errors and reduce local artifacts caused by subtle time delay deviations. By receiving the outputs of the first-level convolutional unit and the spatial attention module from the second-level convolutional unit, the second-level convolutional unit effectively forms a large receptive field through two concatenated convolutional kernels, aggregating neighborhood information over a large spatial range and further compressing the number of channels to reduce redundant channels. After the second-level convolution, the spatial attention module is connected again. By weighting the response consistency between different array elements and different sound velocity channels at the same spatial position, the effect of cross-array phase is synchronized, thereby enhancing the main lobe response and suppressing defocus artifacts caused by phase inconsistency. The third-level convolutional unit receives the output of the second-level convolutional unit, which further compresses the number of channels while maintaining the local convolutional form. This allows the network to reorganize the spatial radio frequency characteristics within a large effective receptive field, correcting macroscopic wavefront distortions caused by factors such as sound velocity mismatch and probe placement. After the third-level convolution, the final spatial attention module is set to selectively suppress the responses in background and high-noise regions, thereby highlighting the coherent signals in the regions where target structures such as blood vessels are located.
4. The photoacoustic image reconstruction method according to claim 3, characterized in that, The stride of each convolutional layer is set to 1, and a non-linear activation function is used in conjunction to enhance the feature representation capability.
5. The photoacoustic image reconstruction method according to claim 3, characterized in that, The workflow of the spatial attention module includes: The input feature map is subjected to strip-shaped average pooling operations in the horizontal and vertical directions respectively. One-dimensional average pooling is performed in the horizontal direction to obtain the width direction feature vector representing the global response of each column; one-dimensional average pooling is performed in the vertical direction to obtain the height direction feature vector representing the global response of each row, thereby extracting global context information in the horizontal and vertical directions respectively. One-dimensional convolution operations are applied to the width-direction feature vector and the height-direction feature vector respectively to model the spatial dependency and response pattern difference between adjacent positions in their respective directions, thereby obtaining intermediate features that include directional correlation. The one-dimensional convolution outputs in both directions are group normalized to eliminate scale differences between different samples and different locations. The weight values in each direction are constrained to between 0 and 1 by the Sigmoid activation function to obtain the normalized horizontal and vertical weight vectors. The horizontal weight vector and the vertical weight vector are multiplied by an outer product to generate a bidirectional coupled two-dimensional spatial weight map. The two-dimensional spatial weight map is then multiplied element-wise with the original feature map to selectively enhance and suppress the response at each spatial location, thereby highlighting the target area with continuous structure and high contrast, and suppressing background noise and scattering artifacts.
6. The photoacoustic image reconstruction method according to claim 1, characterized in that, The multi-scale feature fusion and enhancement network includes a tubular structure feature extraction module and a tubular structure multi-branch convolutional unit. The tubular structure feature extraction module is used to adaptively fit the trajectory of slender structures along different directions using a serpentine convolutional kernel. The tubular structure multi-branch convolutional unit is used to extract complementary features of small blood vessels and large blood vessels under different receptive fields.
7. The photoacoustic image reconstruction method according to claim 6, characterized in that, The tubular structure feature extraction module includes an encoder module, a decoder module, and a skip connection module. The encoder module is used to downsample the input photoacoustic initial reconstructed image or feature map step by step and extract high-level semantic features; the decoder module is used to upsample the features output by the encoder step by step and restore spatial resolution; the skip connection module is used to concatenate the features of each layer of the encoder with the features of the corresponding layer of the decoder in the channel dimension to preserve high-resolution spatial detail information.
8. The photoacoustic image reconstruction method according to claim 7, characterized in that, The tubular multi-branch convolutional unit is embedded in each layer of convolutional operation of the encoder and the decoder to explicitly enhance the vascular tubular structure distributed in different directions and to fuse local and global features at multiple scales.
9. The photoacoustic image reconstruction method according to claim 8, characterized in that, The tubular multi-branch convolutional unit takes the same input feature map as input, and the tubular multi-branch convolutional unit includes parallel standard convolutional branches, horizontal serpentine convolutional branches, and vertical serpentine convolutional branches. The standard convolutional branch uses a standard convolutional kernel to perform local convolution operations on the input feature map, which is used to extract basic local features such as edges and textures, ensuring that the network has the ability to represent conventional image structures. The horizontal serpentine convolution branch is used to employ a serpentine convolution kernel that extends in the horizontal direction. By predicting a set of one-dimensional offsets and deforming the initial straight convolution kernel, the sampling position of the convolution kernel forms a smooth, curved trajectory in the lateral direction, which is used to produce a higher response to slender tubular structures in and around the horizontal direction. The vertical serpentine convolution branch is used to employ a serpentine convolution kernel that extends along the vertical direction. Through offset prediction and deformation operations, the sampling position of the convolution kernel forms a smooth, curved trajectory along the axial direction, which is used to generate a higher response to tubular structures that are oriented vertically and in the vicinity.
10. A photoacoustic image reconstruction system, characterized in that, The system is used to perform the photoacoustic image reconstruction method according to any one of claims 1 to 9, the system comprising: The raw radio frequency signal acquisition unit is used to acquire the raw radio frequency signal collected by the array transducer and determine the probe parameters, imaging range, target resolution and target sound velocity. The time-of-flight interpolation and inversion module unit is used to define array element coordinate vectors in physical space based on the array element spacing using the time-of-flight inversion module; establish an image reconstruction grid covering the region of interest based on the imaging range and the target resolution; calculate the distance between each imaging grid point and each array element; and calculate the flight time from each grid point to each array element based on the target sound velocity; obtain the corresponding amplitude from the original radio frequency signal through interpolation based on the sampling position corresponding to the flight time; and fill the amplitude into the position corresponding to the image reconstruction grid to obtain a spatial radio frequency signal matrix that contains spatial information for a single sound velocity while retaining the original radio frequency signal; repeat the processing for multiple sets of sound velocities and stitch them together along the sound velocity dimension or channel dimension to obtain a multi-sound velocity spatial radio frequency signal matrix; A beamforming enhancement unit is used to perform enhanced beamforming processing on the multisonic spatial radio frequency signal matrix using a beamforming enhancement module. The enhanced beamforming processing includes time alignment and weighted superposition of spatial radio frequency signals of different array elements in each sound speed channel to suppress sidelobes and defocus artifacts. In the imaging plane, a spatial weight distribution of the same size as the imaging grid is formed according to the response characteristics of each position and local neighborhood. The synthesis results at different spatial positions are selectively enhanced and suppressed to obtain multiple initial reconstructed images from the multisonic spatial radio frequency signal matrix. The tubular structure feature extraction unit is used to extract and fuse tubular features from multiple initial reconstructed images using a multi-scale feature fusion and enhancement network to obtain a target photoacoustic reconstructed image. The tubular feature extraction is used to highlight vascular structures that are distributed in tubular or strip-like shapes and to suppress point-like noise and block artifacts. The fusion processing is used to integrate complementary structural information from the initial results of different sound velocities and different beams, so that small blood vessels and large-diameter blood vessels are presented continuously and smoothly in the same reconstructed image.